A method for constructing an endocrine simplified transcriptome based on zebrafish embryo system and application thereof

By constructing a simplified endocrine transcriptome of a zebrafish embryonic system, the problems of low throughput and insufficient mechanism analysis in endocrine disruptor screening were solved, enabling efficient and low-cost endocrine disruptor screening and risk assessment, and providing a comprehensive understanding of endocrine disruptors and high-throughput screening capabilities.

CN116189762BActive Publication Date: 2026-05-01NANJING UNIV
View PDF -1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV
Filing Date
2023-02-28
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Current endocrine disruptor screening methods suffer from low throughput, insufficient mechanism analysis, and high costs, making it difficult to achieve high-throughput screening and effective evaluation.

Method used

We constructed a simplified endocrine transcriptome based on a zebrafish embryonic system. By collecting endocrine-related biological pathways, we constructed endocrine-related harmful outcome pathways, classified endocrine interference mechanisms, assessed gene centrality, redundancy, and information richness, constructed a simplified endocrine gene set, used an improved BORDA counting method to score gene criticality, simplified the number of genes in the gene set, and constructed a simplified endocrine transcriptome for the identification of endocrine disruptors.

Benefits of technology

It significantly reduces the cost of endocrine disruptor screening, improves the high-throughput screening and mechanism analysis capabilities of endocrine disruptors, provides risk management and understanding of endocrine disruptors, comprehensively covers endocrine-related pathways, and supports the prediction of harmful outcomes at a high biological level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189762B_ABST
    Figure CN116189762B_ABST
Patent Text Reader

Abstract

The application discloses a method for constructing an endocrine simplified transcriptome based on a zebrafish embryo system and application thereof, and belongs to the field of high-throughput identification and screening of endocrine disruptors. By using the BORDA counting method, the evidence scores of three dimensions of gene centrality, information richness and redundancy degree are integrated, a specific endocrine interference simplified gene set is constructed, and the sequencing depth required for endocrine interference screening of each substance is greatly reduced, thereby reducing the test cost and improving the screening throughput of endocrine disruptors. The application provides help and support for rapid screening, mechanism analysis and harmful outcome prediction of new pollutants related to endocrine interference in China and the world.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of high-throughput identification and screening of endocrine disruptors, and more specifically, relates to a method and application for constructing a simplified endocrine transcriptome based on a zebrafish embryo system. Background Technology

[0002] The endocrine system is a crucial component of living organisms, regulating normal growth, metabolism, and reproduction. Endocrine-disrupting chemicals (EDCs) can inadvertently disrupt this complex system, leading to a range of adverse consequences such as cognitive decline, cancer, and metabolic disorders. However, existing EDC screening programs by agencies like the European Chemicals Agency (ECHA) and the U.S. Environmental Protection Agency (EPA) often only address a limited number of endocrine-disrupting mechanisms and suffer from low throughput, high cost, and the unreliability of assessment methods. Therefore, there is an urgent need to develop new methods and frameworks for high-throughput screening and mechanistic analysis of current EDCs.

[0003] Whole transcriptomics can capture differential changes in gene expression levels in test organisms, elucidate the mechanisms of interference at the molecular level, and predict harmful outcomes at a higher biological level, possessing the ability to perform high-throughput, parallel testing of a large number of biological endpoints. However, current whole transcriptomics tests remain expensive. Furthermore, due to the complexity of endocrine disruption mechanisms, a single cell line is insufficient to characterize the full range of endocrine disruption mechanisms. Zebrafish, as a model animal with high genetic homology to humans, offers advantages such as in vivo modeling, transparency, low cost, and high throughput, making it an excellent system for screening endocrine disruptors. Summary of the Invention

[0004] 1. The problem to be solved

[0005] To address the issues of low throughput and insufficient mechanism analysis in existing endocrine disruptor screening methods, this invention provides a method and application for constructing a simplified endocrine transcriptome based on a zebrafish embryonic system. By constructing a specific simplified gene set of endocrine disruptors, this invention enables high-throughput screening and discovery of the endocrine effects of compounds. It significantly reduces costs while preserving transcriptomic information and effectively improves managers' control and understanding of endocrine disruptors.

[0006] 2. Technical Solution

[0007] To solve the above problems, the technical solution adopted by the present invention is as follows:

[0008] This invention discloses a method for constructing a simplified endocrine transcriptome based on a zebrafish embryonic system. The method includes collecting endocrine-related biological pathways, constructing endocrine-related harmful outcome pathways, and tagging these pathways with endocrine interference mechanisms to obtain a simplified zebrafish homologous gene set. Then, the gene centrality, redundancy, and information richness of each gene in the simplified zebrafish homologous gene set are evaluated, and gene criticality is scored. The criticality scores are then sorted in descending order. Multiple genes are sequentially superimposed according to their criticality scores to enrich biological pathways. The number of genes corresponding to the plateau phase of pathway enrichment is extracted to construct the simplified endocrine gene set based on the zebrafish embryonic system.

[0009] Preferably, the method for constructing a simplified endocrine transcriptome based on a zebrafish embryonic system according to the present invention specifically includes the following steps:

[0010] S10. Collect information on endocrine-related biological pathways:

[0011] Endocrine-related biological pathways were collected from the MSigDB, HALLMARK, and KEGG databases.

[0012] S20. Constructing endocrine-related harmful outcome pathways:

[0013] Based on literature data and AOPWIKI data, AOP data of endocrine-related harmful outcome pathways are extracted, and the AOP data is mapped with the biological pathways obtained in step S10.

[0014] S30, Endocrine Disruption Mechanism Label Classification:

[0015] Based on the characteristics of endocrine disruptors, the biological pathways obtained in steps S10 and S20 are matched and labeled, and the biological hierarchy and related interference targets and target organs described by the biological pathways obtained in steps S10 and S20 are matched and labeled.

[0016] S40. Organize gene information related to biological pathways:

[0017] Extract all genes of the biological pathways obtained in steps S10 and S20, and map all genes homologously to zebrafish species based on ENTREZ ID to obtain the gene set to be simplified;

[0018] S50, Gene Centrality Assessment:

[0019] Extract the connectivity values ​​between genes in the biological mesh obtained in step S40 of the gene set to be simplified, convert the genes into nodes, use the connectivity values ​​as edge weights between gene nodes to construct an undirected graph model, and calculate the centrality PAGERANK value of each gene in the network using the PAGERANK algorithm with restart.

[0020] S60, Assessment of gene redundancy and information richness:

[0021] Redundancy is assessed based on the total number of genes involved in the biological pathways of each gene, and information richness is assessed based on the number of biological pathways involved in each gene.

[0022] S70, Key Gene Scoring:

[0023] Gene criticality was scored using an improved BORDA counting method based on gene centrality, redundancy, and information richness.

[0024] S80, Simplified gene set size:

[0025] Based on the gene key score results from step S70, the genes are sorted in descending order, and 50 genes are successively stacked according to the size of the key score results to enrich biological pathways. The number of genes corresponding to the plateau phase of pathway enrichment is extracted to construct a simplified endocrine gene set.

[0026] Preferably, in step S20, the biological pathways obtained in step S10 are mapped to the molecular initiation events, key events, and harmful outcomes of AOP to construct a multi-level chain of evidence for the harm of endocrine disruption.

[0027] Preferably, in step S30, the endocrine disruptor characteristics are derived from an expert consensus entitled "Consensus on the key characteristics of endocrine-disrupting chemicals as a basis for hazard identification". Relevant interference characteristics are manually matched for each biological pathway obtained in steps S10 and S20, and the biological level described by the pathway and the relevant target organs are assigned to each biological pathway. The above information is integrated to construct a database of endocrine disruption mechanisms.

[0028] Preferably, in step S50, the connectivity value is exported by the stringAPP module of the Cyctoscape software, and the parameters set by stringAPP include species, network type, and confidence threshold; and the construction of the undirected graph model of the biological network and the PAGERANK algorithm of the random walk model with restart are both implemented using the igraph R package.

[0029] Preferably, in step S60, the information richness is evaluated using the number of biological pathways involved in each gene, and the number of biological pathways is multiplied by a coefficient of 100 to obtain the information richness score; the redundancy is evaluated using the total number of genes contained in the biological pathways involved in each gene, and when the gene involves multiple pathways, the pathway with the fewest total number of genes is used, and the total number of genes contained in the biological pathway is used as the redundancy score.

[0030] Preferably, in step S70, the improved BORDA counting method is used to integrate the gene centrality rank, information richness score, and redundancy rank into a gene criticality value.

[0031] Preferably, in step S80, the enrich function of the clusterProfiler R package is used to successively stack 50 genes according to the size of the key score results to enrich biological pathways, and the plateau value of the model of the number of enriched pathways and the number of stacked genes is extracted, and the simplified endocrine gene set is constructed based on the plateau value.

[0032] This invention discloses a method for identifying endocrine disruptors using simplified endocrine transcriptomics. The method involves constructing a simplified endocrine transcriptome based on a zebrafish embryonic system. After performing simplified endocrine transcriptomics testing on the endocrine disruptors to be screened, the omics data are normalized using DESeq or edgeR. Then, biological pathway enrichment is performed using the GSEA method to obtain endocrine disruption-related biological pathways. The pvalues ​​obtained from these pathways are calibrated using FDR, and the endocrine disruption potential of the endocrine disruptors to be screened is identified based on the FDR pvalues.

[0033] Preferably, the method of the present invention for identifying endocrine disruptors using simplified endocrine transcriptomics further includes extracting endocrine disruption features, potential harmful outcomes, biological hierarchy and target organ information from significantly enriched biological pathway information based on FDR pvalue significance.

[0034] 3. Beneficial effects

[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0036] (1) The present invention provides a method for constructing a simplified endocrine transcriptome based on a zebrafish embryo system. It measures only 1,000 genes, which is 1 / 25 to 1 / 30 of the number of genes tested in the existing full transcriptome and 1 / 5 of the total number of genes related to endocrine pathways. This method can significantly cover more than 90% of endocrine-related pathways. While retaining most of the information, it greatly reduces the cost of screening endocrine disruptors by simplifying the gene set.

[0037] (2) The present invention provides a method for constructing a simplified endocrine transcriptome based on a zebrafish embryo system. In the process of constructing the simplified endocrine transcriptome, the existing biological pathways and multi-level adverse outcome pathways (AOP) data related to endocrine disorders are fully integrated. AOP effectively establishes a clear chain of evidence between molecular-level biological genes and proteins that cause endocrine disturbances and high-level organ and human disease evidence, which is more conducive to the understanding of endocrine disturbance mechanisms and risk management.

[0038] (4) The present invention provides a method for constructing a simplified endocrine transcriptome based on a zebrafish embryonic system. Based on endocrine interference characteristics, biological hierarchy and related target organs, a database was constructed and labeled. This method provides great guidance for the subsequent mining of endocrine interference mechanisms and high-throughput screening of endocrine interferences in omics data.

[0039] (5) The present invention provides a method for constructing a simplified endocrine transcriptome based on a zebrafish embryo system. In the construction of the simplified endocrine gene set, the BORDA counting method is used to integrate the evidence scores of three dimensions: gene centrality, information richness, and redundancy. Compared with the traditional importance scoring rules based only on gene centrality, the effective information density assessment indicators information richness and redundancy scores are introduced, which provides support for maximizing the amount of data retained under the premise of constructing the smallest gene set.

[0040] (6) The present invention provides a method for identifying endocrine disruptors using simplified endocrine transcriptomics. Compared with existing endocrine disruptor identification and screening methods, the method of the present invention tests more comprehensive in vivo endpoints related to endocrine disruption, and can also provide evidence support and guidance for in vitro testing and risk assessment, providing a focus and foothold for high-throughput screening and optimal control of endocrine pollutants in my country and globally. Attached Figure Description

[0041] Figure 1 This is a flowchart illustrating a method for constructing a simplified endocrine transcriptome based on a zebrafish embryonic system according to the present invention.

[0042] Figure 2 This is a schematic diagram showing the distribution of PAGERANK values ​​for each gene in the gene set to be simplified in this invention.

[0043] Figure 3 This is a schematic diagram showing the number of biological pathways (information richness) involved in each gene in the gene set to be simplified in this invention;

[0044] Figure 4 This is a schematic diagram showing the number of genes corresponding to the plateau phase based on the number of biological pathway enrichments in this invention;

[0045] Figure 5This is a schematic diagram showing the proportion of the simplified gene set in endocrine-related biological pathways according to the present invention. Detailed Implementation

[0046] The present invention will be further described below with reference to specific embodiments.

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0048] Example 1

[0049] like Figure 1 As shown in this embodiment, a method for constructing a simplified endocrine transcriptome based on a zebrafish embryonic system includes the following steps:

[0050] S10. Collect information on endocrine-related biological pathways:

[0051] Endocrine-related biological pathways were collected from the MSigDB database, with keywords "HORMONE", "ENDOCR", and "REPRODUC". The keyword collection scope mainly included the eight series H and C1-C7 recorded in the database.

[0052] This study summarizes endocrine-related biological pathways contained in the Hallmark and KEGG databases. For the Hallmark database, relevant biological pathways were manually extracted based on pathway descriptions. For the KEGG database, relevant biological pathways categorized under "Endocrine System" and "Endocrine Metabolic Diseases" were included.

[0053] A total of 172 endocrine-related biological pathways were collected, including 114 related to "HORMONE", 22 related to "ENDOCR", 17 related to "REPRODUC", 5 related to "HALLMARK", and 14 related to the "KEGG" pathway. For details, please see Table 1.

[0054] S20. Constructing endocrine-related harmful outcome pathways:

[0055] Based on literature data and AOPWIKI data, we searched the WEB OF SCIENCE website for literature related to adverse outcome pathways of endocrine disorders according to the following criteria: "Adverse outcome pathway (AOP) (topic) and Endocrine disorder (topic)", and further extracted data on adverse outcome pathways (AOP) of endocrine disorders and the relationship between AOP and biological pathways.

[0056] The biological pathways obtained in step S10 were mapped to the molecular initiation events (MIE), key events (KE), and harmful outcomes (AO) of AOPs. Existing information on endocrine disruption-related biological pathways was integrated to construct a multi-level chain of evidence for the harm caused by endocrine disruption. This part extracted a total of 27 endocrine-related AOPs, and added 57 pathways based on those obtained in step S10. See Table 2 for details.

[0057] S30, Endocrine Disruption Mechanism Label Classification:

[0058] Based on the ten characteristics of endocrine disruptors described in the expert consensus statement entitled "Consensus on the key characteristics of endocrine-disrupting chemicals as a basis for hazard identification," the endocrine-related biological pathways obtained in steps S10 and S20 were matched and labeled. The ten characteristics of endocrine disruptors include:

[0059] KC1: EDCs can activate hormone receptors or interact with hormone receptors;

[0060] KC2: EDCs can antagonize hormone receptors;

[0061] KC3: EDCs can alter the expression of hormone receptors;

[0062] KC4: EDCs can alter signal conduction;

[0063] KC5:EDCs can induce epigenetic modifications in hormone-producing and hormone-responsive cells;

[0064] KC6: EDCs can alter hormone synthesis;

[0065] KC7: EDCs can alter hormone transmembrane transport;

[0066] KC8: EDCs can alter hormone distribution or circulation;

[0067] KC9: EDCs can alter hormone metabolism and clearance;

[0068] KC10:EDCs can alter the normal physiological processes of hormone-producing or hormone-responding cells.

[0069] The biological levels and related interfering targets and target organs described in each biological pathway in steps S10 and S20 were matched and labeled. The biological levels included: molecules, cells, organs, and individuals; the related interfering targets mainly included: precise targets related to androgens, estrogens, and thyroid hormones, as well as broad targets and target organs related to hormones and steroid hormones. The above information was integrated into a database of endocrine interference mechanisms for the mechanism differentiation and substance screening guidance of endocrine substances. Information on endocrine-related biological pathways is shown in Table 1, and information on endocrine-related harmful outcomes is shown in Table 2.

[0070] S40. Organize gene information related to biological pathways:

[0071] All genes contained in the biological pathways obtained in steps S10 and S20 were extracted; and the g:Orth module of the g:Profiler website was used to convert all gene IDs to zebrafish homologous gene ENTREZGENE_ACC IDs, retaining only biological pathways with fewer than 1600 genes and related to endocrine function. Finally, 5048 zebrafish species-related genes, involving 206 biological pathways, were retained. The genes contained in this subset were used to construct a simplified endocrine gene set.

[0072] S50, Gene Centrality Assessment:

[0073] The connectivity values ​​of each gene in the simplified gene set obtained in step S40 are extracted in the biological network. The specific implementation steps are as follows:

[0074] Genes are transformed into nodes, and connectivity values ​​are used as edge weights between gene nodes to construct an undirected graph model. The connectivity values ​​between genes in the biological network are derived from the protein-protein interaction network STRING, and the connectivity values ​​stringdbscore are exported from the stringAPP module of Cyctoscape software. The stringdb score values ​​mentioned above integrate scoring evidence from eight dimensions of protein-protein interactions: cross-species, literature mining, experiments, co-expression, database, co-occurrence, adjacency, and fusion.

[0075] The parameters set by stringAPP mainly include: species (Danio rerio), network type (full STRINGnetwork), and confidence threshold (0.4). This part exported a total of 126,681 interaction relationships.

[0076] The PAGERANK value of each gene in the network is calculated using the PAGERANK algorithm with a restarted random walk model. Both the construction of the undirected graph model of the biological network and the PAGERANK algorithm with a restarted random walk model are implemented using the igraph R package. The stringdb score is used as the edge weight of the undirected graph for gene centrality evaluation. The undirected graph is based on the graph_from_data_frame function, and the PAGERANK algorithm is based on the page_rank function.

[0077] Ultimately, approximately 300 genes had PAGERANK values ​​greater than 0.0005, such as TP53 and ctnnb1. A higher PAGERANK value indicates that the gene may have a more important function in the biological network. The distribution of PAGERANK values ​​for all genes to be simplified is shown below. Figure 2 As shown, the gene centrality scores are ranked in descending order using the `rank` function in R.

[0078] S60, Assessment of gene redundancy and information richness:

[0079] The information richness is assessed based on the number of biological pathways involved in each gene; the more pathways involved, the richer the information. The number of pathways involved is multiplied by a coefficient of 100 to obtain the information richness score. It should be noted that the coefficient is set to 100 because the maximum value after multiplying the number of pathways involved by 100 is roughly close to the maximum value of the gene centrality and redundancy scores. The information richness distribution shows that most of the genes to be simplified involve fewer than 10 biological pathways, with only a small number of genes possessing richer information. The information richness distribution is as follows: Figure 3 As shown.

[0080] In addition, the redundancy level is assessed based on the total number of genes involved in the biological pathways of each gene. The fewer genes involved in the pathway, the lower the redundancy level and the more likely it is to provide effective information.

[0081] S70, Key Gene Scoring:

[0082] Gene criticality values ​​are calculated based on the improved BORDA counting method, which calculates the rank of centrality, the score of information richness, and the rank of redundancy. This improves the scores of genes with high centrality and information richness while penalizing redundancy in genes in a large number of pathways.

[0083] S80, Simplified gene set size:

[0084] Based on the gene key score results from step S70, the genes were sorted in descending order. The `enrich` function in the `clusterProfiler` R package was used to progressively add 50 genes according to their key score results to enrich biological pathways. The minimum number of genes for each pathway was 0, and the maximum was 1600. No thresholds were set for p-value and q-value. The cumulative hypergeometric distribution method (also known as Fisher's exact test) was used for pathway enrichment; pathways with a p-value <= 0.05 were considered significantly enriched.

[0085] Finally, the plateau values ​​of the pathway enrichment number and gene number model are extracted. Based on these values, the number of genes with the highest critical scores in step S70 are selected. This part of the genes uses only a small number of test genes but contains most of the pathway significantly enriched results.

[0086] Simulations revealed that a simplified gene set containing approximately 1000 genes reached a plateau in the model, and the plateau ended after 3000 genes due to loss of sensitivity to pathway enrichment. The plateau phase significantly enriched approximately 180 biological pathways, representing over 90% of all endocrine biological pathways. Figure 4 As shown.

[0087] Furthermore, considering the coverage of simplified gene sets across various biological pathways, selecting the top 1000 genes with the greatest criticality in step S70 is sufficient to cover all endocrine biological pathways. Moreover, the degree of coverage of endocrine-related pathways by the top 1000 genes is not significantly different from that of the top 2000 and top 3000 genes. Detailed information is available as follows... Figure 5 As shown in Table 3. Therefore, this embodiment ultimately uses the top 1000 genes with the highest criticality values ​​to construct a simplified endocrine gene set for high-throughput screening and evaluation of endocrine disruptors.

[0088] The simplified endocrine gene set used in this embodiment was used for the simplified endocrine transcriptome test of the substance to be screened. The omics data were then normalized using DESeq or edgeR, and further enriched using the GSEA (Gene Set Enrichment Analysis) method to obtain biological pathways related to endocrine disruption. The pvalues ​​obtained from these pathways were then calibrated using the FDR (Focused Diagnostic Rate). If any endocrine-related biological pathway was significantly enriched (FDR pvalue <= 0.05), it indicated that the substance to be screened might have endocrine disruption potential. Based on the significantly enriched biological pathway information, endocrine disruption characteristics, potential harmful outcomes, biological hierarchy, and target organs can be extracted from Tables 1 and 2, ultimately achieving mechanism discovery and high-throughput screening of potential endocrine disruptors.

[0089] Table 1. Database of Endocrine-Related Biological Pathways

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097]

[0098] Table 2 Endocrine-Related Harmful Outcome Pathways (AOP) Database

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108] Table 3 shows the simplified endocrine gene set of the 1000 genes with the highest key scores.

[0109]

[0110]

[0111]

[0112]

[0113]

[0114] The present invention and its embodiments have been described above illustratively. This description is not restrictive, and the data used is only one embodiment of the present invention. The actual combination of data is not limited to this. Therefore, if those skilled in the art are inspired by this description and, without departing from the spirit of the present invention, devise similar embodiments and examples of the technical solution without creative design, all such embodiments and examples should fall within the protection scope of the present invention.

Claims

1. A method for constructing a simplified endocrine transcriptome based on a zebrafish embryonic system, characterized in that: Specifically, the following steps are included: S10. Collect information on endocrine-related biological pathways: Endocrine-related biological pathways were collected from the MSigDB, HALLMARK, and KEGG databases. S20. Constructing endocrine-related harmful outcome pathways: Based on literature data and AOPWIKI data, endocrine-related harmful outcome pathway AOP data are extracted, and the AOP data is mapped with the biological pathways obtained in step S10. The biological pathways obtained in step S10 are mapped to the molecular initiation events, key events and harmful outcomes of AOP, and a multi-level endocrine disturbance harm evidence chain is constructed. S30, Endocrine Disruption Mechanism Label Classification: Based on the characteristics of endocrine disruptors, the biological pathways obtained in steps S10 and S20 are matched and labeled, and the biological hierarchy and target organs of the biological pathways described in steps S10 and S20 are matched and labeled. S40. Organize gene information related to biological pathways: Extract all genes of the biological pathways obtained in steps S10 and S20, and map all genes homologously to zebrafish species based on ENTREZ ID to obtain the gene set to be simplified; S50, Gene Centrality Assessment: Extract the connectivity values ​​between genes in the biological mesh obtained in step S40 of the gene set to be simplified, convert the genes into nodes, use the connectivity values ​​as edge weights between gene nodes to construct an undirected graph model, and calculate the centrality PAGERANK value of each gene in the network using the PAGERANK algorithm with restart. S60, Assessment of gene redundancy and information richness: Redundancy is assessed based on the total number of genes involved in the biological pathways involved by each gene, and information richness is assessed based on the number of biological pathways involved by each gene. The information richness is assessed using the number of biological pathways involved by each gene, and the number of biological pathways is multiplied by a coefficient of 100 to obtain the information richness score. Redundancy is assessed using the total number of genes involved in the biological pathways involved by each gene. When a gene involves multiple pathways, the pathway with the fewest total genes is used, and the total number of genes involved in the biological pathway is used as the redundancy score. S70, Key Gene Scoring: Gene criticality was scored using an improved BORDA counting method based on gene centrality, redundancy, and information richness. S80, Simplified gene set size: Based on the gene key score results from step S70, the genes are sorted in descending order, and 50 genes are successively stacked according to the size of the key score results to enrich biological pathways. The number of genes corresponding to the plateau phase of pathway enrichment is extracted to construct a simplified endocrine gene set.

2. The method for constructing a simplified endocrine transcriptome based on a zebrafish embryonic system according to claim 1, characterized in that: In step S30, the endocrine disruptor characteristics are derived from an expert consensus entitled "Consensus on the key characteristics of endocrine-disrupting chemicals as a basis for hazard identification". Interference characteristics are manually matched for each biological pathway obtained in steps S10 and S20, and the biological level and target organ described by the pathway are assigned to each biological pathway. The above information is integrated to construct a database of endocrine disruption mechanisms.

3. The method for constructing a simplified endocrine transcriptome based on a zebrafish embryonic system according to claim 1, characterized in that: In step S50, the connectivity value is exported by the stringAPP module of the Cytoscape software. The parameters set by stringAPP include species, network type, and confidence threshold. Furthermore, the construction of the undirected graph model of the biological network and the PAGERANK algorithm for the random walk model with restart are both implemented using the igraph R package.

4. The method for constructing a simplified endocrine transcriptome based on a zebrafish embryonic system according to claim 1, characterized in that: In step S70, the gene centrality rank, information richness score, and redundancy rank are integrated into a gene criticality value using the improved BORDA counting method.

5. The method for constructing a simplified endocrine transcriptome based on a zebrafish embryonic system according to claim 1, characterized in that: In step S80, the enrich function of the clusterProfiler R package is used to successively stack 50 genes according to the size of the key score results to enrich biological pathways. The plateau value of the model of the number of enriched pathways and the number of stacked genes is extracted, and the simplified endocrine gene set is constructed based on the plateau value.

6. A method for identifying endocrine disruptors using simplified endocrine transcriptomics, characterized in that: Using the method described in any one of claims 1-5, a simplified endocrine transcriptome based on a zebrafish embryonic system was constructed. After performing the simplified endocrine transcriptome test on the endocrine disruptors to be screened, the omics data were normalized using DESeq or edgeR. Then, biological pathway enrichment was performed based on the GSEA method to obtain biological pathways related to endocrine disruption. The pvalues ​​obtained from the pathways were calibrated using FDR. Based on the FDR pvalues, the endocrine disruption potential of the endocrine disruptors to be screened was identified.

7. A method for identifying endocrine disruptors using simplified endocrine transcriptomics according to claim 6, characterized in that: It also includes extracting endocrine disruption features, potential harmful outcomes, biological hierarchy and target organ information from significantly enriched biological pathway information based on FDR pvalue significance.