Environmental endocrine disrupter multi-omics data integration analysis system and method
Through a hybrid database architecture and a dual-stream deep neural network, the problems of low data traceability efficiency and difficulty in data integration in multiomics research on environmental endocrine disruptors are solved, and efficient multimodal data integration analysis and visualization are realized, supporting accurate prediction and query of EDCs toxicity mechanisms.
Patent Information
- Application Number
- CN202510848718.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In the prior art, multiomic research on environmental endocrine disruptors (EDCs) lacks a standardized metadata framework and a systematic data governance architecture, resulting in low data traceability efficiency, limited cross-platform data collaborative utilization, and lack of a data integration paradigm for specific exposure conditions and biological target systems, making it difficult to accurately analyze EDCs toxicity pathways and personalized risk assessment.
The hybrid database architecture is adopted, combined with MySQL relational database and Neo4j graph database, and multimodal storage and management of chemical data and multiomic data of environmental endocrine disturbances are realized. Cross-modal correlation modeling is carried out through dual-stream deep neural networks, and automated conversion and structured storage of multiple data formats are supported. Data acquisition and preprocessing are combined with Django REST Framework and Scrapy-Redis crawler frameworks.
It significantly improves the efficiency of data traceability, shortens the data preprocessing cycle, can process the characteristics of multimodal data and adaptively learn the correlation between modals, supports supervised and unsupervised training modes, and realizes accurate prediction and query analysis of the toxicity mechanism of EDCs.
Smart Images

Figure CN120356536A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical fields of biological science research and big data mining, and particularly relates to a multi-omics data integration and analysis system and method for environmental endocrine disruptors. Background Technique
[0002] Currently, several representative technical platforms have been formed in the multi-omics research system of environmental toxicology. Among them, the Comparative Toxicogenomics Database (CTD) developed by the US National Institute of Environmental Health Sciences is particularly prominent. This platform constructs a knowledge network covering multi-dimensional interaction relationships of chemicals, genes, diseases, and biological processes by systematically extracting and integrating toxicogenomics data from published scientific literature. At the same time, international well-known bioinformatics repositories such as the US National Center for Biotechnology Information (NCBI) and the European Molecular Biology Laboratory European Bioinformatics Institute (EMBL-EBI) have continuously accumulated and standardized the storage of a large number of high-throughput sequencing datasets related to EDCs through their core databases such as Gene Expression Omnibus (GEO) and ArrayExpress, including multi-omics toxicity test data such as transcriptomics and epigenomics.
[0003] There are two core bottlenecks in the current multi-omics research system for environmental endocrine disruptors (EDCs): (1) For the multi-omics research on environmental endocrine disruptors (EDCs), the current technology has not established a standardized metadata framework and a systematic data governance architecture, and the original data generally exists in a scattered storage state. There is a lack of a standardized conversion mechanism for multi-source heterogeneous data such as chemical exposure profiles, transcriptomics, and epigenetics, resulting in a decrease in data traceability efficiency and a significant increase in the time required for the structured data extraction process. In addition, the data island phenomenon between different research institutions severely restricts the collaborative utilization of cross-platform data.
[0004] (2) Although existing public databases (such as CTD, NCBI, and EEMBL-EBI) have achieved the standardized storage of basic data, they lack a data integration paradigm for EDCs and their specific exposure conditions (such as dose-effect relationships, time-dependent exposures), biological target systems (such as molecular networks at the cross-species / organ / cell levels), and toxicity endpoints. This technical shortcoming directly restricts the accurate analysis of EDC toxicity pathways and also poses challenges to the development of personalized risk assessment models based on multi-source data. Therefore, it is very promising to develop a multi-omics data integration and analysis system and method for environmental endocrine disruptors. Summary of the Invention
[0005] The object of the present application is to solve the problems of the prior art, and provide a system and method for integrated analysis of multi-omics data of environmental endocrine disruptors.
[0006] To solve the technical problems, the technical solution of the present application is: a method for integrated analysis of multi-omics data of environmental endocrine disruptors, comprising the following steps: Step 1: Data collection, extracting chemical data and multi-omics data of environmental endocrine disruptors from public databases; Step 2: Data storage; A hybrid database architecture is adopted to achieve multi-modal storage and management of chemical data and multi-omics data of environmental endocrine disruptors. The chemical data and multi-omics data of environmental endocrine disruptors are divided into structured data and unstructured data; For structured data, a normalized metadata management system is constructed based on the MySQL relational database, and core data tables are designed through entity relationship modeling to store the structured data; For unstructured data, a dynamic knowledge network is constructed based on the Neo4j graph database to store the unstructured data; Step 3: Integrated analysis; Step 3-1: Preprocess the chemical data and multi-omics data of environmental endocrine disruptors; Step 3-2: Perform cross-modal association modeling on chemical data and multi-omics data to achieve cross-modal association of chemical data and multi-omics data; Step 4: Visualization; Visualize the cross-modal association results of chemical data and multi-omics data for information query and toxicity mechanism prediction of environmental endocrine disruptors.
[0007] Preferably, the data collection in Step 1 is specifically: using Django REST Framework to build a standardized API interface, supporting data upload in JSON / CSV / XML formats, and at the same time integrating the Scrapy-Redis distributed crawler framework to directionally crawl the chemical data and multi-omics data of environmental endocrine disruptors from public databases; The chemical data includes compound tables, SMILES structural formulas, inhibition constants, half maximal effective concentration, and half maximal inhibitory concentration; The multi-omics data includes transcriptome, metabolome, genome, proteome, epigenome, lipidome, spatial transcriptome, single-cell transcriptome, and exposure record tables.
[0008] Preferably, in step 3-1, the chemical data and multi-omics data of environmental endocrine disruptors are preprocessed. The transcriptome preprocessing of the multi-omics data includes the following steps: Step 3-1-1: Raw data quality control. Use the fastp tool to filter the collected transcriptome data to ensure that the Phred score is ≥20, and at the same time remove the adapter sequences. Step 3-1-2: Sequence alignment. Map the read transcriptome data to the reference genome through the STAR tool, and set the maximum mismatch rate to 10%. Step 3-1-3: Expression quantification. Use featureCounts to count the gene count matrix for gene expression of the transcriptome data. Step 3-1-4: Differential analysis. Use DESeq2 for normalization and hypothesis testing for differential analysis.
[0009] Preferably, in step 3-1, the chemical data and multi-omics data of environmental endocrine disruptors are preprocessed. The metabolome data preprocessing includes the following steps: Step 3-1-1: Raw mass spectrometry peak extraction. Use the centWave algorithm of XCMS to detect the chromatographic peaks of the metabolome data, set the ppm tolerance to 10, and the peak width range to 5-20 seconds. Step 3-1-2: Peak alignment and annotation. Use CAMERA for peak alignment and annotation of the chromatographic peaks, match the HMDB 5.0 database to identify the metabolome data features. Step 3-1-3: Normalization.
[0010] Preferably, step 3-2 includes the following steps: Step 3-2-1: Chemical data encoding. Based on the RDKit chemoinformatics toolkit, convert the chemical data of environmental endocrine disruptors into a 2048-bit Morgan fingerprint feature vector to capture the molecular structure features of the chemical data, and at the same time calculate molecular descriptors as supplementary feature vectors to obtain the chemical data stream. Step 3-2-2: Multi-omics data dimensionality reduction. Perform principal component analysis dimensionality reduction on the high-dimensional multi-omics data. Use the standard of cumulative variance contribution rate ≥95% to determine 500 principal components to be retained, and verify the biological consistency of the dimensionality-reduced multi-omics data through t-SNE visualization to obtain the multi-omics data stream. Step 3-2-3: Cross-modal alignment. Construct a two-stream deep neural network. The chemical data stream uses a 3-layer fully connected network to process the fingerprint feature vector, and the multi-omics data stream uses a 1D convolutional network to extract the features after principal component analysis dimensionality reduction. Finally, learn the cross-modal association weights through a cross-attention mechanism and output the low-dimensional feature vector in the joint embedding space to achieve cross-modal association of chemical data and multi-omics data.
[0011] Preferably, the two-stream deep neural network consists of a chemical data stream, multiple omics data streams, and a cross-modal fusion mechanism. The chemical data stream adopts a multi-layer fully connected layer structure, including an input dimension of 1024, a hierarchical structure of hidden layers [512, 256], and is equipped with a ReLU activation function and a BatchNorm1d normalization layer after each layer. At the same time, a Dropout rate of 0.2 is used to provide regularization, and finally, a 128-dimensional feature vector is output; the multiple omics data stream processes sequence data using a one-dimensional convolutional neural network, designs a two-layer convolutional structure with an input dimension of 200, the number of channels is [64, 128] respectively, uses a 3×3 convolutional kernel, combines max pooling and adaptive average pooling layers, and finally obtains a 128-dimensional feature vector through linear projection; the feature vectors of the chemical data stream and the multiple omics data stream are cross-modally aligned through a multi-head attention mechanism, and then pass through a feature fusion layer. The feature fusion layer includes LayerNorm normalization, linear dimensionality reduction, and a Tanh activation function to generate a unified 128-dimensional feature vector.
[0012] Preferably, the construction steps of the two-stream deep neural network are as follows: Step 3-2-3-1: Chemical data stream processing: 1) Fully connected layer calculation: H_l = ReLU(BN(W_l·H_{l - 1}+b_l)); Where: H_l is the output of the l-th layer; H_{l - 1} is the output of the (l - 1)-th layer; W_l is the weight matrix; b_l is the bias vector; BN is batch normalization; ReLU(x) = max(0, x); 2) Dropout regularization: H_dropout = H * mask; mask ~ Bernoulli(p); Where: H_dropout is the output feature vector, the feature after Dropout processing; H is the input feature vector, the output activation value from the previous layer, that is, the output of the fully connected layer; mask is a binary mask matrix; p is the retention probability, controlling the regularization strength; Bernoulli(p) is the Bernoulli distribution; Step 3-2-3-2: Multiple omics data stream processing: 1) One-dimensional convolution operation: Conv1D(x) = σ(W * x + b); Where: Conv1D(x) is one-dimensional convolution; σ is the ReLU activation function; W is the convolution kernel weight; * is the convolution operation; b is the bias term; 2) Max pooling: MaxPool(x) = max(x[i:i + k]); Where: MaxPool(x) performs max pooling on the input data x and returns the maximum value within each pooling window; k is the pooling window size; i is the starting index of the current window; x is the input data; Step 3-2-3-3: Multi-head attention mechanism: 1) Query Q, key K, value V calculation: Q = X_q · W_q; K = X_k · W_k; V = X_v · W_v; Attention weight calculation: Attention(Q, K, V) = softmax(QK^T / √d_k)V; Where: d_k is the attention dimension; softmax is the normalization function; X_q is the input query Q representation; W_q is the weight matrix of the query; X_k is the input key K representation; W_k is the weight matrix of the key; X_v is the input value V representation; W_v is the weight matrix of the value; T is the transpose operation; K^T is the transpose of the key matrix K; 2) Multi-head attention: MultiHead(Q, K, V) = Concat(head_1,..., head_h)W_o; head_i = Attention(QW_i^Q, KW_i^K, VW_i^V); Where: Concat(head_1, ..., head_h) concatenates h heads, and the dimension after concatenation will increase, representing the parallel computation results of multiple attention heads; W_o is a linear transformation matrix used to transform the result after concatenating multiple heads to obtain the final output; head_i is the output of the i-th attention head; Attention(QW_i^Q, KW_i^K, VW_i^V) is for the i-th attention head, which linearly transforms the query Q, key K, and value V using different weight matrices respectively, and then calculates the standard attention; W_i^Q represents the query Q weight matrix of the i-th head, which linearly transforms the query Q and maps the original query to a new representation space; W_i^K represents the key K weight matrix of the i-th head, which linearly transforms the key K and maps the original key to a new representation space; W_i^V represents the value V weight matrix of the i-th head, which linearly transforms the value V and maps the original value to a new representation space; Step 3-2-3-4: Feature fusion: 1) Feature concatenation: H_fused = [H_chem + Attn_out; H_omics]; Where: H_fused is the concatenated new feature vector; H_chem is the chemical-related feature vector; Attn_out is the output from the attention mechanism; H_omics is the feature vector generated based on other omics data; + is the addition operation of H_chem and Attn_out; ; is the concatenation in the feature dimension, concatenating H_chem + Attn_out and H_omics column-wise to form the new feature vector H_fused; 2) LayerNorm normalization: μ = (1 / H)ΣH_i; σ = sqrt((1 / H)Σ(H_i - μ)^2); LN(H) = γ * (H - μ) / σ + β; H_i is the value of the i-th element in the input feature vector H; μ is the arithmetic mean of all elements of the input feature vector H; σ is the standard deviation of the input feature vector H, measuring the degree of data dispersion; sqrt is the square root; LN(H) is the output after normalizing the input feature vector H; γ is the scaling parameter that controls the scale of the output after normalization; β is the offset parameter that adjusts the position of the output after normalization.
[0013] Preferably, an environmental endocrine disruptor multi-omics data integration and analysis system is used for implementing the above-mentioned environmental endocrine disruptor multi-omics data integration and analysis method, including a data access layer, a data storage layer, an integration and analysis layer, and a visualization layer. The data access layer is used to achieve the standardized collection and upload of multi-source data. The standardized API interface built based on Django REST Framework supports the upload of data in JSON / CSV / XML formats; the integrated Scrapy-Redis distributed crawler framework is used to directionally crawl environmental endocrine disruptor-related data in public databases; the browser-side large file sharding upload function implemented based on HTML5 FileAPI has a sharding size of 10MB. The data storage layer uses a MySQL relational database and a Neo4j graph database for multi-modal storage and management of the chemical data and multi-omics data of environmental endocrine disruptors. The integration and analysis layer is used to preprocess the chemical data and multi-omics data of environmental endocrine disruptors and perform cross-modal association modeling to achieve cross-modal association of chemical data and multi-omics data. The visualization layer is used to visualize the cross-modal association results of chemical data and multi-omics data, and realizes dynamic interactive visualization based on the D3.js framework, which is used for information query and toxicity mechanism prediction of environmental endocrine disruptors.
[0014] Compared with the prior art, the advantages of this application are as follows: (1) This application discloses an environmental endocrine disruptor multi-omics data integration and analysis system. Through data collection, data storage, integration and analysis, and visualization, it constructs a dedicated metadata framework for environmental endocrine disruptors, establishes a full-dimensional collection system covering chemical data and multi-omics data. This application supports the automatic conversion and structured storage of multiple data formats, thereby significantly improving the data traceability efficiency and shortening the data preprocessing cycle. (2) The dual-stream deep neural network disclosed in this application can not only process the specific features of two modal data respectively, but also adaptively learn the associations between modalities through the attention mechanism. At the same time, through multi-level regularization measures, it ensures the stability and generalization ability of the model, supports both supervised and unsupervised training modes, and shows good feature extraction and modal fusion performance on the example data. (3) This application integrates heterogeneous network data from different experimental conditions, species, or omics levels through multi-layer network modeling, realizes multi-modal data integration analysis and visualization functions, and promotes the prediction and query analysis of the toxicity mechanism of EDCs. Description of the Drawings
[0015] Figure 1 It is a flowchart of a method for integrating and analyzing multi-omics data of environmental endocrine disruptors in this application; Figure 2 It is a framework diagram of a system for integrating and analyzing multi-omics data of environmental endocrine disruptors in this application; Figure 3 It is a schematic diagram of the online function usage of a method for integrating and analyzing multi-omics data of environmental endocrine disruptors in this application; Figure 4 It is an analysis platform designed by using a method for integrating and analyzing multi-omics data of environmental endocrine disruptors in this application; Figure 5 It is Figure 4 The function module execution interface of Transcriptional Factors in Specific Embodiments
[0016] The following describes this application in detail with reference to the drawings and specific embodiments, but this application is not limited to these embodiments. This application covers any alternatives, modifications, equivalent methods, and solutions made within the essence and scope of this application. In order to enable the public to have a thorough understanding of this application, specific details are described in detail in the following embodiments of this application, and those skilled in the art can fully understand this application without the description of these details.
[0017] As Figure 1 As shown, this application discloses a method for integrating and analyzing multi-omics data of environmental endocrine disruptors, including the following steps: Step 1: Data collection, extracting chemical data and multi-omics data of environmental endocrine disruptors from public databases; For example: extracting chemical molecules and bioactivity data of environmental endocrine disruptors from public databases such as BindingDB, ChEMBL, and PubChem, and extracting multi-omics data and their corresponding experimental record information from public databases such as ToxCast, Tox 21, EMBL-EBI, NCBI, and the National Genomics Data Center of China.
[0018] Step 2: Data storage; Implement multi-modal storage and management of chemical data and multi-omics data of environmental endocrine disruptors by using a hybrid database architecture. The chemical data and multi-omics data of environmental endocrine disruptors are divided into structured data and unstructured data; For structured data, a normalized metadata management system is built based on the MySQL relational database. Core data tables are designed through entity-relationship modeling to store structured data. For unstructured data, a dynamic knowledge network is built based on the Neo4j graph database to store unstructured data. Step 3: Integration and analysis; Step 3-1: Preprocess the chemical data and multi-omics data of environmental endocrine disruptors; Step 3-2: Conduct cross-modal association modeling on chemical data and multi-omics data to achieve cross-modal association of chemical data and multi-omics data; Step 4: Visualization; Visualize the cross-modal association results of chemical data and multi-omics data for information query and toxicity mechanism prediction of environmental endocrine disruptors.
[0019] Preferably, the data collection in Step 1 is specifically as follows: Use Django REST Framework to build a standardized API interface, support data upload in JSON / CSV / XML formats, and at the same time integrate the Scrapy-Redis distributed crawler framework to direc tionally crawl the chemical data and multi-omics data of environmental endocrine disruptors in public databases; The chemical data includes compound tables, SMILES structural formulas, inhibition constants, half maximal effective concentration, and half maximal inhibitory concentration; The multi-omics data includes transcriptome, metabolome, genome, proteome, epigenome, lipidome, spatial transcriptome, single-cell transcriptome, and exposure record tables.
[0020] Preferably, in Step 3-1, the chemical data and multi-omics data of environmental endocrine disruptors are preprocessed. The preprocessing of the transcriptome in the multi-omics data includes the following steps: Step 3-1-1: Raw data quality control. Use the fastp tool to filter the collected transcriptome data to ensure that the Phred score is ≥20, and at the same time remove adapter sequences; Step 3-1-2: Sequence alignment. Map the read transcriptome data to the reference genome through the STAR tool, and set the maximum mismatch rate to 10%; Step 3-1-3: Expression quantification. Use featureCounts to count the gene count matrix for gene expression of transcriptome data; Step 3-1-4: Differential analysis. Use DESeq2 for normalization and hypothesis testing for differential analysis.
[0021] Preferably, in step 3-1, the chemical data and multi-omics data of the environmental endocrine disruptors are preprocessed, and the preprocessing of the metabolomics data includes the following steps: Step 3-1-1: Extraction of original mass spectrometry peaks. Use the centWave algorithm of XCMS to detect the chromatographic peaks of the metabolomics data, set the ppm tolerance to 10, and the peak width range to 5-20 seconds; Step 3-1-2: Peak alignment and annotation. Use CAMERA for chromatographic peak alignment and annotation, match the HMDB 5.0 database to identify the characteristics of the metabolomics data; Step 3-1-3: Normalization.
[0022] Preferably, step 3-2 includes the following steps: Step 3-2-1: Chemical data encoding: Based on the RDKit chemoinformatics toolkit, convert the chemical data of environmental endocrine disruptors into 2048-bit Morgan fingerprint feature vectors to capture the molecular structure characteristics of the chemical data, and at the same time calculate molecular descriptors as supplementary feature vectors to obtain chemical data streams; Step 3-2-2: Dimensionality reduction of multi-omics data. Perform principal component analysis dimensionality reduction on high-dimensional multi-omics data, and determine to retain 500 principal components according to the standard of cumulative variance contribution rate ≥ 95%, and verify the biological consistency of the multi-omics data after dimensionality reduction through t-SNE visualization to obtain multi-omics data streams; Step 3-2-3: Cross-modal alignment: Construct a two-stream deep neural network. The chemical data stream uses a 3-layer fully connected network to process the fingerprint feature vectors, and the multi-omics data stream uses a 1D convolutional network to extract the features after principal component analysis dimensionality reduction. Finally, learn the cross-modal association weights through a cross-attention mechanism and output low-dimensional feature vectors in the joint embedding space to achieve cross-modal association of chemical data and multi-omics data.
[0023] Preferably, the dual-stream deep neural network consists of a chemical data stream, multiple omics data streams, and a cross-modal fusion mechanism. The chemical data stream adopts a multi-layer fully connected layer structure, including an input dimension of 1024, a hierarchical structure of hidden layers [512, 256], and is equipped with a ReLU activation function and a BatchNorm1d normalization layer after each layer. At the same time, a Dropout rate of 0.2 is used to provide regularization, and finally, a 128-dimensional feature vector is output. The multiple omics data streams use a one-dimensional convolutional neural network to process sequence data, design a two-layer convolutional structure with an input dimension of 200, the number of channels is [64, 128] respectively, use a 3×3 convolutional kernel, combine max pooling and adaptive average pooling layers, and finally obtain a 128-dimensional feature vector through linear projection. The feature vectors of the chemical data stream and the multiple omics data streams are cross-modally aligned through a multi-head attention mechanism, and then pass through a feature fusion layer. The feature fusion layer includes LayerNorm normalization, linear dimensionality reduction, and a Tanh activation function to generate a unified 128-dimensional feature vector.
[0024] Preferably, the construction steps of the dual-stream deep neural network are as follows: Step 3-2-3-1: Chemical data stream processing: 1) Fully connected layer calculation: H_l = ReLU(BN(W_l·H_{l - 1}+b_l)); Where: H_l is the output of the l-th layer; H_{l - 1} is the output of the (l - 1)-th layer; W_l is the weight matrix; b_l is the bias vector; BN is batch normalization; ReLU(x)=max(0,x); 2) Dropout regularization: H_dropout = H * mask; mask ~ Bernoulli(p); Where: H_dropout is the output feature vector, the feature after Dropout processing; H is the input feature vector, the output activation value from the previous layer, that is, the output of the fully connected layer; mask is a binary mask matrix, randomly determining which neurons are retained or discarded; p is the retention probability, controlling the regularization strength; Bernoulli(p) is the Bernoulli distribution; Step 3-2-3-2: Multiple omics data stream processing: 1) One-dimensional convolutional operation: Conv1D(x) = σ(W * x + b); Where: Conv1D(x) is a one-dimensional convolution; σ is the ReLU activation function; W is the convolutional kernel weight; * is the convolution operation; b is the bias term; 2) Max pooling: MaxPool(x) = max(x[i:i + k]); Where: MaxPool(x) performs max pooling on the input data x and returns the maximum value within each pooling window; k is the pooling window size; i is the starting index of the current window; x is the input data, usually a one-dimensional or two-dimensional array or tensor; Step 3-2-3-3: Multi-head attention mechanism: 1) Query Q, key K, value V calculation: Q = X_q · W_q; K = X_k · W_k; V = X_v · W_v; Attention weight calculation: Attention(Q, K, V) = softmax(QK^T / √d_k)V; Where: d_k is the attention dimension; softmax is the normalization function; X_q is the input query Q representation, usually a matrix; W_q is the weight matrix of the query; X_k is the input key K representation, usually a matrix; W_k is the weight matrix of the key; X_v is the input value V representation; W_v is the weight matrix of the value; T is the transpose operation; K^T is the transpose of the key matrix K; 2) Multi-head attention: MultiHead(Q, K, V) = Concat(head_1,..., head_h)W_o; head_i = Attention(QW_i^Q, KW_i^K, VW_i^V); Where: Concat(head_1, ..., head_h) concatenates h heads, and the dimension after concatenation will increase, representing the parallel computation results of multiple attention heads; W_o is a linear transformation matrix used to transform the result after concatenating multiple heads to obtain the final output; head_i is the output of the i-th attention head; Attention(QW_i^Q, KW_i^K, VW_i^V) is for the i-th attention head, which linearly transforms the query Q, key K, and value V using different weight matrices respectively, and then calculates the standard attention; W_i^Q represents the query Q weight matrix of the i-th head, which linearly transforms the query Q and maps the original query to a new representation space; W_i^K represents the key K weight matrix of the i-th head, which linearly transforms the key K and maps the original key to a new representation space; W_i^V represents the value V weight matrix of the i-th head, which linearly transforms the value V and maps the original value to a new representation space; Step 3-2-3-4: Feature fusion: 1) Feature concatenation: H_fused = [H_chem + Attn_out; H_omics]; Where: H_fused is the newly formed feature vector after concatenation; H_chem is the chemical-related feature vector; Attn_out is the output from the attention mechanism; H_omics is the feature vector generated based on other omics data; + is the addition operation on H_chem and Attn_out; ; is the concatenation in the feature dimension, concatenating H_chem + Attn_out and H_omics column by column to form the new feature vector H_fused; 2) LayerNorm normalization: μ = (1 / H)ΣH_i; σ = sqrt((1 / H)Σ(H_i - μ)^2); LN(H) = γ * (H - μ) / σ + β; H_i is the value of the i-th element in the input feature vector H; μ is the arithmetic mean of all elements of the input feature vector H; σ is the standard deviation of the input feature vector H, which measures the degree of data dispersion; sqrt is the square root; LN(H) is the output after normalizing the input feature vector H; γ is the scaling parameter, which controls the scale of the output after normalization; β is the offset parameter, which adjusts the position of the output after normalization.
[0025] Preferably, an environmental endocrine disruptor multi-omics data integration and analysis system includes a data access layer, a data storage layer, an integration and analysis layer, and a visualization layer. As Figure 2 shown, preferably, an environmental endocrine disruptor multi-omics data integration and analysis system is used for the implementation of the above-mentioned environmental endocrine disruptor multi-omics data integration and analysis method, and includes a data access layer, a data storage layer, an integration and analysis layer, and a visualization layer. The data access layer is used to achieve the standardized collection and upload of multi-source data. The standardized API interface built based on Django RESTFramework supports the upload of data in JSON / CSV / XML formats; the integrated Scrapy-Redis distributed crawler framework is used to directionally crawl environmental endocrine disruptor-related data in public databases; the browser-side large file sharding upload function implemented based on HTML5 FileAPI, and the sharding size is 10MB. The data storage layer uses a MySQL relational database and a Neo4j graph database for multi-modal storage and management of chemical data and multi-omics data of environmental endocrine disruptors. The integration and analysis layer is used to preprocess chemical data and multi-omics data of environmental endocrine disruptors and perform cross-modal association modeling to achieve cross-modal association of chemical data and multi-omics data. The visualization layer is used to visualize the cross-modal association results of chemical data and multi-omics data, and realizes dynamic interactive visualization based on the D3.js framework.
[0026] The standardized API interface includes NCBI EUtils API and ArrayExpress REST API. The public databases are NCBI GEO database, ArrayExpress, and CTD. NCBI EUtils API is used to retrieve RNA-seq and ChIP-seq datasets related to environmental endocrine disruptors from the NCBI GEO database, and ArrayExpress REST API supports batch downloading of metabolome and epigenome data from the ArrayExpress database.
[0027] Example 1 The system of this application realizes the full - process integration of data collection, storage, analysis, and visualization. The specific architecture is as follows: Data access layer: Based on Python programming, a data collection module for multi - source data is designed; it includes building a standardized API interface (api / v1 / upload) using Django REST Framework, supporting data upload in JSON / CSV / XML formats; integrating the Scrapy - Redis distributed crawler framework to directionally crawl EDCs - related data from public databases such as PubMed and ChEMBL; implementing large - file shard upload on the browser side based on the HTML5 File API (shard size: 10MB).
[0028] The data entry layer of the data access layer system, whose core function is to support the standardized access and integration of multi - source heterogeneous data. Considering the particularity of environmental endocrine disruptor (EDCs) research, this module integrates the API interfaces of bioinformatics public databases (such as NCBI GEO, ArrayExpress, CTD), supporting users to batch - retrieve and download multi - omics datasets related to EDCs exposure (such as RNA - seq, ChIP - seq, metabolomics mass spectrometry data) through keywords (such as compound name, species, experiment type). At the same time, the module designs a flexible data upload interface, compatible with multiple formats such as JSON, CSV, and XML, and builds a built - in data schema parser, which can automatically identify the omics metadata (such as sample number, experimental platform, sequencing depth) and raw data files (such as FASTQ, mzML) uploaded by users, and convert them into a unified Parquet columnar storage format through the ETL (Extract - Transform - Load) process to ensure the efficiency of subsequent analysis. To improve data quality, the module integrates a data verification rule library (such as the MIAME standard) to perform real - time verification on the integrity of metadata (such as missing exposure dose, time parameter), and prompts users to complete necessary fields through a dynamic form. In addition, the module supports docking with the laboratory information management system (LIMS) to achieve traceability of the entire life cycle of experimental samples. All standardized data are stored in a normalized table of the MySQL database with the experiment ID as the primary key, and at the same time, a data lineage graph is generated to record the data source and transformation path, providing support for auditing and reproducibility analysis.
[0029] Data storage layer: In this embodiment, a hybrid database architecture (MySQL + Neo4j) is adopted to achieve multi - modal storage and management of environmental endocrine disruptor (EDCs) chemical data and multi - omics data.
[0030] For structured data, a normalized metadata management system is built based on MySQL 8.0. The core data tables are designed through entity-relationship modeling (ER Modeling): ① The Chemical table stores the basic attributes of EDCs in the third normal form (3NF), including chemical identifiers (PubChem CID, CAS number), physicochemical parameters (LogP, molecular weight), toxicokinetic indicators (biological half-life, ADME characteristics), etc. fields, and stores dynamic experimental parameters through JSON type fields; ② The Omics table organizes multi-dimensional omics data in a star schema, with the experiment number as the primary key to associate sub-tables such as transcriptomics (FPKM matrix), epigenomics (methylation site β value), and metabolomics (mass spectrometry peak intensity), and at the same time ensures data type consistency through CHECK constraints; ③ The Exposure table realizes spatio-temporal dimension modeling, records key metadata such as exposure dose (nM-μM gradient), time window (acute / chronic), biological model (species, organ, cell line), etc., and uses window functions to support the time series analysis of the dose-effect relationship. This architecture ensures data consistency through foreign key constraints and transaction isolation levels (READ COMMITTED), and combines generated columns to achieve real-time calculation of common metrics (such as the number of differentially expressed genes).
[0031] For unstructured data, a dynamic knowledge network is built based on the Neo4j graph database: ① The EDCs-gene-disease associations extracted from the literature are transformed into graph nodes (Node) and relationship edges (Relationship), and complex network topology analysis (such as discovery of the shortest toxicity pathway) is realized using the Cypher query language; ② Multimodal data is integrated through the property graph model, with the RNA-seq expression profile (in JSON format) embedded in the gene node and the mass spectrometry imaging graph (stored as a cloud storage URL) associated with the EDCs node; ③ Graph embedding algorithms (such as node2vec) are used to vectorize unstructured text features (such as experimental conclusion descriptions) and store them as node attributes to support similarity retrieval.
[0032] The two databases are coordinated through the API gateway: The standardized exposure metadata in the MySQL relational database is input as the initial condition for graph traversal, while the potential molecular interaction relationships inferred by Neo4j can be written back in reverse to the hypothesis table (Hypothesis) in MySQL for verification.
[0033] Integration and analysis layer: Implement asynchronous task engines and network topology analysis, etc. Use Celery + Redis to implement a distributed task queue to handle time-consuming operations such as data cleaning and feature extraction. Build a compound-gene association network based on the NetworkX library and store the adjacency matrix in the Redis Graph database. At the same time, deploy thread scripts and analysis environments for preprocessing and analysis of multi-omics data (genomics, transcriptomics, metabolomics, proteomics, epigenomics, spatial omics, and single-cell omics, etc.). Through front-end input instructions, preprocessing of various omics data can be achieved, and analysis results can be obtained.
[0034] The preprocessing of multi-omics data adopts a modular design, providing customized analysis pipelines for omics data such as genomics, transcriptomics, and metabolomics. For example, the transcriptome data analysis process includes: ①Raw data quality control (FastQC + MultiQC); ②Sequence alignment (STAR / HISAT2); ③Expression quantification (featureCounts / Salmon); ④Differential analysis, using DESeq2 for normalization and hypothesis testing for differential analysis. The metabolomics process covers: ①Raw mass spectrometry peak extraction (XCMS); ②Peak alignment and annotation (CAMERA); ③Normalization (PQN method). All process parameters (such as the p-value threshold for differential expression analysis and the PCA dimensionality reduction dimension) can be configured through a visual form and support Docker containerization for execution to ensure environmental consistency. The preprocessing results (such as gene expression matrices and lists of differentially methylated regions) are stored in a MySQL database in a structured form, and at the same time, an interactive HTML report is generated, including data quality indicators (such as sequencing saturation curves and total ion current maps of mass spectrometry) and statistical summaries (such as the number of differentially expressed genes and metabolite pathway enrichment). The module also provides a "resume from breakpoint" function. When the task is interrupted due to hardware failure, it can resume running from the most recently completed step, significantly improving the robustness of large-scale data analysis.
[0035] Conduct cross-modal association modeling of chemical data and multi-omics data to achieve cross-modal association of chemical data and multi-omics data, including four core stages: (1) Chemical feature encoding: Based on the RDKit cheminformatics toolkit, convert the SMILES structural formula of EDCs into a 2048-bit Morgan fingerprint feature vector (radius = 3) to capture molecular structure features (such as benzene rings and hydroxyl functional groups), and at the same time calculate molecular descriptors (such as topological polar surface area and number of hydrogen bond donors) as supplementary feature vectors.
[0036] (2)Omics feature dimensionality reduction: Perform principal component analysis (PCA) on high-dimensional omics data (such as a 20,000 gene expression matrix), determine to retain 500 principal components using the criterion of cumulative variance contribution rate ≥ 95%, and verify the biological consistency of the data after dimensionality reduction through t-SNE visualization (such as samples clustering according to the exposure dose gradient).
[0037] (3)Cross-modal alignment: Design a dual-stream deep neural network (Dual-Stream DNN). The chemical data stream processes fingerprint vectors using a 3-layer fully connected network, and the omics data stream uses a 1D convolutional network to extract features after PCA dimensionality reduction. Finally, learn the association weights between modalities through a cross-attention mechanism, and output low-dimensional representations in the joint embedding space.
[0038] (4)Association network mining: Build a distributed graph computing engine based on the Graph API of Elasticsearch, perform subnet discovery on the fused features, use the modularity optimization algorithm (Louvain) to identify functional modules related to EDCs (such as the "bisphenol A - estrogen receptor - cell proliferation pathway"), and persist the results to the Neo4j graph database through the Cypher query language. This engine supports GPU acceleration and can complete the pattern mining of millions of molecule-gene association pairs within 10 minutes.
[0039] Visualization layer: Achieve dynamic interactive visualization of various graphs. Based on D3.js, implement a force-directed graph to display the cross-modal association results of chemical data and multi-omics data, supporting node dragging, zooming, and detailed information floating windows. Render large-scale omics heatmaps through HTML5 Canvas + WebGL.
[0040] The core technologies of the visualization layer include: ① Force-directed graph layout: Use the Barnes-Hut algorithm to optimize the rendering performance of large-scale networks (>10,000 nodes), support users to adjust the view in real time through dragging and zooming, map node colors to compound types (EDCs are red, metabolites are blue), and the edge width reflects the association strength (Pearson correlation coefficient).
[0041] ② Dynamic data binding: Achieve real-time connection with the Neo4j database through WebSocket. After the user selects exposure conditions (such as "concentration > 10 μM & time > 24 h"), the system automatically generates a Cypher query (such as MATCH (c:Chemical)-[r:AFFECTS]->(g:Gene) WHERE r.dose>10 RETURN c,g,r), and renders the subnet on the front end.
[0042] ③ Multi-level interaction: Clicking on a node can pop up a floating window to display detailed information (such as GO annotations of genes, PubChem links of EDCs), and the right-click menu supports path highlighting (such as showing the shortest path from a certain compound to a disease phenotype), module export (in SVG / PNG format), and data drilling down (such as viewing the differential expression heatmap of associated genes). In addition, the module integrates WebGL acceleration technology to smoothly present the cell-gene co-occurrence network of single-cell transcriptomics (>100,000 nodes).
[0043] Example 2 This example combines high-performance computing hardware with a modular software architecture to efficiently process and analyze multi-omics data of environmental endocrine disruptors (EDCs). The core of the hardware uses a Dell PowerEdge R750xa server, equipped with dual Intel Xeon Gold 6338 processors (each processor has 32 cores / 64 threads, with a main frequency of 2.2 GHz), and is equipped with 256 GB of DDR4 memory to support large-scale concurrent computing tasks. Four NVIDIA A100 GPUs (each with 80 GB of video memory) are selected as the graphics processing unit to accelerate the training of deep learning models and graph computing tasks. The data storage adopts a hierarchical strategy: 10 TB of NVMe solid-state drives are used for high-speed reading and writing of hot data (such as intermediate results of real-time analysis), while 50 TB of mechanical hard drives are used for long-term archiving of cold data (such as raw sequencing data). The network communication is based on a gigabit Ethernet architecture, and the communication permissions of MySQL (port 3306), Neo4j (port 7474), and Elasticsearch (port 9200) are opened through firewall rules to ensure the security of data interaction.
[0044] The software environment is deployed on the Ubuntu 22.04 LTS operating system. The database system selects MySQL 8.0.32 (with the storage engine InnoDB and the transaction isolation level set to READ COMMITTED) and Neo4j 5.11.0 (integrating the APOC plugin 4.4.0.15 to support advanced graph algorithms). The programming language mainly uses Python 3.10.12, relying on RDKit 2023.03.1 for chemical feature calculations, and PyTorch 2.0.1 for building deep learning models. The front-end visualization is based on JavaScript ES6, using D3.js 7.8.2 to achieve dynamic network rendering, and Cytoscape.js 3.23.0 to support complex interactions. The system realizes the containerized encapsulation of the analysis process through Docker 24.0.5, and pre-installs the Snakemake 7.32.4 workflow engine and R 4.3.1 (integrating the Bioconductor 3.17 bioinformatics tool kit) to ensure the reproducibility and cross-platform compatibility of the analysis environment.
[0045] Step 1: Data collection; Data collection is implemented relying on the data access layer, which is divided into two parts: public data access and user-defined data upload. Public data access is achieved through standardized API interfaces. Among them, the NCBI EUtils API (version 2.0) is used to retrieve RNA-seq and ChIP-seq datasets related to EDCs from the GEO database, while the ArrayExpress REST API (version 1.7) supports batch downloading of metabolomics and epigenomics data. Users can input compound names (such as "Bisphenol A"), species classification IDs (such as human Taxonomy ID 9606), and experiment types, and the system will automatically parse and download the corresponding metadata (such as GSE numbers, sample descriptions) and raw data files (such as FASTQ and mzML). To improve data transfer efficiency, FASTQ files are accelerated for downloading through the aspera command-line tool (version 3.11.1), with the speed limit set to 500 Mbps.
[0046] User-defined data upload supports multiple formats, including CSV, JSON, and XML. CSV files need to be encoded in UTF-8, with the first row being column names (such as sample_id, dose), and numerical fields must conform to the preset range (such as the exposure dose unit being μM). JSON files need to follow the predefined Schema specification (Draft-07), and field integrity verification is performed through the jsonschema library (version 4.18.0). The uploaded data is converted into the Parquet columnar storage format (Apache Parquet 2.9.0) through the ETL process, with the partition key being the experiment ID (in the format such as GSE12345_SAMN678901), and the SNAPPY compression algorithm is used to optimize storage efficiency. Data lineage management is achieved through the PROV-O model (W3C standard), recording the source, transformation path, and version information of the data, and storing it in the data_provenance table in MySQL to ensure the traceability of the experiment.
[0047] Step 2: Data storage; A hybrid database architecture is adopted to achieve multi-modal storage and management of chemical data and multi-omics data of environmental endocrine disruptors. The chemical data and multi-omics data of environmental endocrine disruptors are divided into structured data and unstructured data; Step 3: Integrated analysis; Step 3-1: Preprocess the chemical data and multi-omics data of environmental endocrine disruptors; Taking transcriptome analysis as an example, this process is built based on the Snakemake workflow engine and mainly includes multiple key steps.
[0048] First is the quality control of raw data. The fastp tool (version 0.23.2) is used to filter low-quality read data, ensuring a Phred score ≥ 20, and at the same time removing adapter sequences to improve the data quality. Next is sequence alignment. The STAR tool (version 2.7.10b) is used to map the read data to the GENCODE human release 43 reference genome, setting the maximum mismatch rate to 10% to ensure the accuracy of the alignment results.
[0049] In the expression quantification stage, featureCounts (version 2.0.3) is used to count the gene count matrix, and the annotation file used is gencode.v43.annotation.gtf to provide reliable gene expression data. Subsequently, differential analysis is performed, and DESeq2 (version 1.38.3) is used for normalization and hypothesis testing.
[0050] For the metabolome analysis process, first, mass spectrometry peak extraction is carried out. The centWave algorithm of XCMS (version 3.20.0) is used to detect chromatographic peaks, setting the ppm tolerance to 10 and the peak width range to 5 - 20 seconds. Then, CAMERA (version 1.56.0) is used for peak alignment and annotation, matching the HMDB 5.0 database to identify metabolite features.
[0051] All analysis processes are encapsulated in a Docker image (biocontainers / snakemake:7.32.4). In terms of resource limitations, the single-task memory is set to 64 GB, the CPU cores are 16, and GPU acceleration is supported (--gpus all).
[0052] The preprocessing results are stored in a MySQL database in a structured format, and at the same time an interactive HTML report is generated. The report contains a sequencing saturation curve, a total ion chromatogram of the mass spectrum, and a list of differential genes (FDR < 0.05), providing comprehensive support for subsequent data analysis.
[0053] Step 3 - 2: The multimodal data fusion engine module performs cross-modal association modeling on chemical data and multi-omics data to achieve cross-modal association of chemical data and multi-omics data; In the chemical data encoding stage, the GetMorganFingerprintAsBitVect function of RDKit is used to generate 2048-bit Morgan fingerprints (radius 3) to capture molecular structure features such as benzene rings and hydroxyl groups. At the same time, descriptors such as LogP and the number of hydrogen bond donors are calculated, and the results are stored in JSON format. In the omics feature dimensionality reduction stage, principal component analysis (PCA, scikit-learn version 1.2.2) is used to reduce the gene expression matrix to 500 dimensions, ensuring that the cumulative variance contribution rate reaches or exceeds 95%. Subsequently, the t-SNE method is used to verify the biological clustering effect of the dimensionality-reduced data.
[0054] In the cross-modal alignment stage, a two-stream deep neural network architecture is designed, which consists of a chemical data stream, a multi-omics data stream, and a cross-modal fusion mechanism. Among them, the chemical data stream adopts a multi-layer fully connected network structure, including an input dimension of 1024, a hierarchical structure of hidden layers [512, 256], and is equipped with a ReLU activation function and a BatchNorm1d normalization layer after each layer. At the same time, a Dropout rate of 0.2 is used to provide regularization, and finally, a 128-dimensional feature vector is output; the omics data stream processes sequence data using a one-dimensional convolutional neural network. A two-layer convolutional structure with an input dimension of 200 is designed, and the number of channels is [64, 128] respectively. A 3×3 convolutional kernel is used, combined with max pooling and adaptive average pooling layers, and finally, a 128-dimensional feature vector is obtained through linear projection. The features of the two data streams are cross-modally aligned through a multi-head attention mechanism (4 attention heads, 128-dimensional embedding space, a dropout rate of 0.1). After that, it passes through a feature fusion layer, which includes LayerNorm normalization, linear dimensionality reduction, and a Tanh activation function, and finally generates a unified 128-dimensional feature vector. This architecture design can not only process the specific features of the two modal data respectively but also adaptively learn the associations between modalities through the attention mechanism. At the same time, through multi-level regularization measures, the stability and generalization ability of the model are ensured, supporting both supervised and unsupervised training modes, and showing good feature extraction and modal fusion performance on the example data.
[0055] In the associated network mining stage, an Elasticsearch cluster (consisting of 3 nodes and 10 shards) is constructed to implement a distributed graph engine, and the Louvain algorithm (resolution set to 1.0) is used to identify functional modules. The identification results will be persisted to the Neo4j database for subsequent analysis.
[0056] Step 4: Visualization; The visualization module is developed based on the D3.js framework, aiming to support the construction of dynamic networks and interactive exploration. This module adopts a force-directed layout and uses the Barnes-Hut approximation algorithm (parameter θ = 0.8) to optimize the interaction between nodes. In the visualization, the color mapping of nodes represents the compound types, where environmental hormone disruptors (EDCs) are highlighted in red, and the width of the edges reflects the association strength, which is quantified by calculating the Pearson correlation coefficient.
[0057] Users can connect to the Neo4j database in real time via WebSocket, input Cypher queries (such as filtering the associations between EDCs with a dose greater than 10 μM and genes), and the front-end interface will automatically render the corresponding sub-network and support operations such as dragging and zooming of nodes. In the detailed information pop-up window, users can view the GO annotations of genes and the PubChem links of compounds. In addition, the right-click menu provides functions such as path highlighting (such as the shortest toxicity pathway), module export options (supporting SVG / PNG formats), and data drilling-down functions for more in-depth analysis.
[0058] To improve the rendering performance, this module can handle networks with 100,000-level nodes, adopting the WebGL acceleration technology of Three.js (r152 version) to ensure a smooth user experience with an interactive frame rate reaching or exceeding 30 FPS. This design makes the visualization of large-scale data more efficient and intuitive.
[0059] As Figure 3 shown, it is a schematic diagram of the online function usage of a method for integrating and analyzing multi-omics data of environmental endocrine disruptors in this application. Figure 3 In it, a are function items for selectable execution tasks, including chemical data and multi-omics data preprocessing; b is the interface for performing differential gene expression analysis tasks; c is the option for selecting imported data; d is the interface for displaying task execution results.
[0060] As Figure 4 shown, it is an analysis platform designed using a method for integrating and analyzing multi-omics data of environmental endocrine disruptors in this application. The figure shows the settings of the visualization layer for visualizing the results of the integration analysis layer.
[0061] As Figure 5 shown, it is Figure 4 the execution interface of the functional module of Transcriptional Factors in
[0062] This application discloses a multi-omics data integration and analysis system for environmental endocrine disruptors. Through data collection, data storage, integrated analysis, and visualization, it constructs a dedicated metadata framework for environmental endocrine disruptors and establishes a full-dimensional collection system covering chemical data and multi-omics data. This application supports automated conversion and structured storage of multiple data formats, thereby significantly improving data traceability efficiency and shortening the data preprocessing cycle.
[0063] The dual-stream deep neural network disclosed in this application can not only process the specific features of two types of modal data respectively, but also adaptively learn the associations between modalities through the attention mechanism. At the same time, through multi-level regularization measures, it ensures the stability and generalization ability of the model, supports both supervised and unsupervised training modes, and demonstrates good feature extraction and modal fusion performance on the example data.
[0064] This application integrates heterogeneous network data from different experimental conditions, species, or omics levels through multi-layer network modeling, realizes the functions of multi-modal data integration analysis and visualization, and promotes the prediction and query parsing of the toxicity mechanism of EDCs.
[0065] The above has made a detailed description of the preferred implementation mode of this application. However, this application is not limited to the above implementation mode. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can also be made without departing from the purpose of this application.
[0066] Many other changes and modifications can be made without departing from the concept and scope of this application. It should be understood that this application is not limited to a specific implementation mode, and the scope of this application is defined by the appended claims.
Claims
1. A method for integrated analysis of multi-omics data of environmental endocrine disruptors, characterized in that It includes the following steps: Step 1: Data collection, extracting chemical data and multi-omics data of environmental endocrine disruptors from a public database; Step 2: Data storage; Adopt a hybrid database architecture to achieve multi-modal storage and management of the chemical data and multi-omics data of environmental endocrine disruptors. The chemical data and multi-omics data of environmental endocrine disruptors are divided into structured data and unstructured data; For structured data, construct a normalized metadata management system based on the MySQL relational database, design core data tables through entity-relationship modeling, and store the structured data; For unstructured data, construct a dynamic knowledge network based on the Neo4j graph database and store the unstructured data; Step 3: Integrated analysis; Step 3-1: Preprocess the chemical data and multi-omics data of environmental endocrine disruptors; Step 3-2: Conduct cross-modal association modeling on the chemical data and multi-omics data to achieve cross-modal association of the chemical data and multi-omics data; Step 4: Visualization; Visualize the cross-modal association results of the chemical data and multi-omics data for information query and toxicity mechanism prediction of environmental endocrine disruptors.
2. The method for integrated analysis of multi-omics data of environmental endocrine disruptors according to claim 1, wherein The data collection in Step 1 is specifically as follows: Use Django REST Framework to build a standardized API interface, support data upload in JSON / CSV / XML formats, and at the same time integrate the Scrapy-Redis distributed crawler framework to directionally crawl the chemical data and multi-omics data of environmental endocrine disruptors in the public database; The chemical data includes compound tables, SMILES structural formulas, inhibition constants, half-maximal effective concentration, and half-maximal inhibitory concentration; The multi-omics data includes transcriptome, metabolome, genome, proteome, epigenome, lipidome, spatial transcriptome, single-cell transcriptome, and exposure records.
3. The method for integrated analysis of multi-omics data of environmental endocrine disruptors according to claim 2, wherein In Step 3-1, preprocess the chemical data and multi-omics data of environmental endocrine disruptors. The preprocessing of the transcriptome of the multi-omics data includes the following steps: Step 3-1-1: Raw data quality control, use the fastp tool to filter the collected transcriptome data to ensure that the Phred score ≥ 20, and at the same time excise the adapter sequences; Step 3-1-2: Sequence alignment, map the read transcriptome data to the reference genome through the STAR tool, and set the maximum mismatch rate to 10%; Step 3-1-3: Expression quantification, use featureCounts to count the gene count matrix for gene expression of the transcriptome data; Step 3-1-4: Differential analysis, use DESeq2 for normalization and hypothesis testing for differential analysis.
4. The method for integrated analysis of multi-omics data of environmental endocrine disruptors according to claim 2, wherein In Step 3-1, preprocess the chemical data and multi-omics data of environmental endocrine disruptors. The preprocessing of the metabolome data includes the following steps: Step 3-1-1: Extract raw mass spectrometry peaks, use the centWave algorithm of XCMS to detect the chromatographic peaks of the metabolome data, set the ppm tolerance to 10, and the peak width range to 5 - 20 seconds; Step 3-1-2: Peak alignment and annotation. Use CAMERA to perform peak alignment and annotation of chromatograms, match the HMDB 5.0 database to identify metabolomic data features; Step 3-1-3: Normalization.
5. The multi-omics data integration and analysis method for environmental endocrine disruptors according to claim 1, characterized in that The said Step 3-2 includes the following steps: Step 3-2-1: Chemical data encoding: Based on the RDKit chemoinformatics toolkit, convert the chemical data of environmental endocrine disruptors into 2048-bit Morgan fingerprint feature vectors to capture the molecular structure features of chemical data, and at the same time calculate molecular descriptors as supplementary feature vectors to obtain chemical data streams; Step 3-2-2: Multi-omics data dimensionality reduction: Perform principal component analysis dimensionality reduction on high-dimensional multi-omics data, use the criterion of cumulative variance contribution rate ≥ 95% to determine to retain 500 principal components, and verify the biological consistency of the multi-omics data after dimensionality reduction through t-SNE visualization to obtain multi-omics data streams; Step 3-2-3: Cross-modal alignment: Construct a two-stream deep neural network. The chemical data stream processes fingerprint feature vectors using a 3-layer fully connected network, and the multi-omics data stream uses a 1D convolutional network to extract features after principal component analysis dimensionality reduction. Finally, learn the cross-modal association weights through a cross-attention mechanism and output low-dimensional feature vectors in the joint embedding space to achieve cross-modal association of chemical data and multi-omics data.
6. The method for integrative analysis of multi-omics data of environmental endocrine disruptors according to claim 5, wherein The said two-stream deep neural network consists of a chemical data stream, a multi-omics data stream, and a cross-modal fusion mechanism. The chemical data stream adopts a multi-layer fully connected layer structure, including a hierarchical structure with an input dimension of 1024 and hidden layers [512, 256], and is equipped with a ReLU activation function and a BatchNorm1d normalization layer after each layer. At the same time, a Dropout rate of 0.2 is used to provide regularization, and finally outputs a 128-dimensional feature vector; the multi-omics data stream uses a one-dimensional convolutional neural network to process sequence data, designs a two-layer convolutional structure with an input dimension of 200, the number of channels is [64, 128] respectively, uses a 3×3 convolutional kernel, combines max pooling and adaptive average pooling layers, and finally obtains a 128-dimensional feature vector through linear projection; the feature vectors of the chemical data stream and the multi-omics data stream are cross-modally aligned through a multi-head attention mechanism, and then pass through a feature fusion layer. The feature fusion layer includes LayerNorm normalization, linear dimensionality reduction, and a Tanh activation function to generate a unified 128-dimensional feature vector.
7. The method for integrative analysis of multi-omics data of environmental endocrine disruptors according to claim 6, wherein The construction steps of the said two-stream deep neural network are as follows: Step 3-2-3-1: Chemical data stream processing: 1) Fully connected layer calculation: H_l = ReLU(BN(W_l·H_{l - 1}+b_l)); Where: H_l is the output of the l-th layer; H_{l - 1} is the output of the (l - 1)-th layer; W_l is the weight matrix; b_l is the bias vector; BN is batch normalization; ReLU(x)=max(0,x); 2) Dropout regularization: H_dropout = H*mask; mask ~ Bernoulli(p); Where: H_dropout is the output feature vector, the feature after Dropout processing; H is the input feature vector, which is the output activation value from the previous layer, i.e., the output of the fully connected layer; mask is a binary mask matrix; p is the retention probability, which controls the regularization strength; Bernoulli(p) is the Bernoulli distribution; Step 3-2-3-2: Multi-omics data stream processing: 1) One-dimensional convolution operation: Conv1D(x)=σ(W*x + b); Where: Conv1D(x) is the one-dimensional convolution; σ is the ReLU activation function; W is the convolution kernel weight; * is the convolution operation; b is the bias term; 2) Max pooling: MaxPool(x)=max(x[i:i + k]); Where: MaxPool(x) performs max pooling on the input data x and returns the maximum value within each pooling window; k is the pooling window size; i is the starting index of the current window; x is the input data; Step 3-2-3-3: Multi-head attention mechanism: 1) Query Q, key K, and value V calculation: Q = X_q·W_q; K = X_k·W_k; V = X_v·W_v; Attention weight calculation: Attention(Q, K, V) = softmax(QK^T / √d_k)V; Where: d_k is the attention dimension; softmax is the normalization function; X_q is the input query Q representation; W_q is the weight matrix of the query; X_k is the input key K representation; W_k is the weight matrix of the key; X_v is the input value V representation; W_v is the weight matrix of the value; T is the transpose operation; K^T is the transpose of the key matrix K; 2) Multi-head attention: MultiHead(Q, K, V)=Concat(head_1,..., head_h)W_o; head_i = Attention(QW_i^Q, KW_i^K, VW_i^V); Where: Concat(head_1,..., head_h) concatenates h heads, and the dimension after concatenation will increase, representing the parallel calculation results of multiple attention heads; W_o is the linear transformation matrix, which is used to transform the result after concatenating multiple heads to obtain the final output; head_i is the output of the i-th attention head; Attention(QW_i^Q, KW_i^K, VW_i^V) is for the i-th attention head, which linearly transforms the query Q, key K, and value V using different weight matrices respectively, and then calculates the standard attention; W_i^Q represents the query Q weight matrix of the i-th head, which linearly transforms the query Q and maps the original query to a new representation space; W_i^K represents the key K weight matrix of the i-th head, which linearly transforms the key K and maps the original key to a new representation space; W_i^V represents the value V weight matrix of the i-th head, which linearly transforms the value V and maps the original value to a new representation space; Step 3-2-3-4: Feature fusion: 1) Feature concatenation: H_fused = [H_chem + Attn_out; H_omics]; Where: H_fused is formed by splicing to form a new feature vector; H_chem is a feature vector related to chemistry; Attn_out is the output from the attention mechanism; H_omics is a feature vector generated based on other omics data; + is to perform an addition operation on H_chem and Attn_out; is to perform splicing in the feature dimension, splicing H_chem + Attn_out and H_omics column by column to form a new feature vector H_fused; 2) LayerNorm normalization: μ=(1 / H)ΣH_i; σ =sqrt((1 / H)Σ(H_i - μ)^2); LN(H)=γ*(H - μ) / σ+β; H_i is the value of the i-th element in the input feature vector H; μ is the arithmetic mean of all elements of the input feature vector H; σ is the standard deviation of the input feature vector H, measuring the degree of data dispersion; sqrt is the square root; LN(H) is the output after normalizing the input feature vector H; γ is a scaling parameter, controlling the scale of the output after normalization; β is an offset parameter, adjusting the position of the output after normalization.
8. An environmental endocrine disruptor multi-omics data integration and analysis system, characterized in that, For the implementation of the method for integrated analysis of environmental endocrine disruptor multi-omics data according to any one of claims 1 to 7, it includes a data access layer, a data storage layer, an integrated analysis layer, and a visualization layer. The data access layer is used to achieve the standardized collection and upload of multi-source data. The standardized API interface built based on Django RESTFramework supports the upload of data in JSON / CSV / XML formats; the integrated Scrapy-Redis distributed crawler framework is used to directionally crawl environmental endocrine disruptor-related data in public databases; the browser-side large file sharding upload function implemented based on HTML5 FileAPI, with a shard size of 10MB. The data storage layer uses a MySQL relational database and a Neo4j graph database for multi-modal storage and management of the chemical data and multi-omics data of environmental endocrine disruptors. The integrated analysis layer is used to preprocess the chemical data and multi-omics data of environmental endocrine disruptors and perform cross-modal association modeling to achieve cross-modal association of chemical data and multi-omics data. The visualization layer is used to visualize the cross-modal association results of chemical data and multi-omics data, and realizes dynamic interactive visualization based on the D3.js framework, for information query and toxicity mechanism prediction of environmental endocrine disruptors.
Citation Information
Patent Citations
Medical big data center based on hybrid database structure and construction method thereof
CN112199425A
Construction method of multi-modal information fusion medical big data online service platform
CN117790005A
Single-cell multi-omics data cross-modal generation method and system
CN118173173A
Single-cell multi-omics clustering method based on depth information fusion
CN119763677A
Single cell cross-omics multi-modal prediction method
CN119889434A