A system and method for integrated analysis of multi-omics data of environmental endocrine disruptors

Through a hybrid database architecture and a dual-stream deep neural network, data integration problems in multiomics research on environmental endocrine disruptors are solved, efficient data traceability and cross-modal correlation are achieved, and multimodal data visualization and toxicity mechanism prediction are supported, and data utilization efficiency and analysis capabilities are improved.

CN120356536BActive Publication Date: 2025-08-29INNER MONGOLIA UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510848718.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-29
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

In the prior art, multiomic research on environmental endocrine disruptors (EDCs) lacks a standardized metadata framework and a systematic data governance architecture, resulting in inefficient data traceability, limited cross-platform data collaborative utilization, and lack of a data integration paradigm for EDCs, making it difficult to accurately analyze toxicity pathways and personalized risk assessment.

Method used

The hybrid database architecture is adopted, combined with MySQL relational database and Neo4j graph database, and multimodal storage and management of chemical data and multiomic data of environmental endocrine disturbances is realized, and cross-modal correlation modeling is carried out through dual-stream deep neural network, supporting data upload in JSON/CSV/XML format, and using Django REST Framework to build a standardized API interface and Scrapy-Redis crawler framework for data acquisition and preprocessing, and using specific tools for data quality control and dimensionality reduction, ultimately realizing cross-modal correlation and visualization of chemical data and multiomic data.

Benefits of technology

It significantly improves the efficiency of data traceability, shortens the data preprocessing cycle, can process the characteristics of multimodal data and adaptively learn the correlation between modals, supports supervised and unsupervised training modes, realizes multimodal data integration analysis and visualization, and promotes the prediction and query analysis of EDCs toxicity mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356536B_ABST
    Figure CN120356536B_ABST
Patent Text Reader

Abstract

This application belongs to the field of biological science research and big data mining technology, and discloses a system and method for integrating and analyzing multi-omics data of environmental endocrine disruptors. Through data collection, data storage, integrated analysis and visualization, a dedicated metadata framework for environmental endocrine disruptors is constructed, and a full-dimensional collection system covering chemical data and multi-omics data is established to realize multimodal data integration analysis and visualization functions, thereby promoting the prediction and query analysis of EDCs toxicity mechanisms. This application supports automatic conversion and structured storage of multiple data formats, thereby significantly improving data traceability efficiency and shortening the data preprocessing cycle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of biological science research and big data mining technology, and specifically relates to a system and method for integrating and analyzing multi-omics data of environmental endocrine disruptors. Background Art

[0002] Currently, the multi-omics research system for environmental toxicology has formed several representative technical platforms, among which the Comparative Toxicogenomics Database (CTD), developed by the National Institute of Environmental Health Sciences (NIEHS), is particularly prominent. This platform systematically extracts and integrates toxicogenomics data from published scientific literature, constructing a knowledge network encompassing the multidimensional interactions between chemicals, genes, diseases, and biological processes. Meanwhile, internationally renowned bioinformatics repositories such as the National Center for Biotechnology Information (NCBI) and the European Molecular Biology Laboratory Bioinformatics Institute (EMBL-EBI), through their core databases such as Gene Expression Omnibus (GEO) and ArrayExpress, continue to accumulate and standardize a large number of high-throughput sequencing datasets related to EDCs, including transcriptomics, epigenomics, and other multi-omics toxicity testing data.

[0003] The current multi-omics research system for environmental endocrine disrupting chemicals (EDCs) has two core bottlenecks:

[0004] (1) For multi-omics research on environmental endocrine disrupting chemicals (EDCs), current technologies have not yet established a standardized metadata framework and a systematic data governance architecture, and raw data are generally stored in a decentralized manner. The lack of standardized conversion mechanisms for multi-source heterogeneous data such as chemical exposure profiles, transcriptomics, and epigenetics has led to reduced data traceability efficiency and a significant increase in the time required for structured data extraction processes. In addition, the data silos between different research institutions have seriously restricted the collaborative use of cross-platform data.

[0005] (2) Although existing public databases (such as CTD, NCBI, and EEMBL-EBI) have achieved standardized storage of basic data, they lack data integration paradigms for EDCs and their specific exposure conditions (such as dose-effect relationships and time-dependent exposure), biological target systems (such as molecular networks across species / organs / cell levels), and toxicity endpoints. This technical shortcoming directly restricts the accurate analysis of EDCs toxicity pathways and also makes the development of personalized risk assessment models based on multi-source data challenging. Therefore, the development of a multi-omics data integration and analysis system and method for environmental endocrine disruptors has great market prospects. Summary of the Invention

[0006] The purpose of this application is to solve the problems of the prior art and to provide a system and method for integrated analysis of multi-omics data of environmental endocrine disruptors.

[0007] In order to solve the technical problem, the technical solution of this application is: a method for integrated analysis of multi-omics data of environmental endocrine disruptors, comprising the following steps:

[0008] Step 1: Data collection: extract chemical data and multi-omics data of environmental endocrine disruptors from public databases;

[0009] Step 2: Data storage;

[0010] A hybrid database architecture is used to achieve multimodal storage and management of chemical data and multi-omics data of environmental endocrine disruptors. The chemical data and multi-omics data of environmental endocrine disruptors are divided into structured data and unstructured data.

[0011] For structured data, a normalized metadata management system is built based on the MySQL relational database. Core data tables are designed through entity-relationship modeling to store structured data.

[0012] For unstructured data, a dynamic knowledge network is built based on the Neo4j graph database to store unstructured data;

[0013] Step 3: Integrated analysis;

[0014] Step 3-1: Preprocessing chemical data and multi-omics data of environmental endocrine disruptors;

[0015] Step 3-2: Perform cross-modal correlation modeling on chemical data and multi-omics data to achieve cross-modal correlation between chemical data and multi-omics data;

[0016] Step 4: Visualization;

[0017] Visualize cross-modal correlation results of chemical data and multi-omics data for information query and toxicity mechanism prediction of environmental endocrine disruptors.

[0018] Preferably, the data collection in step 1 is specifically as follows: using Django REST Framework to build a standardized API interface, supporting data uploading in JSON / CSV / XML formats, and integrating the Scrapy-Redis distributed crawler framework to crawl chemical data and multi-omics data of environmental endocrine disruptors from public databases;

[0019] The chemical data includes compound table, SMILES structural formula, inhibition constant, half maximal effect concentration and half inhibitory concentration;

[0020] The multi-omics data include transcriptome, metabolome, genome, proteome, epigenome, lipidome, spatial transcriptome, single-cell transcriptome and exposure record table.

[0021] Preferably, in step 3-1, the chemical data and multi-omics data of environmental endocrine disruptors are preprocessed, wherein the transcriptome preprocessing of the multi-omics data comprises the following steps:

[0022] Step 3-1-1: Quality control of raw data: use fastp tool to filter the collected transcriptome data to ensure that the Phred score is ≥20, and remove the adapter sequence;

[0023] Step 3-1-2: Sequence alignment: Use the STAR tool to map the read transcriptome data to the reference genome, setting the maximum mismatch rate to 10%;

[0024] Step 3-1-3: Expression quantification, using featureCounts statistical gene count matrix to perform gene expression analysis on transcriptome data;

[0025] Step 3-1-4: Differential analysis, standardization and hypothesis testing were performed using DESeq2 to perform differential analysis.

[0026] Preferably, in step 3-1, the chemical data and multi-omics data of environmental endocrine disruptors are preprocessed, wherein the metabolomics data preprocessing comprises the following steps:

[0027] Step 3-1-1: Extract the original mass spectrum peaks. Use the centWave algorithm of XCMS to detect the chromatographic peaks of the metabolomics data. Set the ppm tolerance to 10 and the peak width range to 5~20 seconds.

[0028] Step 3-1-2: Peak alignment and annotation: Use CAMERA to align and annotate chromatographic peaks and match them against the HMDB 5.0 database to identify metabolomics data features.

[0029] Step 3-1-3: Normalization.

[0030] Preferably, the step 3-2 includes the following steps:

[0031] Step 3-2-1: Chemical data encoding: Based on the RDKit chemical data science toolkit, the chemical data of environmental endocrine disruptors are converted into 2048-bit Morgan fingerprint feature vectors to capture the molecular structure characteristics of the chemical data. At the same time, molecular descriptors are calculated as supplementary feature vectors to obtain chemical data streams.

[0032] Step 3-2-2: Multi-omics data dimensionality reduction: Principal component analysis (PCA) was performed on the high-dimensional multi-omics data to reduce the dimensionality. The criterion of cumulative variance contribution rate ≥ 95% was used to determine the retention of 500 principal components. The biological consistency of the multi-omics data after dimensionality reduction was verified by t-SNE visualization to obtain the multi-omics data flow.

[0033] Step 3-2-3: Cross-modal alignment: Construct a two-stream deep neural network. The chemical data stream uses a three-layer fully connected network to process the fingerprint feature vector. The multi-omics data stream uses a 1D convolutional network to extract the features after principal component analysis dimensionality reduction. Finally, the cross-attention mechanism is used to learn the inter-modal association weights and output the low-dimensional feature vector in the joint embedding space to achieve cross-modal association of chemical data and multi-omics data.

[0034] Preferably, the dual-stream deep neural network consists of a chemical data stream, a multi-omics data stream and a cross-modal fusion mechanism. The chemical data stream adopts a multi-layer fully connected layer structure, including an input dimension of 1024 and a hidden layer of [512, 256]. Each layer is equipped with a ReLU activation function and a BatchNorm1d normalization layer, and a Dropout rate of 0.2 is used to provide regularization, and finally outputs a 128-dimensional feature vector; the multi-omics data stream uses a one-dimensional convolutional neural network to process sequence data, and designs a two-layer convolution structure with an input dimension of 200. The number of channels is [64, 128], using a 3×3 convolution kernel, combined with maximum pooling and adaptive average pooling layers, and finally obtains a 128-dimensional feature vector through linear projection; the feature vectors of the chemical data stream and the multi-omics data stream are cross-modally aligned through a multi-head attention mechanism, and then pass through a feature fusion layer. The feature fusion layer includes LayerNorm normalization, linear dimensionality reduction and Tanh activation function to generate a unified 128-dimensional feature vector.

[0035] Preferably, the steps of constructing the dual-stream deep neural network are as follows:

[0036] Step 3-2-3-1: Chemical data stream processing:

[0037] 1) Fully connected layer calculation:

[0038] H_l=ReLU(BN(W_l·H_{l-1}+b_l));

[0039] in:

[0040] H_l is the output of layer l;

[0041] H_{l-1} is the output of the l-1th layer;

[0042] W_l is the weight matrix;

[0043] b_l is the bias vector;

[0044] BN is batch normalization;

[0045] ReLU(x)=max(0,x);

[0046] 2) Dropout regularization:

[0047] H_dropout=H*mask;

[0048] mask~Bernoulli(p);

[0049] in:

[0050] H_dropout is the output feature vector, the feature after Dropout processing;

[0051] H is the input feature vector, which is the output activation value from the previous layer, that is, the output of the fully connected layer;

[0052] mask is a binary mask matrix;

[0053] p is the retention probability, which controls the regularization strength;

[0054] Bernoulli(p) is the Bernoulli distribution;

[0055] Step 3-2-3-2: Multi-omics data flow processing:

[0056] 1) One-dimensional convolution operation:

[0057] Conv1D(x)=σ(W*x+b);

[0058] in:

[0059] Conv1D(x) is a one-dimensional convolution;

[0060] σ is the ReLU activation function;

[0061] W is the convolution kernel weight;

[0062] * is the convolution operation;

[0063] b is the bias term;

[0064] 2) Max Pooling:

[0065] MaxPool(x)=max(x[i:i+k]);

[0066] in:

[0067] MaxPool(x) performs maximum pooling on the input data x and returns the maximum value in each pooling window;

[0068] k is the pooling window size;

[0069] i is the starting index of the current window;

[0070] x is the input data;

[0071] Step 3-2-3-3: Multi-head attention mechanism:

[0072] 1) Query Q, key K, value V calculation:

[0073] Q=X_q·W_q;

[0074] K=X_k·W_k;

[0075] V=X_v·W_v;

[0076] Attention weight calculation:

[0077] Attention(Q,K,V) = softmax(QK^T / √d_k)V;

[0078] in:

[0079] d_k is the attention dimension;

[0080] Softmax is a normalization function;

[0081] X_q is the input query Q representation;

[0082] W_q is the weight matrix of the query;

[0083] X_k is the key K representation of the input;

[0084] W_k is the bond weight matrix;

[0085] X_v is the input value V;

[0086] W_v is the weight matrix of values;

[0087] T is the transpose operation;

[0088] K^T is the transpose of the key matrix K;

[0089] 2) Multi-head attention:

[0090] MultiHead(Q,K,V)=Concat(head_1,...,head_h)W_o;

[0091] head_i=Attention(QW_i^Q, KW_i^K, VW_i^V);

[0092] in:

[0093] Concat(head_1, ..., head_h) concatenates h heads. The concatenated dimensions will be larger, representing the parallel computation results of multiple attention heads.

[0094] W_o is a linear transformation matrix used to transform the result of splicing multiple heads to obtain the final output;

[0095] head_i is the output of the i-th attention head;

[0096] Attention(QW_i^Q, KW_i^K, VW_i^V) is for the i-th attention head, using different weight matrices to linearly transform the query Q, key K and value V, and then calculate the standard attention;

[0097] W_i^Q represents the query Q weight matrix of the i-th head, which performs a linear transformation on the query Q and maps the original query into the new representation space;

[0098] W_i^K represents the key K weight matrix of the i-th head, which performs a linear transformation on the key K and maps the original key to the new representation space;

[0099] W_i^V represents the weight matrix of the value V of the i-th head, which performs a linear transformation on the value V and maps the original value to the new representation space;

[0100] Step 3-2-3-4: Feature fusion:

[0101] 1) Feature stitching:

[0102] H_fused=[H_chem+Attn_out;H_omics];

[0103] in:

[0104] H_fused is the new feature vector formed by splicing;

[0105] H_chem is the chemical-related feature vector;

[0106] Attn_out is the output from the attention mechanism;

[0107] H_omics is a feature vector generated based on other omics data;

[0108] + is the addition operation of H_chem and Attn_out;

[0109] ; It is splicing on the feature dimension, splicing H_chem + Attn_out and H_omics by column to form a new feature vector H_fused;

[0110] 2) LayerNorm normalization:

[0111] μ=(1 / H)ΣH_i;

[0112] σ =sqrt((1 / H)Σ(H_i-μ)^2);

[0113] LN(H)=γ*(H-μ) / σ+β;

[0114] H_i is the value of the i-th element in the input feature vector H;

[0115] μ is the arithmetic mean of all elements of the input eigenvector H;

[0116] σ is the standard deviation of the input eigenvector H, which measures the degree of discreteness of the data;

[0117] sqrt is the square root;

[0118] LN(H) is the output after normalizing the input feature vector H;

[0119] γ is a scaling parameter that controls the scale of the normalized output;

[0120] β is the offset parameter that adjusts the position of the normalized output.

[0121] Preferably, a multi-omics data integration and analysis system for environmental endocrine disruptors is used to implement the above-mentioned multi-omics data integration and analysis method for environmental endocrine disruptors, comprising a data access layer, a data storage layer, an integration and analysis layer, and a visualization layer.

[0122] The data access layer is used to implement standardized collection and upload of multi-source data. It uses a standardized API interface built on Django REST Framework and supports data upload in JSON / CSV / XML formats. It integrates the Scrapy-Redis distributed crawler framework for targeted crawling of environmental endocrine disruptor-related data from public databases. It also supports browser-side large file fragment upload based on HTML5 FileAPI, with a fragment size of 10MB.

[0123] The data storage layer uses MySQL relational database and Neo4j graph database to perform multimodal storage and management of chemical data and multi-omics data of environmental endocrine disruptors;

[0124] The integrated analysis layer is used to preprocess the chemical data and multi-omics data of environmental endocrine disruptors and perform cross-modal association modeling to achieve cross-modal association between chemical data and multi-omics data;

[0125] The visualization layer is used to visualize the cross-modal association results of chemical data and multi-omics data, and realizes dynamic interactive visualization based on the D3.js framework for information query and toxicity mechanism prediction of environmental endocrine disruptors.

[0126] Compared with the prior art, the advantages of this application are:

[0127] (1) This application discloses a multi-omics data integration and analysis system for environmental endocrine disruptors. Through data collection, data storage, integrated analysis and visualization, a dedicated metadata framework for environmental endocrine disruptors is constructed, and a full-dimensional collection system covering chemical data and multi-omics data is established. This application supports automatic conversion and structured storage of multiple data formats, thereby significantly improving data traceability efficiency and shortening the data preprocessing cycle;

[0128] (2) The two-stream deep neural network disclosed in this application can not only process the specific features of the two modal data separately, but also adaptively learn the association between the modalities through the attention mechanism. At the same time, it ensures the stability and generalization ability of the model through multi-level regularization measures, supports both supervised and unsupervised training modes, and demonstrates good feature extraction and modality fusion performance on example data;

[0129] (3) This application integrates heterogeneous network data from different experimental conditions, species or omics levels through multi-layer network modeling, realizing multimodal data integration analysis and visualization functions, and promoting the prediction and query analysis of EDCs toxicity mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS

[0130] Figure 1 A flowchart of a multi-omics data integration and analysis method for environmental endocrine disruptors for this application;

[0131] Figure 2 This is a framework diagram of a multi-omics data integration and analysis system for environmental endocrine disruptors in this application;

[0132] Figure 3 A schematic diagram of the online functions of a multi-omics data integration analysis method for environmental endocrine disruptors in this application;

[0133] Figure 4An analytical platform designed for the application of a multi-omics data integration analysis method for environmental endocrine disruptors;

[0134] Figure 5 for Figure 4 Functional module execution interface of Transcriptional Factors. DETAILED DESCRIPTION

[0135] The present application is described in detail below with reference to the accompanying drawings and specific examples, but the present application is not limited to these examples. This application covers any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this application. To provide a thorough understanding of this application, specific details are described in detail in the following examples of this application, but those skilled in the art can fully understand this application without these detailed descriptions.

[0136] like Figure 1 As shown, the present application discloses a method for integrated analysis of multi-omics data of environmental endocrine disruptors, comprising the following steps:

[0137] Step 1: Data collection: extract chemical data and multi-omics data of environmental endocrine disruptors from public databases;

[0138] For example, chemical molecules and biological activity data of environmental endocrine disruptors are extracted from BindingDB, ChEMBL, and PubChem public databases, and multi-omics data and their corresponding experimental record information are extracted from ToxCast, Tox 21, EMBL-EBI, NCBI public databases, and the China National Genome Science Data Center.

[0139] Step 2: Data storage;

[0140] A hybrid database architecture is used to achieve multimodal storage and management of chemical data and multi-omics data of environmental endocrine disruptors. The chemical data and multi-omics data of environmental endocrine disruptors are divided into structured data and unstructured data.

[0141] For structured data, a normalized metadata management system is built based on the MySQL relational database. Core data tables are designed through entity-relationship modeling to store structured data.

[0142] For unstructured data, a dynamic knowledge network is built based on the Neo4j graph database to store unstructured data;

[0143] Step 3: Integrated analysis;

[0144] Step 3-1: Preprocessing chemical data and multi-omics data of environmental endocrine disruptors;

[0145] Step 3-2: Perform cross-modal correlation modeling on chemical data and multi-omics data to achieve cross-modal correlation between chemical data and multi-omics data;

[0146] Step 4: Visualization;

[0147] Visualize cross-modal correlation results of chemical data and multi-omics data for information query and toxicity mechanism prediction of environmental endocrine disruptors.

[0148] Preferably, the data collection in step 1 is specifically as follows: using Django REST Framework to build a standardized API interface, supporting data uploading in JSON / CSV / XML formats, and integrating the Scrapy-Redis distributed crawler framework to crawl chemical data and multi-omics data of environmental endocrine disruptors from public databases;

[0149] The chemical data includes compound table, SMILES structural formula, inhibition constant, half maximal effect concentration and half inhibitory concentration;

[0150] The multi-omics data include transcriptome, metabolome, genome, proteome, epigenome, lipidome, spatial transcriptome, single-cell transcriptome and exposure record table.

[0151] Preferably, in step 3-1, the chemical data and multi-omics data of environmental endocrine disruptors are preprocessed, wherein the transcriptome preprocessing of the multi-omics data comprises the following steps:

[0152] Step 3-1-1: Quality control of raw data: use fastp tool to filter the collected transcriptome data to ensure that the Phred score is ≥20, and remove the adapter sequence;

[0153] Step 3-1-2: Sequence alignment: Use the STAR tool to map the read transcriptome data to the reference genome, setting the maximum mismatch rate to 10%;

[0154] Step 3-1-3: Expression quantification, using featureCounts statistical gene count matrix to perform gene expression analysis on transcriptome data;

[0155] Step 3-1-4: Differential analysis, standardization and hypothesis testing were performed using DESeq2 to perform differential analysis.

[0156] Preferably, in step 3-1, the chemical data and multi-omics data of environmental endocrine disruptors are preprocessed, wherein the metabolomics data preprocessing comprises the following steps:

[0157] Step 3-1-1: Extract the original mass spectrum peaks. Use the centWave algorithm of XCMS to detect the chromatographic peaks of the metabolomics data. Set the ppm tolerance to 10 and the peak width range to 5~20 seconds.

[0158] Step 3-1-2: Peak alignment and annotation: Use CAMERA to align and annotate chromatographic peaks and match them against the HMDB 5.0 database to identify metabolomics data features.

[0159] Step 3-1-3: Normalization.

[0160] Preferably, the step 3-2 includes the following steps:

[0161] Step 3-2-1: Chemical data encoding: Based on the RDKit chemical data science toolkit, the chemical data of environmental endocrine disruptors are converted into 2048-bit Morgan fingerprint feature vectors to capture the molecular structure characteristics of the chemical data. At the same time, molecular descriptors are calculated as supplementary feature vectors to obtain chemical data streams.

[0162] Step 3-2-2: Multi-omics data dimensionality reduction: Principal component analysis (PCA) was performed on the high-dimensional multi-omics data to reduce the dimensionality. The criterion of cumulative variance contribution rate ≥ 95% was used to determine the retention of 500 principal components. The biological consistency of the multi-omics data after dimensionality reduction was verified by t-SNE visualization to obtain the multi-omics data flow.

[0163] Step 3-2-3: Cross-modal alignment: Construct a two-stream deep neural network. The chemical data stream uses a three-layer fully connected network to process the fingerprint feature vector. The multi-omics data stream uses a 1D convolutional network to extract the features after principal component analysis dimensionality reduction. Finally, the cross-attention mechanism is used to learn the inter-modal association weights and output the low-dimensional feature vector in the joint embedding space to achieve cross-modal association of chemical data and multi-omics data.

[0164] Preferably, the dual-stream deep neural network consists of a chemical data stream, a multi-omics data stream and a cross-modal fusion mechanism. The chemical data stream adopts a multi-layer fully connected layer structure, including an input dimension of 1024 and a hidden layer of [512, 256]. Each layer is equipped with a ReLU activation function and a BatchNorm1d normalization layer, and a Dropout rate of 0.2 is used to provide regularization, and finally outputs a 128-dimensional feature vector; the multi-omics data stream uses a one-dimensional convolutional neural network to process sequence data, and designs a two-layer convolution structure with an input dimension of 200. The number of channels is [64, 128], using a 3×3 convolution kernel, combined with maximum pooling and adaptive average pooling layers, and finally obtains a 128-dimensional feature vector through linear projection; the feature vectors of the chemical data stream and the multi-omics data stream are cross-modally aligned through a multi-head attention mechanism, and then pass through a feature fusion layer. The feature fusion layer includes LayerNorm normalization, linear dimensionality reduction and Tanh activation function to generate a unified 128-dimensional feature vector.

[0165] Preferably, the steps of constructing the dual-stream deep neural network are as follows:

[0166] Step 3-2-3-1: Chemical data stream processing:

[0167] 1) Fully connected layer calculation:

[0168] H_l=ReLU(BN(W_l·H_{l-1}+b_l));

[0169] in:

[0170] H_l is the output of layer l;

[0171] H_{l-1} is the output of the l-1th layer;

[0172] W_l is the weight matrix;

[0173] b_l is the bias vector;

[0174] BN is batch normalization;

[0175] ReLU(x)=max(0,x);

[0176] 2) Dropout regularization:

[0177] H_dropout=H*mask;

[0178] mask~Bernoulli(p);

[0179] in:

[0180] H_dropout is the output feature vector, the feature after Dropout processing;

[0181] H is the input feature vector, which is the output activation value from the previous layer, that is, the output of the fully connected layer;

[0182] Mask is a binary mask matrix that randomly determines which neurons are retained or discarded;

[0183] p is the retention probability, which controls the regularization strength;

[0184] Bernoulli(p) is the Bernoulli distribution;

[0185] Step 3-2-3-2: Multi-omics data flow processing:

[0186] 1) One-dimensional convolution operation:

[0187] Conv1D(x)=σ(W*x+b);

[0188] in:

[0189] Conv1D(x) is a one-dimensional convolution;

[0190] σ is the ReLU activation function;

[0191] W is the convolution kernel weight;

[0192] * is the convolution operation;

[0193] b is the bias term;

[0194] 2) Max Pooling:

[0195] MaxPool(x)=max(x[i:i+k]);

[0196] in:

[0197] MaxPool(x) performs maximum pooling on the input data x and returns the maximum value in each pooling window;

[0198] k is the pooling window size;

[0199] i is the starting index of the current window;

[0200] x is the input data, usually a one-dimensional or two-dimensional array or tensor;

[0201] Step 3-2-3-3: Multi-head attention mechanism:

[0202] 1) Query Q, key K, value V calculation:

[0203] Q=X_q·W_q;

[0204] K=X_k·W_k;

[0205] V=X_v·W_v;

[0206] Attention weight calculation:

[0207] Attention(Q,K,V) = softmax(QK^T / √d_k)V;

[0208] in:

[0209] d_k is the attention dimension;

[0210] Softmax is a normalization function;

[0211] X_q is the input query Q representation, usually a matrix;

[0212] W_q is the weight matrix of the query;

[0213] X_k is the key K representation of the input, usually a matrix;

[0214] W_k is the bond weight matrix;

[0215] X_v is the input value V;

[0216] W_v is the weight matrix of values;

[0217] T is the transpose operation;

[0218] K^T is the transpose of the key matrix K;

[0219] 2) Multi-head attention:

[0220] MultiHead(Q,K,V)=Concat(head_1,...,head_h)W_o;

[0221] head_i=Attention(QW_i^Q, KW_i^K, VW_i^V);

[0222] in:

[0223] Concat(head_1, ..., head_h) concatenates h heads. The concatenated dimensions will be larger, representing the parallel computation results of multiple attention heads.

[0224] W_o is a linear transformation matrix used to transform the result of splicing multiple heads to obtain the final output;

[0225] head_i is the output of the i-th attention head;

[0226] Attention(QW_i^Q, KW_i^K, VW_i^V) is for the i-th attention head, using different weight matrices to linearly transform the query Q, key K and value V, and then calculate the standard attention;

[0227] W_i^Q represents the query Q weight matrix of the i-th head, which performs a linear transformation on the query Q and maps the original query into the new representation space;

[0228] W_i^K represents the key K weight matrix of the i-th head, which performs a linear transformation on the key K and maps the original key to the new representation space;

[0229] W_i^V represents the weight matrix of the value V of the i-th head, which performs a linear transformation on the value V and maps the original value to the new representation space;

[0230] Step 3-2-3-4: Feature fusion:

[0231] 1) Feature stitching:

[0232] H_fused=[H_chem+Attn_out;H_omics];

[0233] in:

[0234] H_fused is the new feature vector formed by splicing;

[0235] H_chem is the chemical-related feature vector;

[0236] Attn_out is the output from the attention mechanism;

[0237] H_omics is a feature vector generated based on other omics data;

[0238] + is the addition operation of H_chem and Attn_out;

[0239] ; It is splicing on the feature dimension, splicing H_chem + Attn_out and H_omics by column to form a new feature vector H_fused;

[0240] 2) LayerNorm normalization:

[0241] μ=(1 / H)ΣH_i;

[0242] σ =sqrt((1 / H)Σ(H_i-μ)^2);

[0243] LN(H)=γ*(H-μ) / σ+β;

[0244] H_i is the value of the i-th element in the input feature vector H;

[0245] μ is the arithmetic mean of all elements of the input eigenvector H;

[0246] σ is the standard deviation of the input eigenvector H, which measures the degree of discreteness of the data;

[0247] sqrt is the square root;

[0248] LN(H) is the output after normalizing the input feature vector H;

[0249] γ is a scaling parameter that controls the scale of the normalized output;

[0250] β is the offset parameter that adjusts the position of the normalized output.

[0251] Preferably, a multi-omics data integration and analysis system for environmental endocrine disruptors comprises a data access layer, a data storage layer, an integration and analysis layer, and a visualization layer.

[0252] like Figure 2 As shown, preferably, a multi-omics data integration and analysis system for environmental endocrine disruptors is used to implement the above-mentioned multi-omics data integration and analysis method for environmental endocrine disruptors, including a data access layer, a data storage layer, an integration and analysis layer, and a visualization layer.

[0253] The data access layer is used to implement standardized collection and upload of multi-source data. It uses a standardized API interface built on Django REST Framework and supports data upload in JSON / CSV / XML formats. It integrates the Scrapy-Redis distributed crawler framework for targeted crawling of environmental endocrine disruptor-related data from public databases. It also supports browser-side large file fragment upload based on HTML5 FileAPI, with a fragment size of 10MB.

[0254] The data storage layer uses MySQL relational database and Neo4j graph database to perform multimodal storage and management of chemical data and multi-omics data of environmental endocrine disruptors;

[0255] The integrated analysis layer is used to preprocess the chemical data and multi-omics data of environmental endocrine disruptors and perform cross-modal association modeling to achieve cross-modal association between chemical data and multi-omics data;

[0256] The visualization layer is used to visualize the cross-modal correlation results of chemical data and multi-omics data, and realize dynamic interactive visualization based on the D3.js framework.

[0257] The standardized API interfaces include the NCBI EUtils API and the ArrayExpress REST API. The public databases are the NCBI GEO database, ArrayExpress, and CTD. The NCBI EUtils API is used to retrieve RNA-seq and ChIP-seq datasets related to environmental endocrine disruptors from the NCBI GEO database. The ArrayExpress REST API supports batch downloading of metabolome and epigenomic data from the ArrayExpress database.

[0258] Example 1

[0259] This application system integrates the entire process of data collection, storage, analysis, and visualization. The specific architecture is as follows:

[0260] Data access layer: A multi-source data collection module was designed based on Python programming. This included using the DjangoREST Framework to build a standardized API interface (api / v1 / upload), which supports data upload in JSON / CSV / XML formats. The Scrapy-Redis distributed crawler framework was integrated to capture EDCs-related data from public databases such as PubMed and ChEMBL. Large file segment uploads (segment size: 10MB) were implemented in the browser based on the HTML5 File API.

[0261] The core function of the data entry layer of the data access system is to support the standardized access and integration of multi-source heterogeneous data. To address the specificities of environmental endocrine disruptor (EDC) research, this module integrates APIs from public bioinformatics databases (such as NCBI GEO, ArrayExpress, and CTD), enabling users to batch search and download multi-omics datasets related to EDC exposure (such as RNA-seq, ChIP-seq, and metabolomics mass spectrometry data) using keywords (such as compound name, species, and experimental type). The module also features a flexible data upload interface compatible with various formats such as JSON, CSV, and XML. It also includes a built-in data schema parser that automatically identifies user-uploaded omics metadata (such as sample number, experimental platform, and sequencing depth) and raw data files (such as FASTQ and mzML), converting them into a unified Parquet columnar storage format through an ETL (Extract-Transform-Load) process to ensure efficient subsequent analysis. To improve data quality, the module integrates a data validation rule library (such as the MIAME standard) to perform real-time verification of metadata integrity (e.g., missing exposure dose and time parameters) and prompts users to complete required fields through dynamic forms. Furthermore, the module supports integration with Laboratory Information Management Systems (LIMS) to achieve traceability of experimental samples throughout their lifecycle. All standardized data is stored in a normalized table in a MySQL database, keyed by the experiment ID. A data lineage graph is generated to document data sources and transformation paths, facilitating audit and reproducibility analysis.

[0262] Data storage layer: This embodiment uses a hybrid database architecture (MySQL + Neo4j) to achieve multimodal storage and management of environmental endocrine disruptors (EDCs) chemical data and multi-omics data.

[0263] For structured data, a normalized metadata management system was built based on MySQL 8.0. Core data tables were designed through entity-relationship modeling (ER modeling): ① The chemical table (Chemical) uses third normal form (3NF) to store basic attributes of EDCs, including fields such as chemical identifiers (PubChem CID, CAS number), physicochemical parameters (LogP, molecular weight), and toxicokinetic indicators (biological half-life, ADME properties). JSON-type fields are also used to expand the storage of dynamic experimental parameters. ② The omics table (Omics) uses a star schema to organize multidimensional omics data. Experiment number is used as the primary key to link subtables such as the transcriptome (FPKM matrix), epigenome (methylation site β-values), and metabolome (mass spectrometry peak intensity). Check constraints are used to ensure data type consistency. ③ The exposure record table (Exposure) implements spatiotemporal modeling, recording key metadata such as exposure dose (nM-μM gradient), time window (acute / chronic), and biological model (species, organ, cell line). Window functions are used to support time-series analysis of dose-effect relationships. This architecture ensures data consistency through foreign key constraints and transaction isolation level (READ COMMITTED), and combines virtual columns (Generated Columns) to achieve real-time calculation of common indicators (such as the number of differentially expressed genes).

[0264] For unstructured data, a dynamic knowledge network is constructed based on the Neo4j graph database: ① The EDCs-gene-disease associations extracted from the literature are converted into graph nodes (Node) and relationship edges (Relationship), and the Cypher query language is used to implement complex network topology analysis (such as the discovery of the shortest toxicity pathway); ② Multimodal data are integrated through an attribute graph model, RNA-seq expression profiles (JSON format) are embedded in gene nodes, and mass spectrometry imaging images (stored as cloud storage URLs) are associated with EDCs nodes; ③ Graph embedding algorithms (such as node2vec) are used to vectorize unstructured text features (such as experimental conclusion descriptions) and store them as node attributes to support similarity retrieval.

[0265] The two databases collaborate through an API gateway: the standardized exposed metadata in the MySQL relational database serves as the initial condition input for graph traversal, while the potential molecular interaction relationships inferred by Neo4j can be reversely written into the MySQL hypothesis table (Hypothesis) for verification.

[0266] The integrated analysis layer implements an asynchronous task engine and network topology analysis. Celery and Redis are used to implement a distributed task queue, handling time-consuming operations such as data cleaning and feature extraction. Compound-gene association networks are constructed based on the NetworkX library, and the adjacency matrix is ​​stored in a Redis Graph database. Furthermore, threaded scripts and an analysis environment are deployed for preprocessing and analyzing multi-omics data (genomics, transcriptomics, metabolomics, proteomics, epigenomics, spatial omics, and single-cell omics). By inputting commands through the front-end, various omics data can be preprocessed and analyzed.

[0267] Multi-omics data preprocessing utilizes a modular design, providing customized analysis pipelines for genomics, transcriptomics, metabolomics, and other omics data. For example, the transcriptomics data analysis workflow includes: ① Raw data quality control (FastQC + MultiQC); ② Sequence alignment (STAR / HISAT2); ③ Expression quantification (featureCounts / Salmon); and ④ Differential analysis, using DESeq2 for normalization and hypothesis testing. The metabolomics workflow encompasses: ① Raw mass spectrometry peak extraction (XCMS); ② Peak alignment and annotation (CAMERA); and ③ Normalization (PQN method). All workflow parameters (such as p-value thresholds for differential expression analysis and PCA dimensionality reduction) are configurable through a visual spreadsheet, and Docker containerization is supported for consistent environment execution. Preprocessing results (such as gene expression matrices and differentially methylated region lists) are stored in a structured format in a MySQL database. Interactive HTML reports are also generated, including data quality metrics (such as sequencing saturation curves and mass spectrometry total ion currents) and statistical summaries (such as the number of differentially expressed genes and metabolite pathway enrichment). The module also provides a "breakpoint resume" function. When a task is interrupted due to hardware failure, it can be resumed from the most recently completed step, significantly improving the robustness of large-scale data analysis.

[0268] Modeling cross-modal associations between chemical and multi-omics data involves four core stages:

[0269] (1) Chemical feature encoding: Based on the RDKit cheminformatics toolkit, the SMILES structural formula of EDCs was converted into a 2048-bit Morgan fingerprint feature vector (radius = 3) to capture molecular structural features (such as benzene rings and hydroxyl functional groups), and molecular descriptors (such as topological polar surface area and number of hydrogen bond donors) were calculated as supplementary feature vectors.

[0270] (2) Dimensionality reduction of omics features: Principal component analysis (PCA) was performed on high-dimensional omics data (e.g., 20,000 gene expression matrices). The criterion of cumulative variance contribution ≥ 95% was used to determine the retention of 500 principal components. The biological consistency of the data after dimensionality reduction was verified by t-SNE visualization (e.g., samples were clustered according to the exposure dose gradient).

[0271] (3) Cross-modal alignment: A dual-stream deep neural network (DNN) is designed. The chemical data stream uses a three-layer fully connected network to process the fingerprint vector, and the omics data stream uses a 1D convolutional network to extract the features after PCA dimensionality reduction. Finally, the cross-attention mechanism is used to learn the association weights between modalities and output a low-dimensional representation in the joint embedding space.

[0272] (4) Association network mining: A distributed graph computing engine was built based on Elasticsearch's Graph API. Subnetwork discovery was performed on the fused features. A modularity optimization algorithm (Louvain) was used to identify EDC-related functional modules (e.g., the "bisphenol A-estrogen receptor-cell proliferation pathway"). The results were persisted to the Neo4j graph database using the Cypher query language. This engine supports GPU acceleration and can complete pattern mining of millions of molecule-gene association pairs within 10 minutes.

[0273] Visualization layer: Implements dynamic interactive visualization of various graphs. Based on D3.js, force-directed graphs are implemented to display cross-modal correlation results of chemical and multi-omics data, supporting node dragging, zooming, and detailed information pop-up windows. Large-scale omics heatmaps are rendered using HTML5 Canvas + WebGL.

[0274] The core technologies of the visualization layer include:

[0275] ① Force-directed graph layout: The Barnes-Hut algorithm is used to optimize the rendering performance of large-scale networks (>10,000 nodes). Users can adjust the view in real time by dragging and zooming. Node colors map to compound types (EDCs are red, metabolites are blue), and edge widths reflect association strength (Pearson correlation coefficient).

[0276] ② Dynamic data binding: A real-time connection to the Neo4j database is achieved through WebSocket. After the user selects the exposure conditions (such as "concentration>10μM&time>24h"), the system automatically generates a Cypher query (such as MATCH (c:Chemical)-[r:AFFECTS]->(g:Gene) WHERE r.dose>10 RETURN c,g,r) and renders the subnet on the front end.

[0277] ③ Multi-level Interaction: Clicking a node pops up a floating window displaying detailed information (such as the gene's GO annotation and PubChem links to EDCs). The right-click menu supports path highlighting (e.g., displaying the shortest pathway from a compound to a disease phenotype), module export (SVG / PNG format), and data drill-down (e.g., viewing differential expression heatmaps of associated genes). Furthermore, the module integrates WebGL acceleration technology, enabling smooth presentation of cell-gene co-occurrence networks (>100,000 nodes) for single-cell transcriptomes.

[0278] Example 2

[0279] This implementation combines high-performance computing hardware with a modular software architecture to efficiently process and analyze multi-omics data on environmental endocrine disruptors (EDCs). The core hardware is a Dell PowerEdge R750xa server equipped with dual-socket Intel Xeon Gold 6338 processors (32 cores / 64 threads per processor, clocked at 2.2GHz) and 256GB of DDR4 memory to support large-scale concurrent computing tasks. Four NVIDIA A100 GPUs (80GB of video memory each) are used as graphics processing units to accelerate deep learning model training and graph computing tasks. Data storage utilizes a tiered strategy: a 10TB NVMe solid-state drive for high-speed read and write access to hot data (such as real-time analysis of intermediate results), while a 50TB mechanical hard drive is used for long-term archiving of cold data (such as raw sequencing data). Network communication is based on a Gigabit Ethernet architecture, with firewall rules granting access to MySQL (port 3306), Neo4j (port 7474), and Elasticsearch (port 9200) to ensure secure data exchange.

[0280] The software environment was deployed on the Ubuntu 22.04 LTS operating system. The database systems used were MySQL 8.0.32 (with the InnoDB storage engine and the transaction isolation level set to READ COMMITTED) and Neo4j 5.11.0 (with the APOC plugin 4.4.0.15 integrated to support advanced graph algorithms). Python 3.10.12 was the primary programming language, relying on RDKit 2023.03.1 for chemical feature calculations, and PyTorch 2.0.1 for building deep learning models. Front-end visualization was based on JavaScript ES6, using D3.js 7.8.2 for dynamic network rendering and Cytoscape.js 3.23.0 to support complex interactions. The system containerized the analysis workflow using Docker 24.0.5, pre-installing the Snakemake 7.32.4 workflow engine and R 4.3.1 (with the Bioconductor 3.17 bioinformatics toolkit) to ensure reproducibility and cross-platform compatibility.

[0281] Step 1: Data collection;

[0282] Data collection relies on a data access layer, which is divided into public data access and user-defined data upload. Public data access is achieved through standardized APIs. The NCBI EUtils API (version 2.0) is used to retrieve EDC-related RNA-seq and ChIP-seq datasets from the GEO database, while the ArrayExpress REST API (version 1.7) supports batch download of metabolome and epigenomic data. Users can enter the compound name (e.g., "Bisphenol A"), species ID (e.g., human Taxonomy ID 9606), and experiment type. The system automatically parses and downloads the corresponding metadata (e.g., GSE number, sample description) and raw data files (e.g., FASTQ and mzML). To improve data transfer efficiency, FASTQ files are downloaded using the aspera command-line tool (version 3.11.1), with a speed limit of 500 Mbps.

[0283] User-defined data uploads are supported in multiple formats, including CSV, JSON, and XML. CSV files must be UTF-8 encoded, with column names (e.g., sample_id, dose) in the first row. Numeric fields must conform to predefined ranges (e.g., exposure dose in μM). JSON files must conform to the predefined Schema specification (draft-07) and have field integrity validation performed using the jsonschema library (version 4.18.0). Uploaded data is converted to the Parquet columnar storage format (Apache Parquet 2.9.0) through an ETL process, with the experiment ID as the partition key (e.g., GSE12345_SAMN678901). SNAPPY compression is used for optimal storage efficiency. Data lineage management is implemented using the PROV-O model (W3C standard), which records data source, conversion path, and version information. Data is stored in the data_provenance table in MySQL to ensure traceability.

[0284] Step 2: Data storage;

[0285] A hybrid database architecture is used to achieve multimodal storage and management of chemical data and multi-omics data of environmental endocrine disruptors. The chemical data and multi-omics data of environmental endocrine disruptors are divided into structured data and unstructured data.

[0286] Step 3: Integrated analysis;

[0287] Step 3-1: Preprocessing chemical data and multi-omics data of environmental endocrine disruptors;

[0288] Taking transcriptome analysis as an example, the process is built based on the Snakemake workflow engine and mainly includes several key steps.

[0289] First, we performed quality control on the raw data. We used the fastp tool (version 0.23.2) to filter low-quality reads, ensuring a Phred score ≥ 20, and excised adapter sequences to improve data quality. Next, we performed sequence alignment using the STAR tool (version 2.7.10b) to map the reads to the GENCODE human release 43 reference genome, with a maximum mismatch rate of 10% to ensure accurate alignment.

[0290] During expression quantification, featureCounts (version 2.0.3) was used to generate a gene count matrix, using the annotation file gencode.v43.annotation.gtf to provide reliable gene expression data. Subsequently, differential analysis was performed using DESeq2 (version 1.38.3) for normalization and hypothesis testing.

[0291] The metabolome analysis workflow began with mass spectrometry peak extraction. Chromatographic peaks were detected using the centWave algorithm in XCMS (version 3.20.0), with a ppm tolerance of 10 and a peak width range of 5–20 seconds. Next, peak alignment and annotation were performed using CAMERA (version 1.56.0) against the HMDB 5.0 database to identify metabolite signatures.

[0292] All analysis processes were encapsulated in a Docker image (biocontainers / snakemake:7.32.4). In terms of resource constraints, the single-task memory was set to 64 GB, the CPU cores were set to 16, and GPU acceleration was supported (--gpus all).

[0293] The preprocessing results are stored in a structured format in a MySQL database, and an interactive HTML report is generated. The report includes a sequencing saturation curve, a mass spectrometry total ion current, and a list of differentially expressed genes (FDR < 0.05), providing comprehensive support for subsequent data analysis.

[0294] Step 3-2: The multimodal data fusion engine module performs cross-modal correlation modeling on chemical data and multi-omics data to achieve cross-modal correlation between chemical data and multi-omics data;

[0295] During the chemical data encoding phase, the GetMorganFingerprintAsBitVect function in RDKit was used to generate a 2048-bit Morgan fingerprint (radius 3) to capture molecular structural features such as benzene rings and hydroxyl groups. Descriptors such as LogP and the number of hydrogen bond donors were also calculated and stored in JSON format. During the omics feature dimensionality reduction phase, principal component analysis (PCA, scikit-learn version 1.2.2) was used to reduce the gene expression matrix to 500 dimensions, ensuring a cumulative variance contribution of 95% or greater. t-SNE was then used to verify the biological clustering of the reduced data.

[0296] In the cross-modal alignment phase, a two-stream deep neural network architecture was designed, consisting of a chemical data stream, a multi-omics data stream, and a cross-modal fusion mechanism. The chemical data stream uses a multi-layer fully connected network structure with an input dimension of 1024 and a hierarchy of hidden layers of [512, 256]. Each layer is equipped with a ReLU activation function and a BatchNorm1d normalization layer, and a dropout rate of 0.2 is used for regularization, ultimately outputting a 128-dimensional feature vector. The omics data stream uses a one-dimensional convolutional neural network to process sequence data. A two-layer convolutional structure with an input dimension of 200 and a channel number of [64, 128] is designed. A 3×3 convolution kernel is used, combined with maximum pooling and adaptive average pooling layers, to obtain a 128-dimensional feature vector through linear projection. The features of the two data streams are cross-modally aligned using a multi-head attention mechanism (four attention heads, a 128-dimensional embedding space, and a dropout rate of 0.1). They then pass through a feature fusion layer that includes LayerNorm normalization, linear dimensionality reduction, and the Tanh activation function to generate a unified 128-dimensional feature vector. This architectural design not only processes the specific features of the two modalities separately but also adaptively learns intermodal associations through the attention mechanism. Multi-level regularization ensures model stability and generalization, supporting both supervised and unsupervised training modes, and demonstrating excellent feature extraction and modality fusion performance on sample data.

[0297] During the association network mining phase, we built an Elasticsearch cluster (consisting of three nodes and ten shards) to implement a distributed graph engine. We used the Louvain algorithm (with a resolution of 1.0) to identify functional modules. The identification results were persisted to a Neo4j database for subsequent analysis.

[0298] Step 4: Visualization;

[0299] The visualization module, developed based on the D3.js framework, is designed to support dynamic network construction and interactive exploration. This module employs a force-directed layout and the Barnes-Hut approximation algorithm (with parameter θ = 0.8) to optimize node interactions. In the visualization, node color mapping indicates compound type, with environmental hormone disruptors (EDCs) highlighted in red. Edge width reflects the strength of association, quantified by calculating the Pearson correlation coefficient.

[0300] Users can connect to the Neo4j database in real time via WebSocket and enter a Cypher query (for example, to screen for associations between EDCs at doses greater than 10 μM and genes). The front-end interface automatically renders the corresponding subnetwork and supports dragging and zooming nodes. In the detailed information window, users can view the GO annotation of the gene and the PubChem link of the compound. In addition, the right-click menu provides pathway highlighting (such as the shortest toxicity pathway), module export options (supporting SVG / PNG formats), and data drill-down capabilities for deeper analysis.

[0301] To improve rendering performance, the module can handle networks with up to 100,000 nodes and utilizes WebGL acceleration technology from Three.js (version r152) to ensure a smooth user experience with interactive frame rates exceeding 30 FPS. This design makes large-scale data visualization more efficient and intuitive.

[0302] like Figure 3 , which is a schematic diagram of the online function of the multi-omics data integration analysis method for environmental endocrine disruptors in this application, Figure 3 Figure a is the function item for selecting tasks to perform, including chemical data and multi-omics data preprocessing; b is the interface for performing differential gene expression analysis tasks; c is the option for selecting import data; and d is the interface for displaying task execution results.

[0303] like Figure 4 As shown, an analysis platform designed using the multi-omics data integration analysis method of environmental endocrine disruptors in this application is shown. The figure shows the setting of the visualization layer, which visualizes the results of the integrated analysis layer.

[0304] like Figure 5 As shown, Figure 4 Functional module execution interface of Transcriptional Factors.

[0305] This application discloses a multi-omics data integration and analysis system for environmental endocrine disruptors. Through data collection, data storage, integrated analysis and visualization, a dedicated metadata framework for environmental endocrine disruptors is constructed, and a full-dimensional collection system covering chemical data and multi-omics data is established. This application supports automatic conversion and structured storage of multiple data formats, thereby significantly improving data traceability efficiency and shortening the data preprocessing cycle.

[0306] The dual-stream deep neural network disclosed in this application can not only process the specific features of the two modal data separately, but also adaptively learn the association between the modalities through the attention mechanism. At the same time, it ensures the stability and generalization ability of the model through multi-level regularization measures, supports both supervised and unsupervised training modes, and demonstrates good feature extraction and modal fusion performance on example data.

[0307] This application integrates heterogeneous network data from different experimental conditions, species or omics levels through multi-layer network modeling, realizes multimodal data integration analysis and visualization functions, and promotes the prediction and query analysis of EDCs toxicity mechanism.

[0308] The preferred embodiments of the present application have been described in detail above, but the present application is not limited to the above embodiments. Various changes can be made within the scope of knowledge possessed by ordinary technicians in this field without departing from the purpose of the present application.

[0309] Many other changes and modifications can be made without departing from the concept and scope of the present application. It should be understood that the present application is not limited to the specific embodiments, and the scope of the present application is defined by the appended claims.

Claims

1. A method for integrated analysis of multi-omics data of environmental endocrine disruptors, characterized in that: The following steps are involved: Step 1: Data collection: extract chemical data and multi-omics data of environmental endocrine disruptors from public databases; Step 2: Data storage; A hybrid database architecture is used to achieve multimodal storage and management of chemical data and multi-omics data of environmental endocrine disruptors. The chemical data and multi-omics data of environmental endocrine disruptors are divided into structured data and unstructured data. For structured data, a normalized metadata management system is built based on the MySQL relational database. Core data tables are designed through entity-relationship modeling to store structured data. For unstructured data, a dynamic knowledge network is built based on the Neo4j graph database to store unstructured data; Step 3: Integrated analysis; Step 3-1: Preprocessing chemical data and multi-omics data of environmental endocrine disruptors; Step 3-2: Perform cross-modal correlation modeling on chemical data and multi-omics data to achieve cross-modal correlation between chemical data and multi-omics data; Step 4: Visualization; Visualize cross-modal correlation results of chemical data and multi-omics data for information query and toxicity mechanism prediction of environmental endocrine disruptors.

2. The method for integrated analysis of multi-omics data of environmental endocrine disruptors according to claim 1, characterized in that: The data collection in step 1 is specifically as follows: using Django REST Framework to build a standardized API interface, supporting data upload in JSON / CSV / XML formats, and integrating the Scrapy-Redis distributed crawler framework to crawl chemical data and multi-omics data of environmental endocrine disruptors from public databases; The chemical data includes compound table, SMILES structural formula, inhibition constant, half maximal effect concentration and half inhibitory concentration; The multi-omics data include transcriptome, metabolome, genome, proteome, epigenome, lipidome, spatial transcriptome, single-cell transcriptome and exposure record table.

3. The method for integrated analysis of multi-omics data of environmental endocrine disruptors according to claim 2, characterized in that: In step 3-1, the chemical data and multi-omics data of environmental endocrine disruptors are preprocessed, wherein the transcriptome preprocessing of the multi-omics data includes the following steps: Step 3-1-1: Quality control of raw data: use fastp tool to filter the collected transcriptome data to ensure that the Phred score is ≥20, and remove the adapter sequence; Step 3-1-2: Sequence alignment: Use the STAR tool to map the read transcriptome data to the reference genome, setting the maximum mismatch rate to 10%; Step 3-1-3: Expression quantification, using featureCounts statistical gene count matrix to perform gene expression analysis on transcriptome data; Step 3-1-4: Differential analysis, standardization and hypothesis testing were performed using DESeq2 to perform differential analysis.

4. The method for integrated analysis of multi-omics data of environmental endocrine disruptors according to claim 2, characterized in that: In step 3-1, the chemical data and multi-omics data of environmental endocrine disruptors are preprocessed, wherein the metabolomics data preprocessing includes the following steps: Step 3-1-1: Extract the original mass spectrum peaks. Use the centWave algorithm of XCMS to detect the chromatographic peaks of the metabolomics data. Set the ppm tolerance to 10 and the peak width range to 5~20 seconds. Step 3-1-2: Peak alignment and annotation: Use CAMERA to align and annotate chromatographic peaks and match them against the HMDB 5.0 database to identify metabolomics data features. Step 3-1-3: Normalization.

5. The method for integrated analysis of multi-omics data of environmental endocrine disruptors according to claim 1, characterized in that: The step 3-2 includes the following steps: Step 3-2-1: Chemical data encoding: Based on the RDKit chemical data science toolkit, the chemical data of environmental endocrine disruptors are converted into 2048-bit Morgan fingerprint feature vectors to capture the molecular structure characteristics of the chemical data. At the same time, molecular descriptors are calculated as supplementary feature vectors to obtain chemical data streams. Step 3-2-2: Multi-omics data dimensionality reduction: Principal component analysis (PCA) was performed on the high-dimensional multi-omics data to reduce the dimensionality. The criterion of cumulative variance contribution rate ≥ 95% was used to determine the retention of 500 principal components. The biological consistency of the multi-omics data after dimensionality reduction was verified by t-SNE visualization to obtain the multi-omics data flow. Step 3-2-3: Cross-modal alignment: Construct a two-stream deep neural network. The chemical data stream uses a three-layer fully connected network to process the fingerprint feature vector. The multi-omics data stream uses a 1D convolutional network to extract the features after principal component analysis dimensionality reduction. Finally, the cross-attention mechanism is used to learn the inter-modal association weights and output the low-dimensional feature vector in the joint embedding space to achieve cross-modal association of chemical data and multi-omics data.

6. The method for integrated analysis of multi-omics data of environmental endocrine disruptors according to claim 5, characterized in that: The dual-stream deep neural network consists of a chemical data stream, a multi-omics data stream and a cross-modal fusion mechanism. The chemical data stream adopts a multi-layer fully connected layer structure, including an input dimension of 1024 and a hidden layer hierarchy of [512, 256]. Each layer is equipped with a ReLU activation function and a BatchNorm1d normalization layer, and a Dropout rate of 0.2 is used to provide regularization, and finally outputs a 128-dimensional feature vector; the multi-omics data stream uses a one-dimensional convolutional neural network to process sequence data. A two-layer convolution structure with an input dimension of 200 is designed, and the number of channels is [64, 128], using a 3×3 convolution kernel, combined with maximum pooling and adaptive average pooling layers, and finally a 128-dimensional feature vector is obtained through linear projection; the feature vectors of the chemical data stream and the multi-omics data stream are cross-modally aligned through a multi-head attention mechanism, and then pass through a feature fusion layer, which includes LayerNorm normalization, linear dimensionality reduction and Tanh activation function to generate a unified 128-dimensional feature vector.

7. The method for integrated analysis of multi-omics data of environmental endocrine disruptors according to claim 6, characterized in that: The steps for constructing the two-stream deep neural network are as follows: Step 3-2-3-1: Chemical data stream processing: 1) Fully connected layer calculation: H_l=ReLU(BN(W_l·H_{l-1}+b_l)); in: H_l is the output of layer l; H_{l-1} is the output of the l-1th layer; W_l is the weight matrix; b_l is the bias vector; BN is batch normalization; ReLU(x)=max(0,x); 2) Dropout regularization: H_dropout=H*mask; mask~Bernoulli(p); in: H_dropout is the output feature vector, the feature after Dropout processing; H is the input feature vector, which is the output activation value from the previous layer, that is, the output of the fully connected layer; mask is a binary mask matrix; p is the retention probability, which controls the regularization strength; Bernoulli(p) is the Bernoulli distribution; Step 3-2-3-2: Multi-omics data flow processing: 1) One-dimensional convolution operation: Conv1D(x)=σ(W*x+b); in: Conv1D(x) is a one-dimensional convolution; σ is the ReLU activation function; W is the convolution kernel weight; * is the convolution operation; b is the bias term; 2) Max Pooling: MaxPool(x)=max(x[i:i+k]); in: MaxPool(x) performs maximum pooling on the input data x and returns the maximum value in each pooling window; k is the pooling window size; i is the starting index of the current window; x is the input data; Step 3-2-3-3: Multi-head attention mechanism: 1) Query Q, key K, value V calculation: Q=X_q·W_q; K=X_k·W_k; V=X_v·W_v; Attention weight calculation: Attention(Q,K,V) = softmax( )V; in: d_k is the attention dimension; Softmax is a normalization function; X_q is the input query Q representation; W_q is the weight matrix of the query; X_k is the key K representation of the input; W_k is the bond weight matrix; X_v is the input value V; W_v is the weight matrix of values; T is the transpose operation; K^T is the transpose of the key matrix K; 2) Multi-head attention: MultiHead(Q,K,V)=Concat(head_1,...,head_h)W_o; head_i=Attention(QW_i^Q, KW_i^K, VW_i^V); in: Concat(head_1, ..., head_h) concatenates h heads. The concatenated dimensions will be larger, representing the parallel computation results of multiple attention heads. W_o is a linear transformation matrix used to transform the result of splicing multiple heads to obtain the final output; head_i is the output of the i-th attention head; Attention(QW_i^Q, KW_i^K, VW_i^V) is for the i-th attention head, using different weight matrices to linearly transform the query Q, key K and value V, and then calculate the standard attention; W_i^Q represents the query Q weight matrix of the i-th head, which performs a linear transformation on the query Q and maps the original query into the new representation space; W_i^K represents the key K weight matrix of the i-th head, which performs a linear transformation on the key K and maps the original key to the new representation space; W_i^V represents the weight matrix of the value V of the i-th head, which performs a linear transformation on the value V and maps the original value to the new representation space; Step 3-2-3-4: Feature fusion: 1) Feature stitching: H_fused=[H_chem+Attn_out;H_omics]; in: H_fused is the new feature vector formed by splicing; H_chem is the chemical-related feature vector; Attn_out is the output from the attention mechanism; H_omics is a feature vector generated based on other omics data; + is the addition operation of H_chem and Attn_out; ; It is splicing on the feature dimension, splicing H_chem + Attn_out and H_omics by column to form a new feature vector H_fused; 2) LayerNorm normalization: μ=(1 / H)ΣH_i; σ =sqrt((1 / H)Σ(H_i-μ)^2); LN(H)=γ*(H-μ) / σ+β; H_i is the value of the i-th element in the input feature vector H; μ is the arithmetic mean of all elements of the input eigenvector H; σ is the standard deviation of the input eigenvector H, which measures the degree of discreteness of the data; sqrt is the square root; LN(H) is the output after normalizing the input feature vector H; γ is a scaling parameter that controls the scale of the normalized output; β is the offset parameter that adjusts the position of the normalized output.

8. A multi-omics data integration and analysis system for environmental endocrine disruptors, characterized by: Used for the implementation of a multi-omics data integration and analysis method for environmental endocrine disruptors according to any one of claims 1 to 7, comprising a data access layer, a data storage layer, an integration and analysis layer, and a visualization layer, The data access layer is used to implement standardized collection and upload of multi-source data. It uses a standardized API interface built on Django REST Framework and supports data upload in JSON / CSV / XML formats. It integrates the Scrapy-Redis distributed crawler framework for targeted crawling of environmental endocrine disruptor-related data from public databases. It also supports browser-side large file fragment upload based on HTML5 FileAPI, with a fragment size of 10MB. The data storage layer uses MySQL relational database and Neo4j graph database to perform multimodal storage and management of chemical data and multi-omics data of environmental endocrine disruptors; The integrated analysis layer is used to preprocess the chemical data and multi-omics data of environmental endocrine disruptors and perform cross-modal association modeling to achieve cross-modal association between chemical data and multi-omics data; The visualization layer is used to visualize the cross-modal association results of chemical data and multi-omics data, and realizes dynamic interactive visualization based on the D3.js framework for information query and toxicity mechanism prediction of environmental endocrine disruptors.

Citation Information

Patent Citations

  • Single-cell multi-omics data cross-modal generation method and system

    CN118173173A

  • Single cell cross-omics multi-modal prediction method

    CN119889434A