Rodent-oriented biological information data processing platform
By designing a rodent-oriented bioinformatics data processing platform, integrating multiple types of bioinformatics data and adopting advanced analysis technology, the problem of low data integration and analysis efficiency in the existing technology is solved, and efficient and accurate bioinformatics analysis is achieved, providing strong support for biomedical research.
Patent Information
- Application Number
- CN202510644836.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The prior art is difficult to effectively integrate and analyze multiple types of rodent biological information data, and there is a lack of data processing flow and algorithms optimized for rodent characteristics, resulting in low accuracy and efficiency of analysis results.
A biological information data processing platform for rodents is designed, including data acquisition and integration module, gene analysis module, protein analysis module, metabolomics analysis module, association mining module and user interaction module. Data integration is carried out through unified data format standards, and analysis is carried out using machine learning, deep learning and mass spectrometry data analysis and other technologies.
It realizes efficient integration and accurate analysis of various types of rodent bioinformatics data, improves the efficiency of data processing, and provides a reliable basis for biomedical research.
Smart Images

Figure CN120183508A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics data processing systems, and particularly to a bioinformatics data processing platform for rodents. Background Art
[0002] In biomedical research, rodents (such as mice, rats, etc.) are widely used in various experiments due to their physiological characteristics being somewhat similar to those of humans. By studying the biological information such as genes and proteins of rodents, the pathogenesis of diseases, the action targets of drugs, etc. can be deeply understood. However, currently, for the processing of rodent bioinformatics data, there are the following problems: Difficult data integration: The rodent biological data generated by different research institutions has diverse formats and is stored dispersedly, making it difficult to perform effective integration and analysis. Single analysis tool: Existing bioinformatics analysis tools often focus on the analysis of a certain specific type of data, such as gene sequence analysis or protein structure prediction, lacking a comprehensive analysis platform for all aspects of rodent bioinformation. Lack of pertinence: There is no data processing process and algorithm specifically optimized for the characteristics of rodents, resulting in lower accuracy and efficiency of analysis results. Summary of the Invention
[0003] The technical problem to be solved by the present invention is how to provide a bioinformatics data processing platform for rodents that can integrate various types of rodent biological information, is efficient and has accurate analysis.
[0004] To solve the above technical problem, the technical solution adopted by the present invention is: A bioinformatics data processing platform for rodents, including: Data acquisition and integration module: Used to collect rodent biological information data from multiple data sources, including gene sequence data, protein structure data, and metabolomics data; by formulating a unified data format standard, convert and integrate data in different formats, and store them in the database module; Gene analysis module: Used to implement gene sequence alignment using gene sequence alignment algorithms, and analyze the expression patterns of genes using machine learning algorithms to predict the functions of genes in different physiological states; Protein analysis module: Through a protein structure prediction model based on deep learning, predict the three-dimensional structure of a protein according to its amino acid sequence; use a protein-protein interaction network analysis algorithm to construct a rodent biological protein-protein interaction network and obtain the functional relationships between proteins; Metabolomics analysis module: Adopt a mass spectrometry data parsing algorithm, combined with database matching, to identify the types of metabolites; through metabolic pathway enrichment analysis, determine the metabolic pathways related to specific physiological or pathological states; Association relationship mining module: According to the results of gene analysis, protein analysis, and metabolomics analysis, mine the association relationships between genes, proteins, and metabolomics and this type of rodent, and visually display these association relationships; User interaction module: Used to provide a friendly user interface through which users can set permission management functions, upload data, select analysis functions, and view analysis results; Database module, used to store the data received and generated by the platform.
[0005] The beneficial effects of adopting the above technical solutions are as follows: By integrating various types of rodent bioinformatics data, the platform avoids users from switching between different data sources and analysis tools, greatly improving the efficiency of data processing. Using algorithms and models optimized for the characteristics of rodents can more accurately analyze and interpret rodent bioinformatics data, providing a reliable basis for biomedical research. Brief description of the drawings
[0006] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments.
[0007] Figure 1 is the principle block diagram of the platform described in the embodiment of the present invention; Figure 2 is the implementation flowchart of the gene analysis module in the platform described in the embodiment of the present invention; Figure 3 is the implementation flowchart of the protein analysis module in the platform described in the embodiment of the present invention; Figure 4 is the implementation flowchart of the metabolomics analysis module in the platform described in the embodiment of the present invention; Figure 5 is the implementation flowchart of the association relationship mining module in the platform described in the embodiment of the present invention. Detailed implementation manners
[0008] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0009] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0010] As shown Figure 1 in the figure, an embodiment of the present invention discloses a biological information data processing platform for rodents, including: Data acquisition and integration module 101: used to collect rodent biological information data from multiple data sources, including gene sequence data, protein structure data, and metabolomics data; by formulating a unified data format standard, convert and integrate data in different formats, and store them in the database module; It should be noted that the "unified data format standard" in this application is a prior art, which refers to a standardized file format, field definition, encoding rule, and metadata specification formulated for different types of data (such as gene sequences, protein structures, metabolomics data, etc.) in order to achieve cross-platform and cross-data source data compatibility and interoperability. The specific implementation steps are not elaborated here.
[0011] In addition, in this application, the method of converting and integrating data in different formats is a prior art, and its implementation method generally includes the following steps: data parsing and cleaning, format conversion and standard mapping, data integration and deduplication, and quality control and storage, etc., to convert heterogeneous data into a unified format that can be directly used by each module of the platform (such as gene analysis, association mining). The specific implementation steps are not elaborated here.
[0012] Gene analysis module 102: used to implement gene sequence alignment using a gene sequence alignment algorithm, and analyze the gene expression pattern using a machine learning algorithm to predict the role of genes in different physiological states; It should be noted that the method of the step "and analyze the gene expression pattern using a machine learning algorithm" in this application is a prior art, and its implementation method generally includes the following steps: data preparation and preprocessing, feature selection, machine learning model selection and training, model optimization, and using the optimized model to analyze the gene expression pattern. The specific implementation steps are not elaborated here.
[0013] Protein analysis module 103: through a protein structure prediction model based on deep learning, predict the three-dimensional structure of a protein according to its amino acid sequence; use a protein-protein interaction network analysis algorithm to construct a rodent biological protein-protein interaction network and obtain the functional relationship between proteins; Metabolomics analysis module 104: adopt a mass spectrometry data parsing algorithm, combined with database matching, to identify the types of metabolites; determine the metabolic pathways related to specific physiological or pathological states through metabolic pathway enrichment analysis; Association relationship mining module 105: According to the results of gene analysis, protein analysis, and metabolomics analysis, mine the association relationships between genes, proteins, and metabolomics and this type of rodent, and visually display these association relationships; User interaction module 106: Used to provide a friendly user interface through which users can set permission management functions, upload data, select analysis functions, and view analysis results; Database module 107, used to store the data received and generated by the platform.
[0014] Furthermore, as Figure 2 shown, the processing method of the gene analysis module 101 includes the following steps: 1) Gene sequence alignment: Obtain the gene sequence data to be aligned from the database module, including the query sequence (usually the new sequence input by the user) and the reference sequence (the known rodent gene sequence stored in the platform database), establish an index for the reference sequence for quick search of matching fragments. For large-scale gene sequence databases, efficient index structures such as Suffix Tree or Burrows-Wheeler Transform (BWT) can be used to accelerate the search process; divide the query sequence into seed fragments, and according to the sequence characteristics of rodent genes, adjust the length and position of the seeds, search for matching seeds in the index, perform bidirectional extension alignment on the found seed matches, calculate the alignment score using a scoring matrix for rodent genes, determine whether it is a valid alignment according to the set threshold, and use greedy algorithm or dynamic programming algorithm to merge and optimize the matching fragments to obtain the final sequence alignment result; the alignment result is usually presented in the form of a table or text, including information on the query sequence and the reference sequence, alignment positions, alignment scores, matching percentages, etc.
[0015] 2) Gene expression pattern analysis: Feature selection and extraction: Extract features from gene expression data for the training and prediction of machine learning models. The features can include the expression level of genes, the fold change of expression, and differential expression information between different samples, etc.; Construction and training of machine learning models: Use Support Vector Machine (SVM) as the machine learning model, divide the labeled feature dataset (the expression features of known genes in different physiological states and the corresponding physiological state labels) into a training set and a test set, and use the training set and the test set to train the machine learning model. By optimizing the kernel function (such as linear kernel, radial basis kernel, etc.) and adjusting hyperparameters (such as penalty parameter C and kernel parameter γ), make the model achieve optimal performance on the training set; Prediction and Evaluation: Input the gene expression features to be predicted into the trained machine learning model to obtain the prediction results of the roles of genes under different physiological states, usually represented in the form of probabilities or categories. Evaluate the performance of the model, and use the test set to calculate metrics such as accuracy, recall rate, and F1 value to evaluate the reliability of the model.
[0016] Furthermore, as Figure 3 shown, the processing method of the protein analysis module includes the following steps: 1) Protein Structure Prediction Protein Structure Model Construction: Use the AlphaFold architecture as the basic architecture for protein structure prediction, including an input layer, an intermediate layer, and an output layer; the input layer uses one-hot encoding to represent each amino acid sequence of the protein as a high-dimensional vector. For example, for 20 common amino acids, an amino acid can be represented as a vector of length 20, where the position corresponding to the amino acid is 1 and the rest are 0; the intermediate layer includes multiple Transformer modules for learning long-range dependencies in the amino acid sequence. The Transformer module contains a multi-head self-attention mechanism, and the multi-head self-attention mechanism allows the model to simultaneously focus on information at different positions, enhancing the ability to capture sequence information; the output layer is used to output the predicted protein structure information, usually atomic coordinates or certain feature representations of the structure, such as distance matrices, angles, etc. Finally, through geometric calculations, this information is converted into the three-dimensional structure of the protein.
[0017] Training Data Collection and Preprocessing: Collect protein data of rodent organisms with known structures, clean the amino acid sequences to remove incomplete or incorrect sequence information; for the protein structure, convert it into the input representation required by the model, and increase the diversity of the data through rotation, translation, and adding noise operations to improve the generalization ability of the model; Model Training: Divide the preprocessed data into a training set, a validation set, and a test set, use the training set to train the model, adjust the model hyperparameters according to the performance of the validation set, and use the test set to evaluate the performance of the final model; Structure Prediction: For the amino acid sequences of new rodent organism proteins, encode them into input vectors, input them into the trained model, the model outputs the predicted protein structure features, and then through a structure reconstruction algorithm, these features are converted into the three-dimensional structure of the protein; 2) Protein Interaction Network Analysis Network Construction: Represent proteins as nodes and interactions as edges to construct an undirected graph , where is the set of protein nodes, is the set of edges, and different weights are assigned to the edges according to the strength of the interactions; Network analysis algorithm: Using the Markov Clustering Algorithm (MCL), nodes in an undirected graph are divided into different clusters. Expansion and inflation operations are performed on the adjacency matrix to enhance the connection pattern between nodes. By iteratively updating the matrix, it finally converges to a stable state to obtain different clusters. Functional relationship mining: Based on the clustering results and combined with the functional annotation information of proteins, hypergeometric distribution is used for functional enrichment analysis of each cluster, and the functional relationships between proteins are obtained according to the analysis results. Through the above steps and algorithms, the protein analysis module can achieve the prediction of rodent protein structures and the analysis of functional relationships between proteins, providing important support for in-depth understanding of protein functions and biological processes. Furthermore, as Figure 4 shown, the implementation method of the metabolomics analysis module includes the following steps: 1) Identification of metabolite species Mass spectrometry data acquisition: Through liquid chromatography - mass spectrometry (LC - MS) coupling, metabolomics mass spectrometry data of rodent biological samples are obtained. The metabolomics mass spectrometry data include the mass - to - charge ratio and corresponding signal intensity of metabolites. Data pre - processing: The centroid method is used to identify the peaks of metabolite ion peaks from the original mass spectrometry data. The centroid method finds the local maximum of the signal intensity as the center of the peak, and its formula is: ; where is the mass - to - charge ratio of the i - th metabolite, is the signal intensity of the i - th metabolite, and are the scanning ranges where the peak is located; The dynamic time warping (DTW) algorithm is used to align the obtained peaks. The Savitzky - Golay filtering algorithm is used to remove noise to improve data quality. The total ion current normalization algorithm is used to normalize the signal intensity to make the data of different samples comparable. Database matching: The HMDB (Human Metabolome Database) metabolite database is used as the comparison database. The mass - to - charge ratio of metabolites in the pre - processed mass spectrometry data is matched with the exact mass of known metabolites in the database, and a certain mass error range (usually at the ppm level) is set. The formula is: ; where is the observed mass - to - charge ratio, is the expected mass - to - charge ratio of metabolites in the database; For metabolites that can generate fragment ions, the observed fragment ion patterns are matched with the fragment patterns in the database using a similarity metric, and the formula is: ; where and are the signal intensities of the fragment ions observed for the i-th metabolite and the metabolite fragment ions in the database; Combining exact mass matching and fragment ion matching, a comprehensive score is calculated for each metabolite to determine the final matching result; 2) Metabolic pathway enrichment analysis: Using the KEGG (Kyoto Encyclopedia of Genes and Genomes) metabolic pathway database as the comparison database, which contains information on metabolic pathways and the roles of metabolites in the pathways, the list of identified metabolites is used as input, and at the same time, the relative abundance information of each metabolite in different samples is obtained; Let the overall metabolite set be , where there are metabolites belonging to a certain pathway, and there are in the identified metabolite set, and belong to this pathway. The enrichment significance is calculated using the hypergeometric distribution formula: ; represents the probability of randomly selecting metabolites from the given population, among which there are or more metabolites belonging to this pathway. The smaller the value, the more significant the enrichment of this pathway; Calculate the relative abundance change of metabolites under different physiological or pathological states, and the formula is: ; Combining value and value, screen out the significantly changed and enriched metabolic pathways; According to the enriched metabolic pathways, combined with biological knowledge, explain the functional significance of these pathways under specific physiological or pathological states, such as which metabolic pathways are up-regulated or down-regulated and the possible affected physiological functions, etc.
[0018] Furthermore, as Figure 5 shown, the implementation method of the association relationship mining module includes the following steps: Data integration: Integrate the data from the gene analysis module, protein analysis module, and metabolomics analysis module into a dataset; Association analysis: Calculate the Pearson and Spearman correlation coefficients between different types of data in the above dataset to determine the linear and monotonic relationships between them; perform partial least squares regression to find the relationship between the predictor variables and the response variables; based on the results of correlation and regression analysis, use network analysis methods to construct a comprehensive network and perform module analysis; Visualization presentation: Select visualization tools and graph types, and draw networks, heatmaps, scatter plots, or bar charts according to the analysis results.
[0019] Furthermore, the data integration includes the following steps: Gene analysis data: Obtain information such as gene sequence alignment results, gene expression pattern analysis results, and predicted functions of genes under different physiological states from the gene analysis module. These information may include gene sequence similarity scores, gene expression levels, gene function prediction results, etc.
[0020] Protein analysis data: Include three-dimensional structure information of proteins, protein interaction network information, and functional relationship data between proteins. For example, protein structure features obtained from a deep learning-based protein structure prediction model, and protein interaction strength and network topology information determined by protein interaction network analysis algorithms.
[0021] Metabolomics analysis data: Obtain metabolite species identification results, metabolic pathway enrichment analysis results, and metabolite abundance data, etc. These data are used to understand the status of metabolites and changes in metabolic pathways.
[0022] Data integration: Integrate the data from the gene analysis module, protein analysis module, and metabolomics analysis module into a single dataset.
[0023] Furthermore, the association analysis includes the following steps: Correlation analysis: The Pearson correlation coefficient is used to measure the linear correlation between two variables, and the formula is: ; Where, and are the observed values of two randomly selected variables in the dataset (for example, gene expression level and metabolite abundance, and the specific two values can be set according to needs), and are the means of the above two variables, is the number of observed values; The Spearman rank correlation coefficient is used to measure the monotonic relationship between two variables, does not depend on the distribution of the data, and is applicable to non-linear relationships. The formula is: ; where is the rank difference between two selected variables in the above dataset, is the sample size; Pearson correlation coefficient and Spearman rank correlation coefficient are used to analyze the relationship between gene expression levels and metabolite abundances, or the relationship between protein interaction strengths and metabolic pathway activities, to determine whether there is a linear or monotonic association between the data; Partial least squares regression PLS: Through the partial least squares regression PLS algorithm, a potential structural relationship between the predictor variables and the response variables is constructed: ; ; wherein, is the predictor variable matrix (such as gene expression and protein data), is the response variable matrix (such as metabolomics data or physiological phenotypes), and are latent variables, and are loading matrices, and are residual matrices; PLS is used to find the relationship between the predictor variables and the response variables, such as finding which gene and protein features are most relevant to the changes in metabolites.
[0024] Network analysis: Taking genes, proteins, and metabolites as nodes, a weighted network is constructed based on the correlations obtained from the above analysis; if the correlation exceeds a certain threshold, an edge is added between the corresponding nodes, and the weight of the edge is the absolute value of the correlation or other association strength metrics; Using the Louvain detection algorithm, the network is divided into different modules, and the formula is: ; wherein, is the number of edges, is the degree of node , is the degree of node within the module, and are the module labels of node and node , is an indicator function that is 1 when and 0 otherwise. Module analysis is used to discover sets of genes, proteins, and metabolites with close associations, and these sets may cooperate functionally.
[0025] Furthermore, the visual presentation includes the following steps: Network visualization: Use Cytoscape to visualize the constructed comprehensive network. For the nodes in the network, different shapes and colors are used to represent different types of entities (genes, proteins, metabolites), and the width and color of the edges represent the association strength or correlation magnitude. Heatmap visualization: Use the Matplotlib library in Python to draw a heatmap to show the correlations between genes, proteins, and metabolites; present the correlation matrix or the results of principal component analysis in the form of a heatmap, where the shade of the color represents the strength of the correlation or the contribution degree of the principal component, and the rows and columns can represent different genes, proteins, and metabolites respectively. Scatter plots and bar charts: For paired association relationships, including the relationship between gene expression levels and metabolite abundances, use scatter plots to show, where the X-axis represents the gene expression level, the Y-axis represents the metabolite abundance, and each point represents a sample; for the results of principal component analysis or enrichment analysis, use bar charts to show the proportion of the explained variance of different principal components or the enrichment significance of different metabolic pathways.
[0026] Through the above steps and analysis methods, the association relationship mining module can effectively mine the association relationships between gene, protein, and metabolomics data, and display these relationships in an intuitive way, providing a comprehensive perspective for studying the physiological and pathological states of rodents.
[0027] In summary, the platform can integrate various types of rodent biological information, and is efficient in processing rodent biological information and accurate in analysis.
Claims
1. A biological information data processing platform for rodents, characterized in that include: Data collection and integration module: used to collect rodent bioinformatics data from multiple data sources, including gene sequence data, protein structure data and metabolomics data; By formulating a unified data format standard, data in different formats can be converted and integrated and stored in the database module; Gene analysis module: used to implement gene sequence alignment using a gene sequence alignment algorithm, and to analyze gene expression patterns using a machine learning algorithm to predict the role of genes under different physiological states; Protein analysis module: Through the protein structure prediction model based on deep learning, the three-dimensional structure of the protein is predicted according to the amino acid sequence; using the protein interaction network analysis algorithm, the rodent protein interaction network is constructed to obtain the functional relationship between proteins; Metabolomics analysis module: uses mass spectrometry data analysis algorithm combined with database matching to identify the types of metabolites; Through metabolic pathway enrichment analysis, we can identify metabolic pathways associated with specific physiological or pathological states; Association mining module: Based on the results of gene analysis, protein analysis and metabolomics analysis, the association between genes, proteins and metabolomics and the species of rodents is mined and the association is visualized; User interaction module: used to provide a friendly user interface through which users can set permission management functions, upload data, select analysis functions and view analysis results; The database module is used to store the data received and generated by the platform.
2. A biological information data processing platform for rodents as claimed in claim 1, characterized in that: The processing method of the gene analysis module comprises the following steps: 1) Gene sequence alignment: Obtaining gene sequence data to be compared from a database module, including a query sequence and a reference sequence, and establishing an index for the reference sequence; dividing the query sequence into seed segments, and adjusting the length and position of the seeds according to the sequence characteristics of rodent genes, searching for matching seeds in the index, matching the found seeds, performing a bidirectional extended comparison, using a scoring matrix for rodent genes to calculate the comparison score, determining whether it is a valid comparison according to a set threshold, and using a greedy algorithm or a dynamic programming algorithm to merge and optimize the matching segments to obtain a final sequence comparison result; 2) Analysis of gene expression patterns: Feature selection and extraction: Extract features from gene expression data for training and prediction of machine learning models. The features include gene expression levels, expression change folds, and differences between different samples. Construction and training of machine learning models: Use support vector machine (SVM) as the machine learning model, divide the labeled feature data set into training set and test set, and use the training set and test set to train the machine learning model. By optimizing the kernel function of the model and adjusting the hyperparameters of the model, the model can achieve optimal performance on the training set. Prediction and evaluation: The gene expression features to be predicted are input into the trained machine learning model to obtain the prediction results of the gene's role in different physiological states.
3. A biological information data processing platform for rodents as claimed in claim 1, characterized in that: The processing method of the protein analysis module comprises the following steps: 1) Protein structure prediction: Protein structure model construction: AlphaFold architecture is used as the basic architecture for protein structure prediction, including input layer, middle layer and output layer; the input layer uses one-hot encoding to represent each amino acid sequence of the protein as a high-dimensional vector; the middle layer includes multiple Transformer modules to learn long-range dependencies in the amino acid sequence; the output layer is used to output the predicted protein structure information; Training data collection and preprocessing: Collect rodent protein data with known structures, clean the amino acid sequences, and remove incomplete or erroneous sequence information; convert the protein structure into the input representation required by the model, increase the diversity of the data through rotation, translation, and noise addition operations, and improve the generalization ability of the model; Model training: Divide the preprocessed data into training set, validation set and test set, use the training set to train the model, adjust the model hyperparameters according to the performance of the validation set, and use the test set to evaluate the performance of the final model; Structure prediction: For new rodent protein amino acid sequences, they are encoded as input vectors and input into the trained model. The model outputs predicted protein structural features, which are then converted into protein three-dimensional structures through a structural reconstruction algorithm. 2) Protein interaction network analysis: Network construction: Represent proteins as nodes and interactions as edges to construct an undirected graph ,in is the set of protein nodes, is a set of edges, with different weights assigned to the edges according to the strength of their interactions; Network analysis algorithm: Use the Markov clustering algorithm MCL to divide the nodes in the undirected graph into different clusters, expand and inflate the adjacency matrix, enhance the connection pattern between nodes, and iteratively update the matrix to finally converge to a stable state and obtain different clusters; Functional relationship mining: Based on the clustering results and combined with the functional annotation information of proteins, a functional enrichment analysis is performed on each cluster using hypergeometric distribution, and the functional relationship between proteins is obtained based on the analysis results.
4. A biological information data processing platform for rodents as claimed in claim 1, characterized in that: The implementation method of the metabolomics analysis module comprises the following steps: 1) Identification of metabolites: Mass spectrometry data acquisition: The metabolomics mass spectrometry data of rodent biological samples were obtained by liquid chromatography-mass spectrometry (LC-MS), and the metabolomics mass spectrometry data included the mass-to-charge ratio of metabolites and the corresponding signal intensity; Data preprocessing: The centroid method is used to identify the peaks of metabolite ions from the raw mass spectrometry data. The centroid method finds the local maximum of the signal intensity as the center of the peak. The formula is: ; in, is the mass-to-charge ratio of the ith metabolite, is the signal intensity of the ith metabolite, and is the scan range where the peak is located; The dynamic time warping (DTW) algorithm was used to align the peaks obtained, the Savitzky-Golay filtering algorithm was used to remove noise, and the total ion current normalization algorithm was used to normalize the signal intensity, so that the data of different samples were comparable. Database Matching: Using the HMDB metabolite database as a comparison database, the metabolite mass-to-charge ratio in the preprocessed mass spectrometry data was matched with the accurate mass of the known metabolites in the database, and a certain mass error range was set. The formula is: ; in, is the observed mass-to-charge ratio, is the expected mass-to-charge ratio of the metabolite in the database; For metabolites that can produce fragment ions, the observed fragment ion patterns were matched with the fragment patterns in the database using a similarity metric, which is formulated as: ; in, and is the signal intensity of the observed fragment ion of the ith metabolite and the fragment ion of the metabolite in the database; Combining accurate mass matching and fragment ion matching, a comprehensive score is calculated for each metabolite to determine the final matching result; 2) Metabolic pathway enrichment analysis: The KEGG metabolic pathway database was used as a comparison database, and the identified metabolite list was used as input to obtain the relative abundance information of each metabolite in different samples; Assume that the total metabolite set is , among which the metabolites belonging to a certain pathway are There are a total of , belonging to this channel are The hypergeometric distribution formula was used to calculate the enrichment significance: ; In a given population, randomly select metabolites, including The probability of more than one metabolite belonging to the pathway. The smaller the value, the more significant the enrichment of the pathway. The relative abundance changes of metabolites under different physiological or pathological conditions were calculated using the formula: ; Combination Value and Values were used to screen out significantly changed and enriched metabolic pathways; Based on the enriched metabolic pathways and combined with biological knowledge, the functional significance of these metabolic pathways under specific physiological or pathological conditions can be obtained.
5. A biological information data processing platform for rodents as claimed in claim 1, characterized in that: The implementation method of the association mining module includes the following steps: Data integration: Integrate the data from the gene analysis module, protein analysis module, and metabolomics analysis module into one dataset; Correlation analysis: Pearson and Spearman correlation coefficients were calculated between different types of data for the above data sets to determine the linear and monotonic relationships between them; partial least squares regression was performed to find the relationship between the predictor variables and the response variables; based on the correlation and regression analysis results, a comprehensive network was constructed using network analysis methods, and module analysis was performed; Visualization: Select a visualization tool and graph type to draw a network, heat map, scatter plot, or bar chart based on the analysis results.
6. A biological information data processing platform for rodents as claimed in claim 5, characterized in that: The data integration includes the following steps: Gene analysis data: Obtain gene sequence alignment results, gene expression pattern analysis results, and predicted effects of genes under different physiological states from the gene analysis module; Protein analysis data: including protein three-dimensional structure information, protein interaction network information, and functional relationship data between proteins; Metabolomics analysis data: obtain metabolite species identification results, metabolic pathway enrichment analysis results and metabolite abundance data; Data integration: The data from the gene analysis module, protein analysis module, and metabolomics analysis module were integrated into one dataset.
7. A biological information data processing platform for rodents as claimed in claim 5, characterized in that: The association analysis comprises the following steps: Correlation analysis: The Pearson correlation coefficient is used to measure the linear correlation between two variables. The formula is: ; in, and are the random observations of the two variables in the data set, and is the average of the above two variables, is the number of observations; The Spearman rank correlation coefficient is used to measure the monotonic relationship between two variables. It does not depend on the distribution of the data and is applicable to nonlinear relationships. The formula is: ; in is the rank difference of the two variables selected in the above data set, is the sample size; Pearson correlation coefficient and Spearman rank correlation coefficient were used to analyze the relationship between gene expression level and metabolite abundance, or the relationship between protein interaction strength and metabolic pathway activity to determine whether there was a linear or monotonic association between the data; Partial Least Squares Regression PLS: The potential structural relationship between the predictor variables and the response variables is constructed through the partial least squares regression PLS algorithm: ; ; in, is the predictor variable matrix, is the response variable matrix, and is a latent variable, and is the loading matrix, and is the residual matrix, and the relationship between the predictor and response variables is found through the above formula; Network Analysis: Genes, proteins, and metabolites are used as nodes, and a weighted network is constructed based on the correlations obtained from the above analysis; if the correlation exceeds a certain threshold, an edge is added between the corresponding nodes, and the weight of the edge is the absolute value of the correlation or other association strength indicators; Using the Louvain detection algorithm, the network is divided into different modules. The formula is: ; in, The number of edges, Is a node The degree, Is a node Degree within the module, and Is a node and nodes The module tag, is an indicator function, when 1 if the value is 0, otherwise it is 0.
8. A biological information data processing platform for rodents as claimed in claim 5, characterized in that: The visual presentation comprises the following steps: Network visualization: Use Cytoscape to visualize the constructed comprehensive network. Different shapes and colors are used to represent different types of entities for nodes in the network, and the width and color of the edge represent the strength of association or the size of correlation. Heatmap visualization: Use Python's Matplotlib library to draw heatmaps to show the correlations between genes, proteins, and metabolites; display the correlation matrix or principal component analysis results in the form of a heatmap, with the color depth indicating the strength of the correlation or the contribution of the principal component. Rows and columns can represent different genes, proteins, and metabolites respectively. Scatter plots and bar graphs: For paired associations, including the relationship between gene expression levels and metabolite abundance, scatter plots are used for display, where the X-axis represents gene expression levels, the Y-axis represents metabolite abundance, and each point represents a sample; for principal component analysis results or enrichment analysis results, bar graphs are used to display the explained variance ratio of different principal components or the enrichment significance of different metabolic pathways.
Citation Information
Patent Citations
Cloud computation-based biological information analysis platform
CN108694305A
Multi-dimensional annotation space metabolome database establishment method
CN115201391A
Sequencing and analyzing method for eukaryotic transcriptome of canna edulis and plantain
CN118692561A
Nucleic acid binding protein recognition method based on protein map and protein language model
CN119252348A
Traditional Chinese medicine meridian prescription system for analyzing extracellular vesicles based on multi-omics integration
CN119339972A