Method, apparatus, server, and medium for generating a prediction model of a gene regulatory network

Through the method of combining a comparison learning of gene expression matrix encoder and text encoder, the gene regulation network prediction model is optimized, which solves the problem of insufficient accuracy and generalization capabilities in the existing technology, and achieves higher accuracy and robust gene regulation network prediction.

CN119649894BActive Publication Date: 2025-07-04INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510164703.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-04
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The accuracy and generalization ability of predictive model prediction models of gene regulation networks in the prior art have decreased, mainly because the characteristic information of gene expression data cannot be fully captured when manually selecting characteristic parameters.

Method used

The dynamic changes in the gene expression time series were captured by the gene expression matrix encoder, biological text features were extracted in combination with the text encoder, gene expression and text features were aligned using a comparison learning method, and the model was optimized using the cosine decay learning rate scheduling mechanism.

Benefits of technology

The accuracy and generalization ability of the gene regulation network prediction model are improved, the correlation model between gene expression data and biological text description is realized, and the robustness of the model and the ability to adapt to complex multimodal data is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649894B_ABST
    Figure CN119649894B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, server, and medium for generating a prediction model of a gene regulatory network, which relates to the field of computer technology. The gene expression matrix encoder is used to capture the gene activity changes at different time points in the gene expression time series to obtain global gene expression features. The text encoder is used to analyze the correspondence between gene expression patterns and text semantics. Through the method of contrastive learning, the gene expression feature data at multiple time points is aligned with the biological text feature data. The cosine decay learning rate scheduling mechanism is used to gradually and smoothly reduce the learning rate to ensure that the model is closer to the global optimal solution. The prediction model of the gene regulatory network is used to establish the correlation modeling between the gene expression data at multiple time points and the biological text description, improving the accuracy and generalization ability of the statistical model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a method, device, server, and medium for generating a prediction model of a gene regulatory network. Background Art

[0002] Gene expression data is an important concept in the fields of genetics and bioinformatics. It reflects the abundance of gene transcripts measured directly or indirectly in cells. The gene expression data records the gene expression levels under different conditions and provides resources for studying gene regulation. A gene regulatory network refers to a network formed by the interaction relationships between genes in a cell. The gene regulatory network is used to describe the regulatory relationships between genes. By predicting the gene regulatory network, key regulatory factors and regulatory paths can be obtained therefrom, and then it can be inferred which genes have important effects on the expression of other genes.

[0003] In the current related technologies, the method for predicting a gene regulatory network is mainly machine learning, that is: using gene expression data as input features and inputting them into a prediction model to output the prediction result of the gene regulatory network. However, in the related technologies, the input features are manually selected and optimized. The complexity of gene expression data makes it impossible to comprehensively capture the feature information of gene expression data when manually selecting feature parameters. And the performance of the prediction model depends on the selection and optimization of feature parameters. The method of manually screening feature parameters reduces the analysis accuracy and generalization ability of the prediction model. Summary of the Invention

[0004] The present application provides a method, device, server, and medium for generating a prediction model of a gene regulatory network to at least solve the problem of the decline in the analysis accuracy and generalization ability of the prediction model in the related technologies.

[0005] The present application provides a method for generating a prediction model of a gene regulatory network, including:

[0006] Obtaining a gene expression time series;

[0007] Inputting the gene expression time series into a gene expression matrix encoder to output global gene expression features;

[0008] Obtaining biological text data;

[0009] Inputting the biological text data into a text encoder to output global text features;

[0010] Aligning the global gene expression features with the global text features to obtain a contrast loss;

[0011] Optimizing a prediction model according to the contrast loss to generate an optimized prediction model;

[0012] Train the optimized prediction model through a cosine annealing learning rate scheduling mechanism to generate a trained prediction model.

[0013] This application also provides a device for generating a prediction model of a gene regulatory network, including:

[0014] A first acquisition module for acquiring gene expression time series;

[0015] A first output module for inputting the gene expression time series into a gene expression matrix encoder to output global gene expression features;

[0016] A second acquisition module for acquiring biological text data;

[0017] A second output module for inputting the biological text data into a text encoder to output global text features;

[0018] An alignment module for aligning the global gene expression features with the global text features to obtain a contrastive loss;

[0019] An optimization module for optimizing a prediction model according to the contrastive loss to generate an optimized prediction model;

[0020] A training module for training the optimized prediction model through a cosine annealing learning rate scheduling mechanism to generate a trained prediction model.

[0021] This application also provides a server, including: a memory for storing a computer program; a processor for implementing the steps of any of the above methods for generating a prediction model of a gene regulatory network when executing the computer program.

[0022] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above methods for generating a prediction model of a gene regulatory network.

[0023] This application also provides a computer program product including a computer program, which, when executed by a processor, implements the steps of any of the above methods for generating a prediction model of a gene regulatory network.

[0024] With this application, since the gene expression matrix encoder captures the changes in gene activities at different time points in the gene expression time series, extracts the feature representation of the time series to obtain the global gene expression features, extracts the global semantic features through the text encoder, analyzes the correspondence between the gene expression pattern and the text semantics, and through the method of contrastive learning, aligns the multi-time point gene expression feature data with the biological text feature data. By means of the cosine decay learning rate scheduling mechanism, the learning rate is gradually and smoothly reduced to ensure that the model is closer to the global optimal solution. Through the prediction model of the gene regulatory network, the correlation modeling between the multi-time point gene expression data and the biological text description is realized, improving the accuracy and generalization ability of the statistical model. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0026] Figure 1 Schematic diagram of the system structure of the server provided by the embodiment of this application;

[0027] Figure 2 Schematic diagram of the flow of the method for generating the prediction model of the gene regulatory network provided by the embodiment of this application;

[0028] Figure 3 Schematic diagram of the contrastive learning provided by the embodiment of this application;

[0029] Figure 4 Schematic diagram of the structure of the device for generating the prediction model of the gene regulatory network provided by the embodiment of this application;

[0030] Figure 5 Schematic diagram of the structure of the server provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of this application.

[0032] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0033] First, the nouns involved in this application are explained:

[0034] Gene expression: It refers to the process of synthesizing functional gene products from the genetic information of genes. Gene expression products are usually proteins, and gene expression can be regulated by regulating several of its steps, including transcription, RNA splicing, translation, and post-translational modification, to achieve the regulation of gene expression.

[0035] Attention mechanism: The attention mechanism is a data processing method in machine learning, and its core goal is to select the information that is more critical to the current task goal from numerous information. It is widely used in various types of machine learning tasks such as natural language processing, image recognition, and speech recognition.

[0036] Query vector: A query vector refers to the representation of the input sequence that the model currently wants to focus on or process. It is usually generated from the last hidden state obtained by the encoder or the current decoder state. The query vector represents the information that the model is currently focusing on and can be regarded as a question or content that needs to be judged.

[0037] Multimodal: Multimodal refers to the technology in the fields of artificial intelligence, etc., that simultaneously uses two or more different senses, data types, or information sources for interaction, processing, and analysis. Multimodal technology integrates and utilizes two or more different senses, data types, or information sources, such as text, images, audio, and video, for interaction, processing, and analysis. By integrating and utilizing the relevance and differences between these different modalities, to achieve higher-level tasks such as feature extraction, classification, and retrieval.

[0038] Contrastive loss: Contrastive loss is a metric learning loss function that makes the distance between similar sample pairs smaller in the feature space and the distance between dissimilar sample pairs larger by learning the similarity and difference between sample pairs.

[0039] To solve the problem of the decline in the analysis accuracy and generalization ability of the prediction model in the prior art, the embodiments of the present application propose the following technical concepts: Considering the time series characteristics of gene expression data, the inventor designed a gene expression matrix encoder capable of capturing time dynamic changes. By using the gene expression matrix encoder to capture the gene activity changes at different time points, global gene expression features are obtained. The inventor constructed a text encoder based on a language model, extracted global semantic features through the text encoder, analyzed the correspondence between gene expression patterns and text semantics, and obtained global text features. Considering the method of contrastive learning, the gene expression feature data at multiple time points is aligned with the biological text feature data. The learning rate is gradually and smoothly reduced through the cosine decay learning rate scheduling mechanism to ensure that the model is closer to the global optimal solution.

[0040] To enable those skilled in the art of the present technology to better understand the solution of the present application, the following further detailed description of the present application will be given in conjunction with the accompanying drawings and specific implementation manners.

[0041] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the prediction model generation method of the gene regulatory network depends, the specific application environment architecture or specific hardware architecture will be described here. Refer to Figure 1 , Figure 1 This is a schematic diagram of the system structure of the server provided by the embodiments of the present application. As Figure 1 shown, the server includes: a receiving device 101, a processor 102, and a display device 103.

[0042] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the item recognition method. In other feasible embodiments of the present application, the above architecture may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements, which can be specifically determined according to the actual application scenario and will not be limited here. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0043] In the specific implementation process, the receiving device 101 can be an input / output interface or a communication interface, and can obtain the gene expression time series.

[0044] The processor 102 can generate a trained prediction model.

[0045] The display device 103 can be used to display the above trained prediction model, etc.

[0046] The display device can also be a touch display screen, which is used to receive user instructions while displaying the above content to achieve operation interaction with the user.

[0047] It should be understood that the above-mentioned processor can be implemented by a processor reading instructions in a memory and executing the instructions, or can be implemented by chip circuits.

[0048] In addition, the network architecture and service scenarios described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art can know that with the evolution of the network architecture and the emergence of new service scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0049] Figure 2 It is a schematic flowchart of a method for generating a prediction model of a gene regulatory network provided by an embodiment of the present application. As Figure 2 shown, the method includes:

[0050] S201: Obtain gene expression time series.

[0051] In this embodiment, the gene expression time series data shows the dynamic regulation mechanism of an organism in the time dimension by capturing the changes in gene activity at different time points, covering the entire time frame from the initial stimulus to the end of the biological process.

[0052] Exemplarily, in psoriasis research, imiquimod is applied to mice as an intervention means at the zero time point, and gene expression data is sampled at the acute phase, remission phase, and relapse phase respectively to explore the dynamic changes of key genes in the skin pathological process.

[0053] In this embodiment, the data set representation form of the gene expression time series is as follows:

[0054]

[0055] In the formula, represents the data set of the gene expression time series; represents the gene expression time series matrix of the i-th sample; represents the text annotation.

[0056] Among them, the content of the text annotation includes but is not limited to the experimental design background, gene function annotation, biological significance of time points, and processing conditions.

[0057] In this embodiment, each sample contains time points, where is expressed as:

[0058]

[0059] In the formula, M represents the number of genes analyzed; F represents the feature dimension of each time point; represents a real space matrix of dimension; represents the th sample at a time point.

[0060] Among them, the feature dimensions at each time point include but are not limited to the gene expression level FPKM (Fragments Per Kilobase of exon model per Million mapped fragments), the gene expression level TPM (Transcripts Per Million), and additional statistical values under experimental conditions.

[0061] Specifically, in order to standardize the training process and adapt to the fixed input structure of the model, the sample time points are sampled, and the maximum number of time points for each sample is limited to k, expressed as:

[0062]

[0063] In the formula, represents the set of sampled samples; represents the th sampled sample; represents the index set selected from satisfying .

[0064] In this embodiment, biological key time points will be preferentially preserved during sampling.

[0065] Exemplarily, in psoriasis research, acute-phase samples reflect the strong inflammatory response of the mouse immune system, remission-phase samples reveal the activation of immunosuppressive signals, and relapse-phase samples can show the potential molecular mechanisms of pathological recurrence.

[0066] Specifically, when the number of time points of the sample is less than the maximum number of time points for each sample, that is, , pad with zero-filled time points, expressed as:

[0067]

[0068] In the formula, represents the sample after supplementing time points; represents the th sample at a time point; is the padding matrix, initialized to zero; represents a real space matrix of dimension; represents the number of filled time points, a total of ones.

[0069] In this embodiment, the padding strategy serves as a placeholder and does not introduce additional biological information.

[0070] S202: Input the gene expression time series into the gene expression matrix encoder to output global gene expression features.

[0071] Specifically, input the gene expression matrix in the gene expression time series into the feature extraction model to obtain an embedding vector, calculate the weights of each time point in the gene expression time series through the attention mechanism, and calculate the global gene expression features according to the weights.

[0072] S203: Obtain biological text data.

[0073] In this embodiment, integrate biomedical corpora during the pre-training stage to obtain biological text data.

[0074] Among them, the sources of biomedical corpora include but are not limited to corpora, websites, and articles.

[0075] S204: Input the biological text data into the text encoder to output global text features.

[0076] Specifically, dynamically annotate the time points in the biological text data through the attention mechanism, screen the key time points to retain the core biological text information, and assign different weights to different time points through weighted feature aggregation and normalization of the attention mechanism; update the vocabulary, perform word segmentation on the biological text, input the segmented text into the language model, and extract the hidden state sequence of the output to obtain global text features.

[0077] S205: Align the global gene expression features with the global text features to obtain a contrastive loss.

[0078] Specifically, normalize the global text features, calculate the similarity of the feature vectors according to the global gene expression features and the global text features, and calculate the bidirectional loss function according to the similarity of the feature vectors: the loss function from gene expression to text and the loss function from text to gene expression, and obtain the contrastive loss according to the bidirectional loss function.

[0079] S206: Optimize the prediction model according to the contrastive loss to generate an optimized prediction model.

[0080] Specifically, judge the semantic consistency between the gene expression time series and the corresponding text annotations according to the contrastive loss, and obtain the optimized prediction model by dynamically adjusting the weight parameters.

[0081] Figure 3 It is a schematic diagram of the contrastive learning provided by the embodiment of the present application.

[0082] As shown Figure 3 in the figure, the gene expression matrix encoder integrates gene expression features at different time points according to the attention mechanism to obtain global gene expression features; the text encoder analyzes text semantics, combines context and domain knowledge, obtains biological semantics and temporal relationships, and obtains global text features; the text semantics and gene expression data are aligned through contrastive learning to optimize the prediction model.

[0083] S207: Train the optimized prediction model through the cosine decay learning rate scheduling mechanism to generate the trained prediction model.

[0084] Specifically, by introducing cosine decay, based on the cosine function, calculate the learning rate of each iteration according to the initial learning rate, optimize the prediction model through batch gradient descent, and optimize the matching gene expression and text according to the contrastive loss.

[0085] In this embodiment, during the training process, the prediction model learns the semantic consistency between gene expression features and text descriptions, dynamically adjusts parameters to improve robustness under different data distributions.

[0086] In addition, it should be noted that in the inference stage, the prediction model generates aggregated features according to the new gene expression time series data that has not been processed through the feature representation ability in the training stage , calculates the similarity between the aggregated features and the pre-stored text feature representations to determine the matching degree between the two.

[0087] Among them, the formula for calculating the similarity between the aggregated features and the pre-stored text features is:

[0088]

[0089] In the formula, represents the similarity between the aggregated features and the pre-stored text features; represents the aggregated features; represents the pre-stored text features.

[0090] Specifically, in the inference stage, the prediction model identifies gene expression patterns related to specific biological processes or experimental conditions by calculating the similarity between gene expression features and text descriptions, displays key features in the dynamic regulation process, generates semantic annotations for unlabeled gene expression data, and predicts the potential functions of unknown genes or genes with unclear functions.

[0091] As can be seen from the above embodiments, the gene expression matrix encoder captures the changes in gene activity at different time points in the gene expression time series, extracts the feature representation of the time series to obtain the global gene expression features, extracts the global semantic features through the text encoder, analyzes the correspondence between the gene expression pattern and the text semantics, and aligns the gene expression feature data at multiple time points with the biological text feature data through the method of contrastive learning. The learning rate is gradually and smoothly decreased through the cosine decay learning rate scheduling mechanism to ensure that the model is closer to the global optimal solution. The correlation modeling between the gene expression data at multiple time points and the biological text description is realized through the prediction model of the gene regulatory network, improving the accuracy and generalization ability of the statistical model.

[0092] In one embodiment of the present application, step S202 includes:

[0093] S2021: Obtain the gene expression matrix in the gene expression time series.

[0094] In this embodiment, the gene expression matrix at each time point in the gene expression time series is denoted as .

[0095] S2022: Input the gene expression matrix into the feature extraction model to output the embedding vector.

[0096] In this embodiment, the feature extraction model includes, but is not limited to, a convolutional neural network model, a multi-head attention mechanism model, and a wavelet transform model.

[0097] In this embodiment, the embedding vector is expressed as:

[0098]

[0099] In the formula, represents the embedding vector; represents the number of features; represents the embedding dimension; represents a real space matrix of dimension; represents the gene expression matrix; Encoder() represents encoding.

[0100] Among them, the global feature at each time point is extracted by the embedding of a specific marker as the representative feature corresponding to this time point.

[0101] Exemplarily, the specific markers include, but are not limited to, the [CLS] marker, the [SEP] marker, and the [PAD] marker.

[0102] S2023: Calculate the weight at each time in the gene expression time series according to the embedding vector through the attention mechanism.

[0103] In this embodiment, the formula for calculating the weight of each time in the gene expression time series according to the embedding vector through the attention mechanism is as follows:

[0104]

[0105] In the formula, represents the weight; represents the attention scoring function; represents the feature at the -th time point; represents the query vector; represents the feature at the -th time point; represents the number of time points.

[0106] In this embodiment, the attention scoring function is used to measure the correlation between the time point feature and the query vector.

[0107] In this embodiment, the query vector can be initialized as the global query vector learned during the training process, or can be generated from the prior information of the experimental design.

[0108] Exemplarily, in psoriasis research, the gene expression in the acute phase usually plays a role in studying the immune mechanism of the disease, and its attention weight can be dynamically adjusted through training.

[0109] S2024: Calculate and generate the global gene expression feature according to the embedding vector and the weight of each time in the gene expression time series.

[0110] In this embodiment, the formula for calculating and generating the global gene expression feature according to the embedding vector and the weight of each time in the gene expression time series is as follows:

[0111]

[0112] In the formula, represents the global gene expression feature; represents the weight; represents the time point feature; represents the number of time points.

[0113] In this embodiment, the introduction of the attention mechanism enables the model to dynamically adjust the degree of attention to different time points, highlighting the feature extraction of the key time periods of biological processes.

[0114] Exemplarily, in a mouse model of psoriasis, the acute phase is usually accompanied by a strong immune response, while the remission phase involves an upregulation of immunosuppressive signals. The attention mechanism can better capture important immune signals by assigning higher weights to the acute-phase time points. Meanwhile, for the time-point data including the relapse phase, the attention mechanism can focus on the key changes before and after relapse, revealing potential relapse triggers.

[0115] As can be seen from the above embodiments, according to the temporal characteristics of gene expression data, through the dynamic feature aggregation method based on the attention mechanism, the model can extract the correlations between time points and dynamically adjust the weights of each time point according to the importance of the time points, avoiding over-smoothing and ignoring detailed information, and improving the model's understanding ability of time series.

[0116] In an embodiment of the present application, step S204 includes:

[0117] S2041: Convert the biological text data into a text token sequence.

[0118] In this embodiment, the text token sequence is represented as:

[0119]

[0120] In the formula, represents the text token sequence; represents the maximum sequence length; represents the th word or sub-word unit.

[0121] S2042: Perform text processing on the text token sequence through a word segmentation algorithm to obtain the segmented text.

[0122] Specifically, by recording the co-occurrence frequencies of characters or sub-words in the corpus, updating the initialized vocabulary according to the set number of iterations, and performing text processing on the text token sequence according to the updated vocabulary, the segmented text is obtained.

[0123] S2043: Input the segmented text into a language model to output a hidden state sequence.

[0124] In this embodiment, the formula for inputting the segmented text into a language model to output a hidden state sequence is:

[0125]

[0126] In the formula, represents the hidden state sequence; represents the segmented text sequence; represents a real matrix space of dimension; Indicates the sequence length; Indicates the embedding dimension; BERT() indicates processing by the language model.

[0127] In this embodiment, the language model is a pre-trained language model.

[0128] S2044: Extract the hidden state from the hidden state sequence to output the global text feature.

[0129] In this embodiment, the global text feature is represented as:

[0130]

[0131] In the formula, Indicates the global text feature; Indicates the hidden state sequence.

[0132] As can be seen from the above embodiment, the global semantic feature of the text description is extracted through the text encoder. By introducing biological text data with time context, combining context and domain knowledge, the relationship between the semantics and time of biological text is captured, improving the text encoding ability.

[0133] In an embodiment of the present application, step S2042 includes:

[0134] S421: Create an initial vocabulary.

[0135] In this embodiment, the content recorded in the initial vocabulary is a basic set of character combinations, and the words in the text are decomposed into individual characters.

[0136] Exemplarily, the basic set of character combinations includes, but is not limited to, letters, numbers, and punctuation marks.

[0137] Exemplarily, the word "apple" is decomposed into the character sequence "a", "p", "p", "l", "e".

[0138] S422: Record the co-occurrence frequency of characters or sub-words in the corpus, and merge the sub-words with the highest co-occurrence frequency according to the set number of iterations to generate an updated vocabulary.

[0139] Specifically, the tokenization algorithm counts the co-occurrence frequency of adjacent character or sub-word pairs in the corpus, records the co-occurrence times, identifies common character or sub-word combinations, and in each iteration, merges the sub-words with the highest co-occurrence frequency to generate new sub-words.

[0140] Among them, the upper limit of the number of iterations is to reach the size of the vocabulary or the frequency threshold.

[0141] S423: Perform text processing on the text token sequence according to the updated vocabulary to obtain the segmented text.

[0142] Specifically, segment the text token sequence through the updated vocabulary. For each input word, match the longest sub-word unit from the updated vocabulary through a matching algorithm.

[0143] In this embodiment, the matching algorithm is the maximum matching algorithm or the greedy algorithm.

[0144] Exemplarily, if the vocabulary includes "Bus" and "##pirone", but does not record the word "Buspirone", the algorithm will decompose "Buspirone" into "Bus" and "##pirone" for recognition.

[0145] Among them, "##" represents a sub-word and needs to be connected to the previous sub-word.

[0146] As can be seen from the above embodiments, by decomposing words into smaller sub-word units through the segmentation algorithm, when a complete word is not in the vocabulary of the prediction model, the model can represent it through sub-word combination, reducing the probability that the prediction model fails to recognize out-of-vocabulary words.

[0147] In an embodiment of the present application, before step S2041, it further includes:

[0148] S301: Analyze the static semantics of the biological text data to generate the analyzed biological text data.

[0149] Specifically, analyze the morphology and grammar in the text data, and identify each text type in the text data according to the semantic rules to obtain the analyzed biological text data.

[0150] S302: Identify the biological process text in the analyzed biological text data.

[0151] In this embodiment, the content recorded in the biological text data includes but is not limited to time, experimental objects, and experimental steps.

[0152] S303: Perform time calibration on the biological process text to generate the calibrated biological text data.

[0153] Specifically, identify the sequence of occurrence of biological processes in the biological process text according to the attention mechanism, and calibrate the time information of each biological process to obtain the calibrated biological text data.

[0154] Exemplarily, the biological process text is described as follows: "During the acute phase, the immune response is significantly enhanced; subsequently, it enters the remission phase, and the immunosuppressive signals gradually increase." The time words "acute phase" and "remission phase" are calibrated, and the self-attention mechanism is used to identify the chronological relationship in the text description and assign weights corresponding to different time points.

[0155] As can be seen from the above embodiments, by using the attention mechanism to capture the time information of the chronological relationship of biological processes in the text data, the model can identify the chronological association in the text description, assign corresponding weights to different times, and improve the alignment effect between the text and the time series gene expression data.

[0156] In an embodiment of the present application, step S205 includes:

[0157] S2051: Calculate the similarity of the feature vectors based on the global gene expression feature and the global text feature.

[0158] In this embodiment, the similarity of the feature vectors is used to measure the alignment degree of the global gene expression feature and the global text feature in the multimodal space.

[0159] Among them, the similarity of the feature vectors is expressed as:

[0160]

[0161] In the formula, represents the transpose of the normalized global gene expression feature; represents the normalized global text feature; represents the similarity of the feature vectors.

[0162] S2052: Calculate the contrast loss between each gene expression feature in the global gene expression feature and the text feature according to the similarity of the feature vectors, and obtain the loss function from gene expression to text.

[0163] In this embodiment, the formula for calculating the contrast loss between each gene expression feature in the global gene expression feature and the text feature is:

[0164]

[0165] In the formula, represents the loss function from gene expression to text; represents the temperature parameter; represents the feature vector similarity of the gene expression feature; represents the feature vector similarity of the global gene expression feature and the global text feature; represents the number of gene expression features; represents the number of text features.

[0166] S2053: Calculate the contrast loss between each text feature in the global text feature and the gene expression feature based on the similarity of the feature vectors, and obtain the loss function from text to gene expression.

[0167] In this embodiment, the formula for calculating the contrast loss between each text feature in the global text feature and the gene expression feature is:

[0168]

[0169] In the formula, represents the loss function from text to gene expression; represents the temperature parameter; represents the similarity of the feature vectors of the text features; represents the similarity of the feature vectors of the global gene expression feature and the global text feature; represents the number of gene expression features; represents the number of text features.

[0170] S2054: Calculate the generated contrast loss according to the loss function from gene expression to text and the loss function from text to gene expression.

[0171] In this embodiment, the formula for calculating the generated contrast loss according to the loss function from gene expression to text and the loss function from text to gene expression is:

[0172]

[0173] In the formula, represents the contrast loss; represents the loss function from gene expression to text; represents the loss function from text to gene expression.

[0174] In addition, it should be noted that the contrast loss reflects the corresponding relationship in biology. In psoriasis research, the gene expression feature can capture the dynamic gene regulation patterns at different time points, such as the acute phase, remission phase, and relapse phase, while the text feature records the semantic descriptions of biological processes, such as the enhancement or inhibition of immune responses. According to the contrast loss between the gene expression feature and the text feature, the similarity of the matching is improved. The contrast loss has the following characteristics:

[0175] 1. The contrast loss can reflect the dynamic changes of gene expression over time. For example, high expression of inflammation-related genes may be detected in the acute phase, while downregulation of gene expression may be manifested in the remission phase. This mapping of dynamic changes helps the model capture the semantic consistency between gene regulation patterns and text descriptions.

[0176] 2. The contrastive loss can align the features of heterogeneous biological processes and reveal the synergistic relationships of multiple mechanisms. For example, gene expression features may reflect the activation state of signal transduction pathways, while text features describe the related molecular functions, thus achieving deep integration in the multimodal semantic space.

[0177] 3. The contrastive loss has the ability to detect abnormal patterns. When the model finds that certain gene expression features cannot be aligned with the text descriptions during the contrastive learning process, this scenario is recorded as a potential abnormal pattern, such as a new type of gene expression change not described in the experimental data.

[0178] As can be seen from the above embodiments, by calculating the loss function of the global gene expression features and the global text features, each gene expression feature is aligned with the correctly matched text feature and far from the unmatched text features, and each text feature is aligned with the correctly matched gene expression feature and far from the unmatched gene expression features. By calculating the bidirectional loss function from gene expression to text and from text to gene expression, the model simultaneously focuses on the bidirectional alignment processes from gene expression to text and from text to gene expression, thereby achieving the comprehensive alignment of multimodal features.

[0179] In an embodiment of the present application, before step S2051, it further includes:

[0180] Normalize the global text features to obtain the processed global text features.

[0181] In this embodiment, the process of text feature normalization is expressed as:

[0182]

[0183] In the formula, represents the normalized global text features; represents the unprocessed global text features; represents the absolute value of the global text features.

[0184] As can be seen from the above embodiments, through the normalization operation, it is ensured that the feature scales of the text and gene expression in the multimodal space are consistent, facilitating the optimization of contrastive learning; by automatically associating the corresponding time points in the gene expression data and combining biological semantics and temporal relationships, the text encoder can more accurately parse complex medical texts, providing convenience for diagnosis and prediction.

[0185] In an embodiment of the present application, step S207 includes:

[0186] S2071: Obtain the initial learning rate of the optimized prediction model.

[0187] In this embodiment, Represents the initial learning rate.

[0188] S2072: Calculate the learning rate after iteration according to the set number of iterations and the initial learning rate.

[0189] In this embodiment, the formula for calculating the learning rate after iteration according to the set number of iterations and the initial learning rate is:

[0190]

[0191] In the formula, Represents the learning rate after iteration; Represents the initial learning rate; Represents the number of iterations; Represents the total number of training steps.

[0192] S2073: Train the optimized prediction model according to the learning rate after iteration to generate a trained prediction model.

[0193] Specifically, adjust the weights of the prediction model according to the learning rate after multiple iterations to generate a trained prediction model.

[0194] As can be seen from the above embodiments, by introducing a cosine annealing learning rate scheduling mechanism during the training process, based on the cosine function, the learning rate is gradually decreased during the training process. The lower learning rate enables the model to focus on fine weight adjustments and avoid the model falling into a sub-optimal solution prematurely; by gradually and smoothly decreasing the learning rate, the problem of weight oscillation in the later stage of training can be effectively avoided, thereby improving the stability of the model; by using the cosine annealing learning rate scheduling mechanism to reduce the imbalance problem of gene expression and text features during the fusion process, the training efficiency and the adaptability of the model to complex multi-modal data are improved.

[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0196] Figure 4 It is a schematic structural diagram of a device for generating a prediction model of a gene regulatory network provided by an embodiment of the present application. As Figure 4 shown, the device 40 for generating a prediction model of a gene regulatory network provided by this embodiment includes: a first acquisition module 401, a first output module 402, a second acquisition module 403, a second output module 404, an alignment module 405, an optimization module 406, and a training module 407.

[0197] The first acquisition module 401 is used to acquire gene expression time series.

[0198] The first output module 402 is configured to input the gene expression time series into a gene expression matrix encoder to output global gene expression features.

[0199] The second acquisition module 403 is configured to acquire biological text data.

[0200] The second output module 404 is configured to input the biological text data into a text encoder to output global text features.

[0201] The alignment module 405 is configured to align the global gene expression features with the global text features to obtain a contrastive loss.

[0202] The optimization module 406 is configured to optimize the prediction model according to the contrastive loss to generate an optimized prediction model.

[0203] The training module 407 is configured to train the optimized prediction model through a cosine annealing learning rate scheduling mechanism to generate a trained prediction model.

[0204] In a possible implementation manner, the first output module 402 includes:

[0205] The first acquisition unit is configured to acquire the gene expression matrix in the gene expression time series.

[0206] The first output unit is configured to input the gene expression matrix into a feature extraction model to output an embedding vector.

[0207] The first calculation unit is configured to calculate the weight of each time in the gene expression time series according to the embedding vector through an attention mechanism.

[0208] The second calculation unit is configured to calculate and generate global gene expression features according to the embedding vector and the weight of each time in the gene expression time series.

[0209] In a possible implementation manner, the formula for calculating the weight of each time in the gene expression time series according to the embedding vector by the first calculation unit through the attention mechanism is:

[0210]

[0211] In the formula, represents the weight; represents the attention score function; represents the feature at the th time point; represents the query vector; represents the th time point; represents the number of time points.

[0212] In a possible implementation, the formula for calculating and generating the global gene expression feature in the second calculation unit based on the embedding vector and the weight at each time in the gene expression time series is as follows:

[0213]

[0214] In the formula, represents the global gene expression feature; represents the weight; represents the time point feature; represents the number of time points.

[0215] In a possible implementation, the second output module 404 includes:

[0216] A conversion unit for converting biological text data into a text token sequence.

[0217] A text processing unit for performing text processing on the text token sequence through a word segmentation algorithm to obtain the segmented text.

[0218] A second output unit for inputting the segmented text into a language model to output a hidden state sequence.

[0219] An extraction unit for extracting hidden states from the hidden state sequence to output the global text feature.

[0220] In a possible implementation, the text processing unit includes:

[0221] A creation subunit for creating an initial vocabulary.

[0222] A recording subunit for recording the co-occurrence frequency of characters or subwords in the corpus and merging the subwords with the highest co-occurrence frequency according to the set number of iterations to generate an updated vocabulary.

[0223] A text processing subunit for performing text processing on the text token sequence according to the updated vocabulary to obtain the segmented text.

[0224] In a possible implementation, the formula for inputting the segmented text into a language model in the second output unit to output a hidden state sequence is as follows:

[0225]

[0226] In the formula, represents the hidden state sequence; represents the segmented text sequence; represents a real matrix space of dimension; represents the sequence length; represents the embedding dimension; BERT() represents processing by the language model.

[0227] In a possible implementation, the second output module 404 further includes:

[0228] A parsing unit for parsing the static semantics of biological text data to generate parsed biological text data.

[0229] An identification unit for identifying biological process texts in the parsed biological text data.

[0230] A calibration unit for performing time calibration on the biological process texts to generate calibrated biological text data.

[0231] In a possible implementation, the alignment module 405 includes:

[0232] A third calculation unit for calculating the similarity of feature vectors based on the global gene expression features and the global text features.

[0233] A fourth calculation unit for calculating the contrast loss between each gene expression feature in the global gene expression features and the text features based on the similarity of the feature vectors, to obtain a loss function from gene expression to text.

[0234] A fifth calculation unit for calculating the contrast loss between each text feature in the global text features and the gene expression features based on the similarity of the feature vectors, to obtain a loss function from text to gene expression.

[0235] A sixth calculation unit for calculating and generating a contrast loss based on the loss function from gene expression to text and the loss function from text to gene expression.

[0236] In a possible implementation, the formula for the sixth calculation unit to calculate and generate a contrast loss based on the loss function from gene expression to text and the loss function from text to gene expression is:

[0237]

[0238] In the formula, represents the contrast loss; represents the loss function from gene expression to text; represents the loss function from text to gene expression.

[0239] In a possible implementation, the training module 407 includes:

[0240] A second acquisition unit for acquiring the initial learning rate of the optimized prediction model.

[0241] A seventh calculation unit for calculating the learning rate after iteration according to the set number of iterations and the initial learning rate.

[0242] A training unit for training the optimized prediction model according to the learning rate after iteration to generate a trained prediction model.

[0243] In a possible implementation manner, the formula for calculating the learning rate after iteration in the seventh calculation unit according to the set number of iterations and the initial learning rate is:

[0244]

[0245] In the formula, represents the learning rate after iteration; represents the initial learning rate; represents the number of iterations; represents the total number of training steps.

[0246] For the description of the features in the corresponding embodiment of the prediction model generation device of the gene regulatory network, reference can be made to the relevant description in the corresponding embodiment of the prediction model generation method of the gene regulatory network, which will not be elaborated here one by one.

[0247] Figure 5 This is a schematic structural diagram of the server provided by this application. As Figure 5 shown, the server 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the server 50 further includes a communication component 503. Among them, the processor 501, the memory 502, and the communication component 503 are connected through a bus 504.

[0248] In a specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the above-mentioned embodiment of the prediction model generation method of the gene regulatory network.

[0249] For the specific implementation process of the processor 501, reference can be made to the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0250] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of the hardware and software modules in the processor.

[0251] The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.

[0252] The bus may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0253] The embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any of the embodiments of the above-mentioned method for generating a prediction model of a gene regulatory network when running.

[0254] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other media that can store computer programs.

[0255] The embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the embodiments of the above-mentioned method for generating a prediction model of a gene regulatory network are implemented.

[0256] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the method for generating a prediction model of a gene regulatory network are implemented.

[0257] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0258] The above has introduced in detail a method, apparatus, server, and medium for generating a prediction model of a gene regulatory network provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for generating a prediction model of a gene regulatory network, characterized in that, Including: Obtain gene expression time series; Input the gene expression time series into a gene expression matrix encoder to output global gene expression features; Obtain biological text data; Input the biological text data into a text encoder to output global text features; Align the global gene expression features with the global text features to obtain a contrast loss; Optimize a prediction model according to the contrast loss to generate an optimized prediction model; Train the optimized prediction model through a cosine decay learning rate scheduling mechanism to generate a trained prediction model; The aligning the global gene expression features with the global text features to obtain a contrast loss includes: Calculate the similarity of feature vectors according to the global gene expression features and the global text features; Calculate the contrast loss between each gene expression feature and the text feature in the global gene expression features according to the similarity of the feature vectors to obtain a loss function from gene expression to text; Calculate the contrast loss between each text feature and the gene expression feature in the global text features according to the similarity of the feature vectors to obtain a loss function from text to gene expression; Calculate and generate a contrast loss according to the loss function from gene expression to text and the loss function from text to gene expression; The training the optimized prediction model through a cosine decay learning rate scheduling mechanism to generate a trained prediction model includes: Obtain the initial learning rate of the optimized prediction model; Calculate the learning rate after iteration according to the set number of iterations and the initial learning rate; Train the optimized prediction model according to the learning rate after iteration to generate a trained prediction model.

2. The method for generating a prediction model of a gene regulatory network according to claim 1, wherein The inputting the gene expression time series into a gene expression matrix encoder to output global gene expression features includes: Obtain the gene expression matrix in the gene expression time series; Input the gene expression matrix into a feature extraction model to output an embedding vector; Calculate the weight of each time in the gene expression time series according to the embedding vector through an attention mechanism; Calculate and generate global gene expression features according to the embedding vector and the weight of each time in the gene expression time series.

3. The method for generating a prediction model of a gene regulatory network according to claim 2, wherein The formula for calculating the weight of each time in the gene expression time series according to the embedding vector through an attention mechanism is: In the formula, represents weight; represents the attention score function; Indicates Characteristics of each time point; represents the query vector; Indicates Characteristics of each time point; Indicates the number of time points.

4. The method for generating a prediction model of a gene regulatory network according to claim 2, wherein The formula for calculating and generating global gene expression features according to the embedding vector and the weight of each time in the gene expression time series is: In the formula, represents the global gene expression feature; represents the weight; represents the time point feature; represents the number of time points.

5. The method for generating a prediction model of a gene regulatory network according to claim 1, wherein, The inputting the biological text data into a text encoder to output global text features includes: Convert the biological text data into a text token sequence; Perform text processing on the text token sequence through a word segmentation algorithm to obtain segmented text; Input the segmented text into a language model to output a hidden state sequence; Extract hidden states from the hidden state sequence to output global text features.

6. The method for generating a prediction model of a gene regulatory network according to claim 5, wherein The performing text processing on the text token sequence through a word segmentation algorithm to obtain segmented text includes: Create an initial vocabulary; Record the co-occurrence frequencies of characters or sub-words in the corpus, and merge the sub-words with the highest co-occurrence frequencies according to the set number of iterations to generate an updated vocabulary list; Perform text processing on the text token sequence according to the updated vocabulary list to obtain the segmented text.

7. The method for generating a prediction model of a gene regulatory network according to claim 5, wherein The formula for inputting the segmented text into the language model to output the hidden state sequence is: Wherein, represents the hidden state sequence; represents the text sequence after word segmentation; represents the real matrix space of dimension; represents the sequence length; represents the embedding dimension; BERT() represents processing by the language model.

8. The method for generating a prediction model of a gene regulatory network according to claim 5, wherein, Before converting the biological text data into a text token sequence, it further includes: Parse the static semantics of the biological text data to generate the parsed biological text data; Identify the biological process text in the parsed biological text data; Perform time calibration on the biological process text to generate the calibrated biological text data.

9. The method for generating a prediction model of a gene regulatory network according to claim 1, wherein The formula for calculating the contrastive loss according to the loss function from gene expression to text and the loss function from text to gene expression is: In the formula, represents the contrastive loss; represents the loss function from gene expression to text; represents the loss function from text to gene expression.

10. The method for generating a prediction model of a gene regulatory network according to claim 1, wherein The formula for calculating the learning rate after iteration according to the set number of iterations and the initial learning rate is: In the formula, represents the learning rate after iteration; represents the initial learning rate; represents the number of iterations; represents the total number of training steps.

11. A server, characterized in that, It includes: A memory for storing computer programs; A processor for implementing the steps of the method for generating a prediction model of a gene regulatory network as described in any one of claims 1 to 10 when executing the computer program.

12. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by the processor, implements the steps of the method for generating a prediction model of a gene regulatory network as described in any one of claims 1 to 10.

13. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method for generating a prediction model of a gene regulatory network as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method for deducing gene regulatory network

    CN117831632A

  • Prediction model training method based on transfer learning and program product

    CN119026646A