Multi-source gene expression processing method and related products
Patent Information
- Application Number
- CN202511938257.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-12-22
AI Technical Summary
[0003]然而,在实际研究中, BULK基因表达数据常来自不同测序平台,由于不同测序平台在技术原理、探针设计和检测方法上的差异会引入显著的批次效应,从而会导致相同生物学状态的样本呈现非真实的表达差异,干扰生物学信号的识别,降低结果的可重复性和跨研究可比性
[0015]本公开的实施例提供的多源基因表达量处理方法及相关产品,获取患者样本多源基因表达量集和健康样本多源基因表达量集,其中,患者样本多源基因表达量集包括至少一个患者样本中每个第一基因的第一基因表达量,健康样本多源基因表达量集包括至少一个健康样本中每个第一基因的第二基因表达量;根据患者样本多源基因表达量集和健康样本多源基因表达量集中的至少一个多源基因表达量集,确定各患者样本中每个第一基因的目标基因表达量;获取基因数据集,其中,基因数据集包括至少一个第二基因以及各第二基因之间的关联关系,各第二基因包含于各第一基因之中;根据基因数据集生成与每个第二基因对应的目标基因向量;根据各目标基因向量和各目标基因表达量,确定各患者样本的基因向量序列。本公开首先基于患者样本多源基因表达量集和健康样本多源基因表达量集中的至少一个多源基因表达量集,可以去除各患者样本中每个第一基因的第一基因表达量的噪声以获得目标基因表达量,接着,可以为基因数据集中的每个第二基因生成对应的目标基因向量,由于各第二基因属于各第一基因,因此,可以通过各第二基因的目标基因向量替代相应第一基因的目标基因表达量,获得患者样本的基因向量序列,该基因向量序列可以减少跨测序平台差异引入的批次效应对基因表征的影响,进而显著提升跨平台基因表达数据的测量准确性与可比性,以及,基因向量序列可以更适用于机器学习和深度学习模型(如自然语言处理模型)进行疾病分型、预后预测或生物信息挖掘,有助于提高模型的鲁棒性、泛化能力和预测精度,为精准医学研究提供更可靠的数据基础。
Smart Images

Figure CN121366634B_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of bioinformatics technology, specifically to a method for processing multi-source gene expression levels and related products. Background Technology
[0002] Gene expression analysis is a core technology in biomedical research, widely used in disease mechanism research and drug target screening. Bulk gene expression analysis, in particular, reflects the overall expression level of genes under specific physiological or pathological conditions by sequencing the total RNA (ribonucleic acid) of an entire tissue or cell population. Due to its mature technology, low cost, and standardized analytical procedures, it has become an important tool for large-scale cohort studies and clinical sample analysis.
[0003] However, in actual research, BULK gene expression data often come from different sequencing platforms. Due to the differences in technical principles, probe design and detection methods between different sequencing platforms, significant batch effects can be introduced, which can lead to non-true expression differences in samples with the same biological state, interfering with the identification of biological signals and reducing the reproducibility and cross-study comparability of the results. Summary of the Invention
[0004] The embodiments of this disclosure propose a method and related products for processing multi-source gene expression levels. Specifically, based on at least one multi-source gene expression set from patient samples and healthy samples, noise in the first gene expression level of each first gene in each patient sample is removed to obtain the target gene expression level. Then, a corresponding target gene vector can be generated for each second gene in the gene dataset. Since each second gene belongs to each first gene, the target gene expression level of the corresponding first gene can be replaced by the target gene vector of each second gene to obtain the gene vector sequence of the patient sample. This gene vector sequence can reduce the impact of batch effects introduced by cross-sequencing platform differences on gene characterization, thereby improving the comparability and measurement accuracy of cross-platform gene expression data.
[0005] In a first aspect, this disclosure provides a method for processing multi-source gene expression levels, the method comprising: Obtain a multi-source gene expression set of patient samples and a multi-source gene expression set of healthy samples, wherein the multi-source gene expression set of patient samples includes the first gene expression level of each first gene in at least one patient sample, and the multi-source gene expression set of healthy samples includes the second gene expression level of each first gene in at least one healthy sample; Based on at least one multi-source gene expression set in the multi-source gene expression set of the patient samples and the multi-source gene expression set of the healthy samples, determine the target gene expression level of each of the first genes in each of the patient samples. Obtain a gene dataset, wherein the gene dataset includes at least one second gene and the association between each second gene, and each second gene is contained within each first gene; Generate a target gene vector corresponding to each of the second genes based on the gene dataset; The gene vector sequence of each patient sample is determined based on the target gene vector and the expression level of each target gene.
[0006] In some optional implementations, determining the target gene expression level of each of the first genes in each of the patient samples based on at least one multi-source gene expression set from the patient sample multi-source gene expression set and the healthy sample multi-source gene expression set includes: Based on the multi-source gene expression set of the patient samples, determine the target gene expression level of each of the first genes in each of the patient samples; or, based on the multi-source gene expression set of the patient samples and the multi-source gene expression set of the healthy samples, determine the target gene expression level of each of the first genes in each of the patient samples.
[0007] In some optional implementations, determining the expression level of the target gene for each of the first genes in each of the patient samples based on the multi-source gene expression set of the patient samples includes: For each of the first genes, the average expression level of the first gene in each of the patient samples is determined based on the first gene expression level of the first gene in the multi-source gene expression set of the patient samples; and, For each patient sample, calculate the first gene expression deviation between the first gene expression level in the patient sample and the average gene expression level in the patient sample; and, The deviation of the first gene expression is determined as the expression level of the target gene of the first gene in the patient sample.
[0008] In some optional implementations, determining the target gene expression level of each of the first genes in each of the patient samples based on the multi-source gene expression sets of the patient samples and the multi-source gene expression sets of the healthy samples includes: For each of the first genes, the average expression level of the first gene in healthy samples is determined based on the expression level of the second gene of the first gene in each of the healthy samples in the multi-source gene expression set of the healthy samples; and, For each patient sample, calculate the deviation of the expression level of the first gene in the patient sample from the average gene expression level in the healthy samples; and, The second gene expression deviation is determined as the target gene expression level of the first gene in the patient sample.
[0009] In some optional implementations, generating a target gene vector corresponding to each of the second genes based on the gene dataset includes: Gene knowledge graph is generated based on the gene dataset, wherein each node in the gene knowledge graph represents a second gene, and each edge represents the association between two second genes; Based on the preset gene vector conversion method, an initial gene vector corresponding to each of the second genes is generated; For each second gene, based on the gene knowledge graph, at least one neighboring second gene corresponding to the second gene is determined, wherein each neighboring second gene is included within each second gene; and, Based on the initial gene vector of the second gene and the initial gene vectors of each of the neighboring second genes, the target gene vector corresponding to the second gene is generated.
[0010] In some optional implementations, determining the gene vector sequence of each patient sample based on each target gene vector and each target gene expression level includes: For each patient sample, the expression levels of the target gene for each of the first genes in the patient sample are sorted in descending order to generate the first gene expression level sequence of the patient sample; and, For each of the first genes in the patient sample, determine whether there is a target second gene that matches the first gene in each of the second genes; If it is determined that it exists, the target gene expression level of the first gene in the first gene expression level sequence is replaced with the target gene vector of the target second gene, and the target gene vector of the target second gene is determined as the target gene vector of the first gene; The gene vector sequence of the patient sample is generated based on the target gene vector of each of the first genes that have been replaced in the first gene expression sequence.
[0011] Secondly, this disclosure provides a multi-source gene expression level processing device, the device comprising: A gene expression level acquisition unit is used to acquire a multi-source gene expression level set of patient samples and a multi-source gene expression level set of healthy samples, wherein the multi-source gene expression level set of patient samples includes the first gene expression level of each first gene in at least one patient sample, and the multi-source gene expression level set of healthy samples includes the second gene expression level of each first gene in at least one healthy sample. A gene expression level determination unit is used to determine the target gene expression level of each of the first genes in each of the patient samples based on at least one multi-source gene expression level set in the multi-source gene expression level set of the patient samples and the multi-source gene expression level set of the healthy samples. A gene data acquisition unit is used to acquire a gene dataset, wherein the gene dataset includes at least one second gene and the association between each second gene, and each second gene is contained within each first gene; A gene vector generation unit is used to generate a target gene vector corresponding to each of the second genes based on the gene dataset. A gene vector sequence determination unit is used to determine the gene vector sequence of each patient sample based on each target gene vector and each target gene expression level.
[0012] Thirdly, this disclosure provides an electronic device, including: One or more processors; Storage device, on which one or more programs are stored, When the above-described one or more programs are executed by the above-described one or more processors, the above-described one or more processors implement the method as described in any embodiment of the first aspect of this disclosure.
[0013] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method described in any embodiment of the first aspect of this disclosure.
[0014] Fifthly, this disclosure provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the method described in any embodiment of the first aspect of this disclosure.
[0015] The multi-source gene expression level processing method and related products provided in the embodiments of this disclosure obtain a multi-source gene expression level set of patient samples and a multi-source gene expression level set of healthy samples. The patient sample multi-source gene expression level set includes the first gene expression level of each first gene in at least one patient sample, and the healthy sample multi-source gene expression level set includes the second gene expression level of each first gene in at least one healthy sample. Based on at least one multi-source gene expression level set from the patient sample and healthy sample multi-source gene expression level sets, the target gene expression level of each first gene in each patient sample is determined. A gene dataset is obtained, wherein the gene dataset includes at least one second gene and the association relationships between each second gene, and each second gene is contained within each first gene. A target gene vector corresponding to each second gene is generated based on the gene dataset. The gene vector sequence of each patient sample is determined based on each target gene vector and each target gene expression level. This disclosure firstly uses at least one multi-source gene expression set from a multi-source gene expression set of patient samples and a multi-source gene expression set of healthy samples to remove noise from the first gene expression level of each first gene in each patient sample to obtain the target gene expression level. Then, a corresponding target gene vector can be generated for each second gene in the gene dataset. Since each second gene belongs to each first gene, the target gene expression level of the corresponding first gene can be replaced by the target gene vector of each second gene to obtain the gene vector sequence of the patient sample. This gene vector sequence can reduce the impact of batch effects introduced by cross-sequencing platform differences on gene characterization, thereby significantly improving the measurement accuracy and comparability of cross-platform gene expression data. Furthermore, the gene vector sequence is more suitable for machine learning and deep learning models (such as natural language processing models) for disease classification, prognosis prediction, or bioinformatics mining, which helps to improve the robustness, generalization ability, and prediction accuracy of the model, and provides a more reliable data foundation for precision medicine research. Attached Figure Description
[0016] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are for illustrative purposes only and are not intended to limit the invention. In the drawings: Figure 1 This is a system architecture diagram in which an embodiment of the multi-source gene expression level processing method disclosed herein can be applied; Figure 2 This is a flowchart of an embodiment of the multi-source gene expression level processing method according to the present disclosure; Figure 3 This is a schematic diagram of a structure of an embodiment of the multi-source gene expression level processing device according to the present disclosure; Figure 4This is a schematic diagram of the structure of a computer system suitable for implementing embodiments of the present disclosure. Detailed Implementation
[0017] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0019] Figure 1 An exemplary system architecture 100 is shown, in which embodiments of the multi-source gene expression level processing methods, apparatuses, electronic devices and storage media of this disclosure can be applied.
[0020] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, network 104, and server 105. Network 104 is used to provide communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various communication connection types, such as wired communication links, wireless communication links, etc.
[0021] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as multi-source gene expression processing applications, voice interaction applications, video conferencing applications, short video social applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0022] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with microphones and speakers, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), portable computers, and desktop computers, etc. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as multiple software programs or software modules (e.g., acquiring multi-source gene expression sets from patient samples and multi-source gene expression sets from healthy samples), or as a single software program or software module. No specific limitations are imposed here.
[0023] Server 105 can be a server that provides various services, such as a backend server that processes the multi-source gene expression sets of patient samples and healthy samples obtained from terminal devices 101, 102, and 103. The backend server can perform corresponding processing based on the multi-source gene expression sets of patient samples and healthy samples obtained from the terminal devices.
[0024] In some cases, the multi-source gene expression level processing and generation method provided in this disclosure can be jointly executed by terminal devices 101, 102, 103 and server 105. For example, the step of "obtaining the multi-source gene expression level set of patient samples and the multi-source gene expression level set of healthy samples" can be executed by terminal devices 101, 102, 103, and the step of "determining the target gene expression level of each first gene in each patient sample based on at least one multi-source gene expression level set from the patient sample multi-source gene expression level set and the healthy sample multi-source gene expression level set" can be executed by server 105. This disclosure does not limit this. Correspondingly, the multi-source gene expression level processing device can also be respectively set in terminal devices 101, 102, 103 and server 105.
[0025] In some cases, the multi-source gene expression level processing method provided in this disclosure can be executed by server 105. Accordingly, the multi-source gene expression level processing device can also be set in server 105. In this case, the system architecture 100 may not include terminal devices 101, 102, and 103.
[0026] In some cases, the multi-source gene expression level processing method provided in this disclosure can be executed by terminal devices 101, 102, and 103. Correspondingly, the multi-source gene expression level processing device can also be set in terminal devices 101, 102, and 103. In this case, the system architecture 100 may not include server 105.
[0027] It should be noted that server 105 can be either hardware or software. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules (for example, used to provide distributed services), or as a single software program or software module. No specific limitations are made here.
[0028] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0029] All information, data, and signals disclosed herein are authorized by the user or by all parties, and the collection, use, and processing of such data comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0030] Continue to refer to Figure 2 , Figure 2 A flowchart illustrating an embodiment of the multi-source gene expression level processing and generation method according to this disclosure is shown. Figure 2 The multi-source gene expression level processing method shown can be applied to Figure 1 The terminal device or server shown. This method for generating multi-source gene expression levels includes at least the following steps 201-205.
[0031] Step 201: Obtain the multi-source gene expression sets of patient samples and the multi-source gene expression sets of healthy samples.
[0032] In this embodiment, the multi-source gene expression set of patient samples includes the first gene expression level of each first gene in at least one patient sample, and the multi-source gene expression set of healthy samples includes the second gene expression level of each first gene in at least one healthy sample.
[0033] Here, the first gene can refer to at least one gene of common interest in both patient and healthy samples whose gene expression levels are detected and quantified in a Bulk transcriptome sequencing experiment. The expression level of the first gene can refer to the gene expression level of the first gene in the patient sample, and the expression level of the second gene can refer to the gene expression level of the first gene in the healthy sample. Due to the different biological states of patient and healthy samples, there may be significant differences in the expression level of the first gene in patient samples and the expression level of the second gene in healthy samples. This difference can serve as an important basis for identifying disease-related genes. The values of the first and second gene expression levels are derived from Bulk transcriptome sequencing data. The multi-source gene expression sets of patient samples or healthy samples are usually stored in the form of a two-dimensional matrix. For example, each row in the two-dimensional matrix can correspond to a first gene, and each column can correspond to a patient sample or a healthy sample, or each row in the two-dimensional matrix can correspond to a patient sample or a healthy sample, and each column can correspond to a first gene. There are no restrictions here.
[0034] For example, the multi-source gene expression set of a patient sample can be represented as ,in, It can represent a set of multi-source gene expression levels in a patient sample (e.g., a multi-source gene expression matrix in a patient), where m can represent the number of the first gene and p can represent the number of patient samples.
[0035] The multi-source gene expression set of healthy samples can be represented as Where H can represent the set of multi-source gene expression levels of healthy samples (e.g., the multi-source gene expression matrix of healthy individuals), where m represents the number of the first gene, and n can represent the number of healthy samples.
[0036] Step 202: Determine the target gene expression level of each first gene in each patient sample based on at least one multi-source gene expression set from the multi-source gene expression set of the patient sample and the multi-source gene expression set of the healthy sample.
[0037] In this embodiment, the original first gene expression level may be subject to data noise due to factors such as technical noise, batch effect, or individual outliers. Therefore, this step can use at least one multi-source gene expression set from the patient sample multi-source gene expression set and the healthy sample multi-source gene expression set to denoise the original patient multi-source gene expression level, thereby obtaining a more stable and reliable target gene expression level, improving the accuracy of gene expression, and providing a high-quality data foundation for subsequent analysis.
[0038] In some alternative implementations, the target gene expression level of each first gene in each patient sample can be determined based on the multi-source gene expression level set of patient samples, or the target gene expression level of each first gene in each patient sample can be determined based on the multi-source gene expression level set of patient samples and the multi-source gene expression level set of healthy samples.
[0039] In this embodiment, the target gene expression level of each first gene in each patient sample can be determined solely based on the multi-source gene expression level set of the patient sample, or the target gene expression level of each first gene in each patient sample can be determined jointly based on the multi-source gene expression level set of the patient sample and the multi-source gene expression level set of the healthy sample.
[0040] In some alternative implementations, the target gene expression level of each first gene in each patient sample is determined solely based on the multi-source gene expression level set of the patient sample, which may include the following A1-A3.
[0041] A1. For each first gene, determine the average expression level of the first gene in each patient sample based on the expression level of the first gene in the multi-source gene expression level set of the patient samples.
[0042] In this embodiment, the average expression level of the first gene in patient samples can refer to the average gene expression level of the first gene in each patient sample, which can be used to measure the overall expression level of the first gene in the patient sample population.
[0043] For each first gene, the average expression level of that first gene in patient samples can be determined using the following formula:
[0044] Where i represents the i-th first gene, i=1,2,...,m, and m is the total number of first genes; j represents the j-th patient sample, j=1,2,...,p, and p is the total number of patient samples; This can represent the expression level of the first gene of the i-th patient in the j-th patient sample. This represents the average gene expression level of the i-th first gene in patient samples.
[0045] A2, for each patient sample, calculate the deviation of the expression level of the first gene in the patient sample from the average expression level of the first gene in the patient sample. After obtaining the average gene expression level of the i-th first gene in the patient sample, for each patient sample, the deviation of the first gene expression level of the i-th first gene in the patient sample from the average gene expression level of the patient sample can be calculated.
[0046] The deviation of the first gene expression can be used to measure the magnitude and direction (upregulation or downregulation) of the deviation of the first gene expression level in a patient sample from the average gene expression level of the entire patient sample population. The larger the deviation of the first gene expression, the more abnormal the expression of the first gene in the patient sample is, and the smaller the deviation of the first gene expression, the more normal the expression of the first gene in the patient sample is.
[0047] Specifically, the deviation in first gene expression can be calculated using the following formula:
[0048] in, This can represent the deviation of the expression of the first gene of the i-th gene in the j-th patient sample. This can represent the original expression level of the first gene of the i-th gene in the j-th patient sample. This represents the average gene expression level of the i-th first gene in patient samples.
[0049] A3 determines the deviation of the first gene expression as the target gene expression level of the first gene in the patient sample.
[0050] After obtaining the first gene expression deviation, the first gene expression deviation can be determined as the target gene expression level of the first gene in the patient sample. Through A1-A3 above, the original first gene expression level of each patient sample can be converted into a relative deviation value relative to the average gene expression level of the patient sample population.
[0051] In samples of patients with the same disease, certain primary genes may be significantly overexpressed or underexpressed only in some patient samples. By analyzing the deviation of primary gene expression (relative deviation value), key genes with heterogeneous expression can be effectively identified. This can help in the mining of disease subtype or personalized gene information, and can also eliminate the batch effect caused by sequencing platform to a certain extent, thereby improving the comparability and accuracy of data.
[0052] In some optional implementations, the target gene expression level of each first gene in each patient sample is determined jointly based on the multi-source gene expression level set of patient samples and the multi-source gene expression level set of healthy samples, which may include the following B1-B3.
[0053] B1. For each first gene, determine the average expression level of the first gene in healthy samples based on the expression level of the second gene of the first gene in each healthy sample from the multi-source gene expression level set of healthy samples.
[0054] In this embodiment, the average gene expression level of the first gene in healthy samples can refer to the average gene expression level of the second gene of the first gene in each healthy sample, which can be used to measure the overall gene expression level of the first gene in the healthy sample population.
[0055] For each first gene, the average expression level of that first gene in healthy samples can be determined using the following formula:
[0056] Where i represents the i-th first gene, i = 1, 2, ..., m, and m is the total number of first genes; s represents the s-th healthy sample, s = 1, 2, ..., n, and n is the total number of healthy samples; This can represent the expression level of the second gene of the i-th first gene in the s-th healthy sample. This represents the average gene expression level of the i-th healthy sample with the first gene.
[0057] B2, for each patient sample, calculate the deviation of the expression level of the first gene in the patient sample from the average expression level of the second gene in the healthy sample.
[0058] After obtaining the average expression level of the first gene in healthy samples for the i-th first gene, for each patient sample, the deviation of the expression level of the first gene of the i-th first gene in the patient sample from the average expression level of the second gene in healthy samples can be calculated.
[0059] The deviation of the second gene expression can be used to measure the magnitude and direction of the deviation of the gene expression level of the first gene in the patient sample from the average gene expression level (normal range) of the entire healthy sample population. The larger the deviation of the second gene expression, the more abnormal the expression of the first gene in the patient sample is, and the smaller the deviation of the second gene expression, the more normal the expression of the first gene in the patient sample is.
[0060] Specifically, the deviation in second gene expression can be calculated using the following formula:
[0061] in, This can represent the deviation of the expression of the second gene of the i-th gene in the j-th patient sample. This can represent the original expression level of the first gene of the i-th gene in the j-th patient sample. This represents the average gene expression level of the i-th healthy sample with the first gene.
[0062] B3, the deviation of the second gene expression is determined as the target gene expression level of the first gene in the patient sample.
[0063] After obtaining the second gene expression deviation, the second gene expression deviation can be determined as the target gene expression level of the first gene in the patient sample. Through B1-B3 above, the original expression level of each first gene in each patient sample can be converted into a relative deviation value relative to the average gene expression level of healthy samples in the healthy sample group.
[0064] In this way, the expression level of the target gene can characterize the degree of deviation of each first gene in each patient sample from the normal state, making the expression level of the target gene more biologically and clinically interpretable, enhancing the identification of potential pathogenic genes, and also eliminating the batch effect caused by the sequencing platform to a certain extent, improving the comparability and accuracy of the data.
[0065] Step 203: Obtain the gene dataset.
[0066] In this embodiment, the gene dataset includes at least one second gene and the association between each second gene, with each second gene contained within each first gene.
[0067] In this embodiment, the data in the gene dataset can be gene pathway data from biological databases such as KEGG (Kyoto Encyclopedia of Genes and Genomes). Gene pathway data includes molecular interaction information in biological pathways. In the gene pathway data, nodes represent genes involved in the pathway (i.e., second genes), and edges represent the association between genes, such as upstream and downstream regulatory relationships or joint participation in the same signaling pathway.
[0068] Each second gene is contained within each first gene. In other words, all second genes in the gene dataset are subsets of the first genes detected in the multi-source gene expression set of patient samples. In other words, the second gene is a subset of the first genes that are screened or extracted from the first genes and have specific biological significance (such as participating in known pathways, being functionally related, or being associated with diseases).
[0069] Here, by introducing gene datasets, genes can be endowed with structured features containing biological context information, which helps to generate more functionally interpretable target gene vectors and improve the ability to identify disease-related pathways and key genes.
[0070] Step 204: Generate the target gene vector corresponding to each second gene based on the gene dataset.
[0071] In this embodiment, each second gene can correspond to a target gene vector, which can represent the vectorized representation of the functional characteristics and biological context information of the second gene.
[0072] In some alternative implementations, step 204 may include the following C1-C4.
[0073] C1 generates a gene knowledge graph based on a gene dataset.
[0074] In this embodiment, the gene knowledge graph can be composed of a node set V and an edge set E, where each node V represents a second gene, and each edge ( u , v )∈ E This represents the association (also known as functional or regulatory relationship) between two second genes. Specifically, each second gene in the gene dataset is treated as a node in the graph, and connecting edges are established based on their association relationships. Each edge is then assigned a specific name. u , v The relationships between them are recorded, and a structured gene knowledge graph G=(V,E) is finally constructed.
[0075] C2 generates an initial gene vector corresponding to each second gene according to a preset gene vector conversion method.
[0076] In this embodiment, the gene vector conversion method is used to convert the second gene into the corresponding initial gene vector.
[0077] Pre-defined gene vector transformation methods can include, for example, using a pre-trained BERT (Bidirectional Encoder Representations from Transformers) model to convert the second gene into its corresponding initial gene vector, or using graph neural networks or graph embedding algorithms to convert the second gene into its corresponding initial gene vector. The initial gene vector of the second gene can be represented, for example, as follows: .
[0078] C3: For each second gene, determine at least one neighboring second gene corresponding to the second gene based on the gene knowledge graph.
[0079] In this embodiment, each neighboring second gene is contained within each second gene. For each second gene, the neighboring genes can be represented in the gene knowledge graph as other second genes that are directly connected to the second gene through edges, that is, second genes that have a direct relationship or interaction with it.
[0080] Here, for each second gene, after obtaining the initial gene vector, a fixed number of K neighboring second genes can be randomly selected from all neighboring second genes of the second gene or selected according to rules based on the gene knowledge graph and a preset sampling method, wherein K does not exceed the total number of neighboring second genes of the second gene.
[0081] The preset sampling method can be, for example, random sampling, weighted sampling based on edge weights, or priority selection of highly important genes involved in key pathways, etc., and there are no restrictions here.
[0082] C4 generates the target gene vector corresponding to the second gene based on the initial gene vector of the second gene and the initial gene vectors of each neighboring second gene.
[0083] After obtaining the initial gene vector and at least one neighboring second gene as described above, the target gene vector corresponding to the second gene can be generated based on the initial gene vector of the second gene and the initial gene vectors of each neighboring second gene.
[0084] In some alternative implementations, the features of the second gene and at least one neighboring second gene can be aggregated using an aggregation method to capture its local structural and functional context information in the gene knowledge graph and obtain the target gene vector corresponding to the second gene.
[0085] For example, common aggregation methods include mean aggregation, LSTM (Long Short-Term Memory) aggregation, pooling aggregation, etc., without limitation, and mean aggregation will be used as an example:
[0086] in, This can represent the mean aggregation function used in the k-th layer, which is used to average the embedding vectors of the second gene itself and its neighboring second genes. This can represent the total number of neighboring second genes of a given second gene at layer k in a gene knowledge graph. It can represent the embedding vector of neighbor node u at the (k-1)th layer. This can represent the embedding vector of the second gene itself in the (k-1)th layer.
[0087] Subsequently, through learnable linear transformations and nonlinear activations, the target gene vector of the second gene at layer k is obtained: σ(W* ) in, σ can represent the target gene vector of the second gene in the k-th layer, where k≥1 and can be determined according to actual needs without restriction. σ can represent a non-linear activation function, such as ReLU (linear rectified function), Sigmoid (logistic function), or Tanh (hyperbolic tangent function), used to introduce non-linear transformations and enhance the expressive power of the function. W can represent a learnable weight matrix, used to perform linear transformations on the aggregated features.
[0088] It should be noted that when k=1, the aggregation operation is applied to the initial gene vector. When selecting the neighboring second genes of the second gene, the direct neighbors of the second gene in the gene knowledge graph can be selected for aggregation. When k>1, the neighboring second genes of each layer will expand layer by layer. The neighboring second genes of the kth layer contain the neighboring nodes of the second gene in the k-order range of the gene knowledge graph, that is, the neighboring second genes of the neighboring second genes, ..., the neighboring second genes of the neighboring second genes, thereby gradually capturing higher-order structural and functional association information.
[0089] Therefore, as the number of layers increases, the generated target gene vectors can incorporate a wider range of contextual information, enabling gene function semantic modeling from local to global, which can be used for subsequent tasks such as gene function prediction and disease association analysis.
[0090] In some alternative implementations, the target gene vector corresponding to the second gene can also be generated based on methods such as word2Vec (Word to Vector), GCN (Graph Convolutional Network), and GAT (Graph Attention Network).
[0091] Taking the word2Vec method as an example, each gene pathway in biological databases such as KEGG can be regarded as a "sentence", and each second gene in the gene pathway can be regarded as a "word".
[0092] For example, the ordered arrangement of the second gene in a certain gene pathway can be represented as: ,in It can represent the t-th second gene.
[0093] A training corpus can be constructed based on gene pathway information in databases such as KEGG, and unsupervised training can be performed using the Word2Vec algorithm to learn the vector representation of each second gene in the context co-occurrence pattern, that is, to learn the distributed representation of genes.
[0094] The word2vec algorithm mainly includes two main models: CBOW (Continuous Bag of Words) and Skip-Gram.
[0095] This section uses the Skip-Gram model as an example. In the Skip-Gram model, the goal is to pass the current second gene... Predict its context-second gene within a fixed context window .
[0096] For the second gene Predict its contextual second gene The conditional probability can be:
[0097] in, This can be represented in genes Under the conditions of occurrence, its contextual second gene is The predicted probability, for The output vector, It can represent the context of the second gene. The transpose of the output vector. For the second gene The input vector, It can represent the transpose of the output vector of any second gene g. The total number of the second gene in the training corpus.
[0098] The training objective of the model is to maximize the joint likelihood probability of all second genes in the training corpus and their context second genes, i.e., to optimize the following loss function:
[0099] in, The context window size, c, is greater than or equal to 1; for example, c can be 2. These are the model parameters (including the second gene vector and the context second gene vector). It can represent the x-th second gene in the gene sequence, and q represents the relative to The position offset, when j>0. It means yes The second gene on the right, when j < 0, express The second gene on the left, c ≤ q ≤ c The size of the context window is determined to be 2c+1, and j≠0 indicates that... It is not part of the contextual second gene.
[0100] Here, when the loss function When the value of the second gene converges to a stable value, the training process can be considered complete, and the model training is finished. input vector Its functional semantics and biological associations within the context of a second gene have been fully captured, and it can be used as the second gene. The corresponding target gene vector.
[0101] Step 205: Determine the gene vector sequence of each patient sample based on the target gene vector and the expression level of each target gene.
[0102] After obtaining the target gene vectors of each second gene, this step can determine the gene vector sequence of each patient sample based on the target gene vectors of each second gene and the expression levels of the target genes of each first gene in each patient sample.
[0103] In some alternative implementations, step 205 may include the following D1-D4.
[0104] D1: For each patient sample, the expression levels of the target genes of each primary gene in the patient sample are sorted in descending order to generate the expression level sequence of the primary gene of the patient sample.
[0105] In this embodiment, for each patient sample, the expression levels of the target genes of each first gene in the patient sample can be sorted in descending order to construct the first gene expression level sequence corresponding to the patient sample. The first gene expression level sequence describes the molecular expression characteristics of the patient sample in the relative order of gene expression levels, which can effectively reflect the transcriptomic differences between individuals and has good individual discrimination ability.
[0106] D2, for each first gene in the patient sample, determines whether there is a target second gene that matches the first gene in each second gene.
[0107] In this embodiment, for each first gene in the patient sample, each second gene in the gene dataset can be traversed to determine whether there is a target second gene that matches the first gene in each second gene. For example, whether there is a second gene with the same gene name or gene symbol as the first gene. If there is, it is determined that there is a target second gene that matches the first gene in each second gene. If not, it is determined that there is no target second gene that matches the first gene in each second gene.
[0108] D3, if it is determined to exist, replace the target gene expression level of the first gene in the first gene expression sequence with the target gene vector of the second gene, and determine the target gene vector of the second gene as the target gene vector of the first gene.
[0109] In this embodiment, when there is a target second gene that matches the first gene in each second gene, the target gene expression level of the first gene in the first gene expression level sequence can be replaced with the target gene vector of the target second gene, and the target gene vector of the target second gene can be determined as the target gene vector of the first gene. In this way, the conversion of the target gene expression level to the target gene vector with semantic function can be realized.
[0110] D4. Gene vector sequence of patient sample is generated based on the target gene vector of each replaced first gene in the first gene expression sequence.
[0111] In this embodiment, the gene vector sequence of the patient sample can be constructed by arranging the target gene vectors corresponding to each mapped first gene in the original sorting order based on the first gene expression sequence that has completed vector replacement. Alternatively, the target gene expression of the first gene that has not completed vector replacement can be directly deleted to construct the gene vector sequence of the patient sample. This gene vector sequence not only retains the gene expression level of key genes, but also integrates the distributed representation of gene functional semantics, forming a multimodal input representation that combines individual expression characteristics and biological knowledge. It can be used for subsequent deep learning models to realize tasks such as disease identification, disease classification, subtype identification, and survival analysis in precision medicine.
[0112] In some alternative implementations, since the first z (z ≥ 1) first genes in the first gene expression sequence often play an important regulatory or driving role in patient samples and have strong molecular characterization capabilities, they can be regarded as potential key genes for preliminary identification of core molecular features related to the disease.
[0113] The multi-source gene expression level processing method provided in the embodiments of this disclosure includes: obtaining a multi-source gene expression level set for patient samples and a multi-source gene expression level set for healthy samples; wherein the multi-source gene expression level set for patient samples includes the first gene expression level of each first gene in at least one patient sample, and the multi-source gene expression level set for healthy samples includes the second gene expression level of each first gene in at least one healthy sample; determining the target gene expression level of each first gene in each patient sample based on at least one multi-source gene expression level set in both the patient sample and healthy sample multi-source gene expression level sets; obtaining a gene dataset, wherein the gene dataset includes at least one second gene and the association relationships between each second gene, and each second gene is contained within each first gene; generating a target gene vector corresponding to each second gene based on the gene dataset; and determining the gene vector sequence of each patient sample based on each target gene vector and each target gene expression level. This disclosure firstly uses at least one multi-source gene expression set from a multi-source gene expression set of patient samples and a multi-source gene expression set of healthy samples to remove noise from the first gene expression level of each first gene in each patient sample to obtain the target gene expression level. Then, a corresponding target gene vector can be generated for each second gene in the gene dataset. Since the second gene is included in the first gene of the patient sample, the target gene expression level of the corresponding first gene can be replaced by the target gene vector of each second gene to determine the gene vector sequence of the patient sample. This can effectively eliminate the batch effect caused by cross-sequencing platforms, significantly improve the accuracy and comparability of gene expression data, and the gene vector sequence is more suitable for machine learning and deep learning models (such as natural language processing models) for disease classification, prognosis prediction or bioinformatics mining, which helps to improve the robustness, generalization ability and prediction accuracy of the model, and provides a more reliable data foundation for precision medicine research.
[0114] Further reference Figure 3 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a multi-source gene expression level processing device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various terminal devices.
[0115] like Figure 3As shown, the multi-source gene expression level processing device 300 of this embodiment includes: a gene expression level acquisition unit 301, a gene expression level determination unit 302, a gene data acquisition unit 303, a gene vector generation unit 304, and a gene vector sequence determination unit 305. The system includes the following components: a gene expression level acquisition unit 301, used to acquire a multi-source gene expression level set for patient samples and a multi-source gene expression level set for healthy samples; the patient sample multi-source gene expression level set includes the first gene expression level of each first gene in at least one patient sample, and the healthy sample multi-source gene expression level set includes the second gene expression level of each first gene in at least one healthy sample; a gene expression level determination unit 302, used to determine the target gene expression level of each first gene in each patient sample based on at least one multi-source gene expression level set from the patient sample multi-source gene expression level set and the healthy sample multi-source gene expression level set; a gene data acquisition unit 303, used to acquire a gene dataset, wherein the gene dataset includes at least one second gene and the association relationships between the second genes, and each second gene is contained within each first gene; a gene vector generation unit 304, used to generate a target gene vector corresponding to each second gene based on the gene dataset; and a gene vector sequence determination unit 305, used to determine the gene vector sequence of each patient sample based on each target gene vector and each target gene expression level.
[0116] In this embodiment, the specific processing of the gene expression level acquisition unit 301, the gene expression level determination unit 302, the gene data acquisition unit 303, the gene vector generation unit 304, and the gene vector sequence determination unit 305, and the resulting technical effects, can be found in the following references: Figure 2 The relevant descriptions of steps 201 to 205 in the corresponding embodiments will not be repeated here.
[0117] In some alternative implementations, the gene expression level determination unit 302 may be further configured as follows: Based on the multi-source gene expression set of patient samples, determine the target gene expression level of each first gene in each patient sample; or, based on the multi-source gene expression set of patient samples and the multi-source gene expression set of healthy samples, determine the target gene expression level of each first gene in each patient sample.
[0118] In some alternative implementations, the gene expression level determination unit 302 may be further configured as follows: For each first gene, the average expression level of the first gene in each patient sample is determined based on the expression level of the first gene in the multi-source gene expression dataset of patient samples; and, For each patient sample, calculate the deviation of the expression level of the first gene in the patient sample from the average gene expression level in the patient sample; and, The deviation of the first gene expression was determined as the target gene expression level of the first gene in the patient sample.
[0119] In some alternative implementations, the gene expression level determination unit 302 may be further configured as follows: For each first gene, the average expression level of the first gene in healthy samples is determined based on the expression level of the second gene of the first gene in each healthy sample from the multi-source gene expression set; and, For each patient sample, calculate the deviation of the expression level of the first gene in the patient sample from the average expression level of the second gene in the healthy sample; and, The deviation of the second gene expression was determined as the target gene expression level of the first gene in the patient sample.
[0120] In some alternative implementations, the gene vector generation unit 304 may be further configured as follows: Gene knowledge graphs are generated based on gene datasets, where each node in the gene knowledge graph represents a second gene, and each edge represents the association between two second genes. Based on the preset gene vector conversion method, an initial gene vector corresponding to each second gene is generated; For each second gene, based on the gene knowledge graph, identify at least one neighboring second gene corresponding to the second gene, wherein each neighboring second gene is contained within each second gene; and, Based on the initial gene vector of the second gene and the initial gene vectors of each neighboring second gene, generate the target gene vector corresponding to the second gene.
[0121] In some alternative implementations, the gene vector sequence determination unit 305 may be further configured as follows: For each patient sample, the expression levels of the target gene for each primary gene in the patient sample are sorted in descending order to generate the primary gene expression sequence of the patient sample; and, For each first gene in the patient sample, determine whether there is a target second gene that matches the first gene in each second gene; If it is determined to exist, replace the target gene expression level of the first gene in the first gene expression sequence with the target gene vector of the second gene, and determine the target gene vector of the second gene as the target gene vector of the first gene; Gene vector sequences of patient samples are generated based on the target gene vectors of the first gene that has been replaced in the first gene expression sequence.
[0122] It should be noted that the implementation details and technical effects of each unit in the multi-source gene expression processing device provided in the embodiments of this disclosure can be referred to the descriptions of other embodiments in this disclosure, and will not be repeated here.
[0123] The following is for reference. Figure 4 It shows a schematic diagram of the structure of a computer system 400 suitable for implementing the terminal device of this disclosure. Figure 4 The computer system 400 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0124] like Figure 4 As shown, the computer system 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the computer system 400. The processing device 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0125] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows computer system 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 A computer system 400 with various electronic devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0126] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 409, or installed from a storage device 408, or installed from a ROM 402. When the computer program is executed by a processing device 401, it performs the functions defined in the methods of embodiments of this disclosure.
[0127] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0128] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0129] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the following functions: Figure 2 The embodiments shown and their alternative implementations illustrate a method for processing multi-source gene expression levels.
[0130] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and Python, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0132] The units described in the embodiments of this disclosure can be implemented in software or hardware. The name of a unit does not necessarily limit the unit itself; for example, a gene expression level acquisition unit can also be described as an "acquisition unit".
[0133] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
Claims
1. A method for processing multi-source gene expression levels, characterized in that, The method includes: Obtain a multi-source gene expression set of patient samples and a multi-source gene expression set of healthy samples, wherein the multi-source gene expression set of patient samples includes the first gene expression level of each first gene in at least one patient sample, and the multi-source gene expression set of healthy samples includes the second gene expression level of each first gene in at least one healthy sample; Based on at least one multi-source gene expression set in the multi-source gene expression set of the patient samples and the multi-source gene expression set of the healthy samples, determine the target gene expression level of each of the first genes in each of the patient samples. Obtain a gene dataset, wherein the gene dataset includes at least one second gene and the association between each second gene, and each second gene is contained within each first gene; Generate a target gene vector corresponding to each of the second genes based on the gene dataset; Based on the target gene vectors and the expression levels of the target genes, the gene vector sequences of each patient sample are determined. The step of determining the gene vector sequence of each patient sample based on each target gene vector and each target gene expression level includes: For each patient sample, the expression levels of the target gene for each of the first genes in the patient sample are sorted in descending order to generate the first gene expression level sequence of the patient sample; and, For each of the first genes in the patient sample, determine whether there is a target second gene that matches the first gene in each of the second genes; If it is determined that it exists, the target gene expression level of the first gene in the first gene expression level sequence is replaced with the target gene vector of the target second gene, and the target gene vector of the target second gene is determined as the target gene vector of the first gene; Based on the target gene vectors of each of the first genes that have been replaced in the first gene expression sequence, the gene vector sequence of the patient sample is generated. The gene vector sequence retains the gene expression level of key genes and integrates the distributed representation of gene functional semantics, forming a multimodal input representation that combines individual expression characteristics and biological knowledge.
2. The method according to claim 1, characterized in that, The step of determining the target gene expression level of each of the first genes in each of the patient samples based on at least one multi-source gene expression set from the patient sample multi-source gene expression set and the healthy sample multi-source gene expression set includes: Based on the multi-source gene expression set of the patient samples, determine the target gene expression level of each of the first genes in each of the patient samples; or, based on the multi-source gene expression set of the patient samples and the multi-source gene expression set of the healthy samples, determine the target gene expression level of each of the first genes in each of the patient samples.
3. The method according to claim 2, characterized in that, The step of determining the expression level of the target gene for each of the first genes in each of the patient samples based on the multi-source gene expression set of the patient samples includes: For each of the first genes, the average expression level of the first gene in each of the patient samples is determined based on the first gene expression level of the first gene in the multi-source gene expression set of the patient samples; and, For each patient sample, calculate the first gene expression deviation between the first gene expression level in the patient sample and the average gene expression level in the patient sample; and, The deviation of the first gene expression is determined as the expression level of the target gene of the first gene in the patient sample.
4. The method according to claim 2, characterized in that, The step of determining the target gene expression level of each of the first genes in each of the patient samples based on the multi-source gene expression level set of the patient samples and the multi-source gene expression level set of the healthy samples includes: For each of the first genes, the average expression level of the first gene in healthy samples is determined based on the expression level of the second gene of the first gene in each of the healthy samples in the multi-source gene expression set of the healthy samples; and, For each patient sample, calculate the deviation of the expression level of the first gene in the patient sample from the average gene expression level in the healthy samples; and, The second gene expression deviation is determined as the target gene expression level of the first gene in the patient sample.
5. The method according to claim 1, characterized in that, The step of generating a target gene vector corresponding to each of the second genes based on the gene dataset includes: Gene knowledge graph is generated based on the gene dataset, wherein each node in the gene knowledge graph represents a second gene, and each edge represents the association between two second genes; Based on the preset gene vector conversion method, an initial gene vector corresponding to each of the second genes is generated; For each second gene, based on the gene knowledge graph, at least one neighboring second gene corresponding to the second gene is determined, wherein each neighboring second gene is included within each second gene; and, Based on the initial gene vector of the second gene and the initial gene vectors of each of the neighboring second genes, the target gene vector corresponding to the second gene is generated.
6. A multi-source gene expression level processing device, characterized in that, The device includes: A gene expression level acquisition unit is used to acquire a multi-source gene expression level set of patient samples and a multi-source gene expression level set of healthy samples, wherein the multi-source gene expression level set of patient samples includes the first gene expression level of each first gene in at least one patient sample, and the multi-source gene expression level set of healthy samples includes the second gene expression level of each first gene in at least one healthy sample. A gene expression level determination unit is used to determine the target gene expression level of each of the first genes in each of the patient samples based on at least one multi-source gene expression level set in the multi-source gene expression level set of the patient samples and the multi-source gene expression level set of the healthy samples. A gene data acquisition unit is used to acquire a gene dataset, wherein the gene dataset includes at least one second gene and the association between each second gene, and each second gene is contained within each first gene; A gene vector generation unit is used to generate a target gene vector corresponding to each of the second genes based on the gene dataset. A gene vector sequence determination unit is used to determine the gene vector sequence of each patient sample based on each target gene vector and each target gene expression level. The gene vector sequence determination unit is further used for: For each patient sample, the expression levels of the target gene for each of the first genes in the patient sample are sorted in descending order to generate the first gene expression level sequence of the patient sample; and, For each of the first genes in the patient sample, determine whether there is a target second gene that matches the first gene in each of the second genes; If it is determined that it exists, the target gene expression level of the first gene in the first gene expression level sequence is replaced with the target gene vector of the target second gene, and the target gene vector of the target second gene is determined as the target gene vector of the first gene; Based on the target gene vectors of each of the first genes that have been replaced in the first gene expression sequence, the gene vector sequence of the patient sample is generated. The gene vector sequence retains the gene expression level of key genes and integrates the distributed representation of gene functional semantics, forming a multimodal input representation that combines individual expression characteristics and biological knowledge.
7. An electronic device, characterized in that, include: One or more processors; Storage device, on which one or more programs are stored, When one or more programs are executed by one or more processors, the one or more processors implement the method as claimed in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by one or more processors, implements the method as claimed in any one of claims 1-5.
9. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Association prediction method for drugs and pathways of knowledge graph attention network
CN114842927A
Phenotype analysis method, electronic equipment and storage medium
CN121171354A
Method for determining the presence of disease
US20110106739A1