Method, device and equipment for determining causal parameters, and storage medium

CN115662510BActive Publication Date: 2026-08-18JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211115933.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-14
Publication Date
2026-08-18
Estimated Expiration
2042-09-14

AI Technical Summary

Technical Problem

但是,相关技术往往忽略了由于混杂因子的存在而产生的混杂偏差,导致预测基因突变和癌症生物过程的状态变化之间的因果关系的准确性较低

Benefits of technology

[0028]On one hand, a computer program product or computer program is provided, comprising program code stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium and executes the program code, causing the computer device to perform the aforementioned method for determining causal parameters. Through the technical solution provided in this application, by processing gene expression data from multiple biological tissues, reference biological process activity data indicating that the multiple biological tissues have reached a target state is obtained. Using the reference biological process activity data, the final target biological process activity data can be determined. All multiple biological tissues carry the target gene and are in the target state. Encoding somatic mutation data, first-type confounding factor data, and reference biological process activity data from multiple biological tissues yields second-type confounding factor data. The first-type and second-type confounding factor data have different observability. Through the above process, estimation of the unobservable second-type confounding factor data is achieved. Decoding the second-type confounding factor data yields the target biological process activity data indicating that the multiple biological tissues have reached the target state. By using data on the activity of the target biological process, the causal parameters between the target gene and the target state can be determined. The process of determining these causal parameters eliminates the confounding effects of confounding factors and has high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115662510B_ABST
    Figure CN115662510B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for determining a causal parameter, an equipment and a storage medium, and belongs to the technical field of computers. The technical scheme provided by the embodiment of the application processes gene expression data of multiple biological tissues to obtain reference biological process activity data of the multiple biological tissues changing into a target state. The somatic mutation data of the multiple biological tissues, the first type of confounder data of the multiple biological tissues, and the reference biological process activity data are encoded to obtain the second type of confounder data, and the first type of confounder data and the second type of confounder data have different observability. Decoding the second type of confounder data can obtain target biological process activity data of the multiple biological tissues changing into the target state. Through the target biological process activity data, the causal parameter between a target gene and the target state can be determined, and the process of determining the causal parameter eliminates the confounding influence of confounders, and is high in accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for determining causal parameters. Background Technology

[0002] With the development of computer technology, research on genes has become increasingly in-depth, and computer technology can greatly improve the efficiency of gene research. Changes in the state of biological tissues may be caused by gene mutations, and studying the correlation between tissue state and gene mutations is of great significance. For example, numerous studies have shown that cancer is often caused by gene mutations. However, due to various technological limitations, we still do not fully understand which gene mutations lead to the occurrence and development of cancer.

[0003] In related technologies, large amounts of multi-omics data are typically used to identify gene mutations driving cancer by calculating mutation frequencies. However, these technologies often overlook confounding biases caused by the presence of confounding factors, resulting in low accuracy in predicting the causal relationship between gene mutations and changes in the state of cancer biological processes. Summary of the Invention

[0004] This application provides a method, apparatus, device, and storage medium for determining causal parameters, which can improve the accuracy of predicting the causal relationship between gene mutations and state changes. The technical solution is as follows:

[0005] On the one hand, a method for determining causal parameters is provided, the method comprising:

[0006] Gene expression data from multiple biological tissues are processed to obtain reference biological process activity data of the multiple biological tissues as they become the target state, wherein the multiple biological tissues all carry the target gene and are in the target state.

[0007] Somatic mutation data of the multiple biological tissues, first type confounding factor data of the multiple biological tissues, and reference biological process activity data are encoded to obtain second type confounding factor data of the multiple biological tissues. The first type confounding factor data and the second type confounding factor data have different observability.

[0008] Decode the second type of confounding factor data of the multiple biological tissues to obtain the target biological process activity data of the multiple biological tissues transforming into the target state;

[0009] Based on the target biological process activity data, a causal parameter is determined between the target gene and the target state, wherein the causal parameter is used to represent the probability that a mutation in the target gene will cause the biological tissue to be in the target state.

[0010] On the one hand, an apparatus for determining causal parameters is provided, the apparatus comprising:

[0011] The reference biological process data acquisition module is used to process gene expression data of multiple biological tissues to obtain reference biological process activity data of the multiple biological tissues in the target state, wherein the multiple biological tissues carry the target gene and are in the target state.

[0012] The encoding module is used to encode the somatic mutation data of the plurality of biological tissues, the first type of confounding factor data of the plurality of biological tissues, and the reference biological process activity data to obtain the second type of confounding factor data of the plurality of biological tissues. The first type of confounding factor data and the second type of confounding factor data have different observability.

[0013] The decoding module is used to decode the second type of confounding factor data of the multiple biological tissues to obtain the target biological process activity data of the multiple biological tissues becoming the target state;

[0014] The causal parameter determination module is used to determine the causal parameters between the target gene and the target state based on the target biological process activity data. The causal parameters are used to represent the probability that a mutation in the target gene will cause the biological tissue to be in the target state.

[0015] In one possible implementation, the reference biological process data acquisition module is used to determine the correlation between multiple genes in the multiple biological tissues based on the gene expression data of the multiple biological tissues; to determine the core genes of the multiple biological tissues from the multiple genes based on the correlation between the multiple genes in the multiple biological tissues; and to regress the average expression vector of the core genes of the multiple biological tissues on the gene expression data of the core genes of the multiple biological tissues to obtain reference biological process activity data of the multiple biological tissues changing to the target state.

[0016] In one possible implementation, the reference biological process data acquisition module is used to acquire multiple gene expression vectors corresponding to the multiple genes from the gene expression data of the multiple biological tissues; and to determine the correlation between multiple genes in the multiple biological tissues based on the correlation between the multiple gene expression vectors.

[0017] In one possible implementation, the reference biological process data acquisition module is used to determine the global correlation between each gene and other genes in the plurality of genes based on the correlation between multiple genes in the plurality of biological tissues; and to determine the core gene of the plurality of biological tissues from the plurality of genes based on the global correlation between each gene and other genes.

[0018] In one possible implementation, the reference biological process data acquisition module is used to, for any gene among the plurality of genes, fuse the correlation between the gene and other genes among the plurality of genes with the corresponding significance to obtain the target correlation between the gene and other genes; and to perform a weighted summation of the target correlation between the gene and other genes to obtain the global correlation between the gene and other genes.

[0019] In one possible implementation, the reference biological process data acquisition module is used to sort the plurality of genes in descending order of global relevance; and to determine the top target number of genes among the plurality of genes as the core genes of the plurality of biological tissues.

[0020] In one possible implementation, the reference biological process data acquisition module is used to determine the regression coefficient of the expression vector of the core gene of the plurality of biological tissues on the average expression vector of the core gene in the plurality of biological tissues, wherein the average expression vector is the average value of the gene expression vector of the core gene in the plurality of biological tissues; and to determine the regression coefficient as reference biological process activity data for the plurality of biological tissues to become the target state.

[0021] In one possible implementation, the encoding module is configured to input somatic mutation data of the plurality of biological tissues, first type confounding factor data of the plurality of biological tissues, and reference biological process activity data into an encoder; the encoder encodes the first type confounding factor data of the plurality of biological tissues and the reference biological process activity data to obtain a first encoding vector for each of the biological tissues; the encoder further encodes the first encoding vectors based on the somatic mutation data of the plurality of biological tissues to obtain second type confounding factor data of the plurality of biological tissues.

[0022] In one possible implementation, the encoding module is configured to perform at least one full connection on the first type of confounding factor data and reference biological process activity data of any of the plurality of biological tissues to obtain a first encoding vector of the biological tissue.

[0023] In one possible implementation, the encoding module is configured to, for any one of the plurality of biological tissues, when the somatic mutation data of the biological tissue indicates that no gene mutation has occurred in the biological tissue, encode the first encoding vector of the biological tissue through the first neural network of the encoder to obtain the second type of confounding factor data of the biological tissue; and when the somatic mutation data of the biological tissue indicates that a gene mutation has occurred in the biological tissue, encode the first encoding vector of the biological tissue through the second neural network of the encoder to obtain the second type of confounding factor data of the biological tissue.

[0024] In one possible implementation, the decoding module is used to input the second type of confounding factor data of the plurality of biological tissues into the generator; and the generator generates data based on the second type of confounding factor to obtain the target biological process activity data of the plurality of biological tissues.

[0025] In one possible implementation, the target biological process activity data includes first biological process activity data and second biological process activity data. The first biological process activity data is the biological process activity data of the biological tissue when the target gene has not mutated, and the second biological process activity data is the biological process activity data of the biological tissue when the target gene has mutated. The causal parameter determination module is used to perform a weighted summation of the target difference to obtain the causal parameter between the target gene and the target state. The target difference is the difference between the first biological process activity data and the second biological process activity data.

[0026] On one hand, a computer device is provided, the computer device including one or more processors and one or more memories, the one or more memories storing at least one computer program, the computer program being loaded and executed by the one or more processors to implement the method for determining the causal parameters.

[0027] On one hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, the computer program being loaded and executed by a processor to implement the method for determining the causal parameters.

[0028] On one hand, a computer program product or computer program is provided, comprising program code stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium and executes the program code, causing the computer device to perform the aforementioned method for determining causal parameters. Through the technical solution provided in this application, by processing gene expression data from multiple biological tissues, reference biological process activity data indicating that the multiple biological tissues have reached a target state is obtained. Using the reference biological process activity data, the final target biological process activity data can be determined. All multiple biological tissues carry the target gene and are in the target state. Encoding somatic mutation data, first-type confounding factor data, and reference biological process activity data from multiple biological tissues yields second-type confounding factor data. The first-type and second-type confounding factor data have different observability. Through the above process, estimation of the unobservable second-type confounding factor data is achieved. Decoding the second-type confounding factor data yields the target biological process activity data indicating that the multiple biological tissues have reached the target state. By using data on the activity of the target biological process, the causal parameters between the target gene and the target state can be determined. The process of determining these causal parameters eliminates the confounding effects of confounding factors and has high accuracy. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram of the implementation environment of a method for determining causal parameters provided in an embodiment of this application;

[0031] Figure 2 This is a flowchart of a method for determining causal parameters provided in an embodiment of this application;

[0032] Figure 3 This is a flowchart of another method for determining causal parameters provided in an embodiment of this application;

[0033] Figure 4 This is a schematic diagram illustrating the principle of a method for determining causal parameters provided in an embodiment of this application;

[0034] Figure 5 This is an architecture diagram of a method for determining causal parameters provided in an embodiment of this application;

[0035] Figure 6This is a schematic diagram of the structure of a device for determining causal parameters provided in an embodiment of this application;

[0036] Figure 7 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application;

[0037] Figure 8 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0039] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.

[0040] Artificial intelligence (AI) is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0041] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge sub-models to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.

[0042] The Gaussian distribution, also known as the normal distribution, has a bell-shaped curve, high in the middle and low at both ends. The expected value μ determines the position of the Gaussian distribution curve, while the standard deviation σ determines its range. The Gaussian distribution with μ = 0 and σ = 1 is the standard Gaussian distribution.

[0043] Embedded coding, mathematically speaking, represents a correspondence, that is, mapping data in space X to space Y using a function F. This function F is injective, and the mapping result preserves the structure. An injective function means that the mapped data uniquely corresponds to the original data, and preserving the structure means that the order of the original data remains the same. For example, if there are data X1 and X2 before mapping, after mapping we get Y1 corresponding to X1 and Y2 corresponding to X2. If the original data X1 > X2, then correspondingly, the mapped data Y1 > Y2. For words, this means mapping words to another space to facilitate subsequent machine learning and processing.

[0044] A key task in causal relationship research is causal effect estimation, which involves estimating the degree of change in the outcome variable if a different value were assigned to the treatment variable. A fundamental challenge in causal effect estimation is eliminating confounding effects, especially with highly dimensional data. Confounding effects arise from improper modeling of confounding factors, which are variables that simultaneously affect both the treatment and outcome variables. The presence of confounding factors in the model can distort the relationship between the treatment variable (e.g., mutation) and the outcome (e.g., cell proliferation), leading to erroneous results. For example, when estimating the causal effect of mutated TP53 on cell proliferation, oxidative stress level may be a confounding factor because it affects both the TP53 mutation probability and the degree of cell proliferation. When the distribution of oxidative stress levels differs between the TP53-mutated and non-mutated sample groups, it can bias the true effect of TP53 mutation on cell proliferation. Traditional statistical causal models reduce the influence of confounding factors by balancing confounding factors across groups, standardizing and stratifying the data, or performing regression analysis between confounding and treatment variables on observational data. However, these causal models are all based on the assumption of non-confounding, that is, all confounding factors are observable, which is unlikely in many complex biological systems studies. For example, we cannot know exactly which microenvironmental factors affect mutations, nor can we measure the indices of most microenvironmental factors. The technical solutions provided in this application consider the influence of confounding factors when studying causal relationships, and try to eliminate the influence of the presence of confounding factors on causal relationships as much as possible.

[0045] Figure 1 This is a schematic diagram illustrating the implementation environment of a method for determining causal parameters provided in this application embodiment. See also... Figure 1 The implementation environment may include terminal 110 and server 140.

[0046] Terminal 110 is connected to server 140 via a wireless or wired network. Optionally, terminal 110 may be a smartphone, tablet, laptop, desktop computer, smartwatch, etc., but is not limited to these. Terminal 110 has an application installed and running that supports the determination of causal parameters.

[0047] Server 140 is a standalone physical server, or a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.

[0048] Those skilled in the art will understand that the number of terminals and servers described above can be more or less. For example, there may be only one terminal and server, or dozens or hundreds of terminals, or even more. In this case, the implementation environment may also include other terminals and servers. This application does not limit the number or type of terminals and servers in its embodiments.

[0049] In the embodiments of this application, the technical solutions provided in the embodiments of this application can be implemented by a server or a terminal as the execution subject, or the technical methods provided in this application can be implemented through the interaction between the terminal and the server. The embodiments of this application do not limit this.

[0050] After introducing the implementation environment of the embodiments of this application, the application scenarios of the embodiments of this application will be described below. In the following description, the terminal is the terminal 110 in the above implementation environment, and the server is the server 140 in the above implementation environment.

[0051] The technical solutions provided in this application can be applied to scenarios where the causal relationship between gene mutation and disease can be determined, such as the causal relationship between gene mutation and cancer, or the causal relationship between gene mutation and other diseases. This application does not limit these applications.

[0052] In determining the causal relationship between gene mutation and cancer, gene expression data from multiple biological tissues are processed to obtain reference biological process activity data for the transformation of these tissues into a cancerous state. This reference biological process activity data is an estimated biological process activity data. Using this reference biological process activity data, the final target biological process activity data can be determined. Since multiple biological tissues all carry the target gene and are in a cancerous state, they serve as a sample for exploring causality. Somatic mutation data from multiple biological tissues, first-type confounding factor data from multiple biological tissues, and the reference biological process activity data are encoded to obtain second-type confounding factor data. The first-type and second-type confounding factor data have different observability; in some embodiments, the first-type confounding factor data is observable, while the second-type confounding factor data is unobservable. Through the above process, the unobservable second-type confounding factor data is estimated. Decoding the second-type confounding factor data yields the target biological process activity data for the transformation of multiple biological tissues into a cancerous state. Using the target biological process activity data, the causal parameters between the target gene and the cancerous state can be determined, that is, the probability that a mutation in the target gene leads to a cancerous state in the biological tissue can be determined.

[0053] It should be noted that the above description is based on the example of determining the causal relationship between gene mutation and cancer. In other possible implementations, the technical solutions provided in this application can also be applied to scenarios where the causal relationship between gene mutation and other states is being determined. This application does not limit the scope of these applications.

[0054] It should be noted that the following description of the technical solution provided in this application uses a terminal as the execution subject as an example. In other possible implementations, the technical solution provided in this application can also be executed jointly by a terminal and a server. The embodiments of this application do not limit the type of execution subject.

[0055] After introducing the implementation environment and application scenarios of the embodiments of this application, the technical solutions provided by the embodiments of this application are described below. (See also...) Figure 2 Taking the terminal as the executing entity as an example, the method includes the following steps.

[0056] 202. The terminal processes the gene expression data of multiple biological tissues to obtain reference biological process activity data of the multiple biological tissues to become the target state. All of the multiple biological tissues carry the target gene and are in the target state.

[0057] In this study, the multiple biological tissues serve as samples for investigating the causal relationship between gene mutations and a target state. In some embodiments, the target state refers to cancer, in which case all the multiple biological tissues are cancerous tissues, and the target gene is a gene within the biological tissue selected to investigate the causal relationship between gene mutations and cancer. Reference biological process activity data are estimated biological process activity data, which refers to the transcriptional level of a biological tissue in relation to a biological process, and this transcriptional level is associated with the expression values ​​of a set of genes related to that biological process.

[0058] 204. The terminal encodes the somatic mutation data of the multiple biological tissues, the first type of confounding factor data of the multiple biological tissues, and the reference biological process activity data to obtain the second type of confounding factor data of the multiple biological tissues. The first type of confounding factor data and the second type of confounding factor data have different observability.

[0059] The somatic mutation data of biological tissues is used to indicate whether a target gene in a biological tissue has mutated. In some embodiments, when a target gene in a biological tissue has mutated, the somatic mutation data of that biological tissue is a first value; when the target gene in a biological tissue has not mutated, the somatic mutation data of that biological tissue is a second value. The first value and the second value are different. The first type of confounding factor data is an abstract expression of the first type of confounding factor. In some embodiments, the first type of confounding factor refers to an observable confounding factor. Correspondingly, the second type of confounding factor data is an abstract expression of the second type of confounding factor. In some embodiments, the second type of confounding factor refers to an unobservable confounding factor.

[0060] 206. The terminal decodes the second type of confounding factor data of the multiple biological tissues to obtain the target biological process activity data of the multiple biological tissues becoming the target state.

[0061] The target biological process activity data is biological process activity data regenerated based on the second type of confounding factor data, which eliminates the influence of the second type of confounding factor.

[0062] 208. Based on the target biological process activity data, the terminal determines the causal parameter between the target gene and the target state. The causal parameter is used to represent the probability that a mutation in the target gene will cause the biological tissue to be in the target state.

[0063] The technical solution provided in this application involves processing gene expression data from multiple biological tissues to obtain reference biological process activity data for the transformation of these tissues into the target state. This reference biological process activity data allows for the determination of the final target biological process activity data, where all biological tissues carry the target gene and are in the target state. Encoding somatic mutation data, first-type confounding factor data, and reference biological process activity data from multiple biological tissues yields second-type confounding factor data. The first and second types of confounding factor data have different observability. This process allows for the estimation of unobservable second-type confounding factor data. Decoding this second-type confounding factor data yields the target biological process activity data for the transformation of multiple biological tissues into the target state. Using the target biological process activity data, the causal parameter between the target gene and the target state can be determined. This determination of the causal parameter eliminates the confounding effects of the confounding factors, resulting in high accuracy.

[0064] The following is combined Figure 3 The principles of the embodiments of this application will be explained.

[0065] See Figure 3 In this diagram, nodes represent variables, and arrows indicate the direction of causal relationships. Specifically, if we want to estimate the causal effect of genes on cancer biological processes, then the outcome variable of the causal system is denoted as Y (target biological process activity data), i.e., the biological process activity of the cancer sample; the treatment variable is denoted as M (gene mutation data), i.e., the somatic mutation data of gene g, which is also the target gene; observable confounding factors are denoted as X (first-class confounding factor data), i.e., somatic data of genes other than gene g; and unobservable confounding factors are denoted as Z (second-class confounding factor data), such as oxidative stress levels and other difficult-to-measure confounding factors. Although we cannot directly act on the unobserved confounding factors Z, we can find their proxy variables and recover the posterior probability distribution of Z from the observed data using generative models, such as variational autoencoders. A key step in inferring a causal relationship (M→Y) is eliminating confounding effects caused by confounding factors. These confounding factors simultaneously affect both the intervention variable (M, via Z→M) and the outcome variable (Y, via Z→Y), leading to spurious statistical correlations between M and Y. The technical solution provided in the embodiments of this application can eliminate such spurious statistical correlations.

[0066] Steps 202-208 above are a brief description of the technical solutions provided in the embodiments of this application. The following will combine some examples with the above... Figure 3 The principles described herein will be explained in more detail to illustrate the technical solutions provided in the embodiments of this application. See also... Figure 4Taking the terminal as the executing entity as an example, the method includes the following steps.

[0067] 402. The terminal acquires gene expression data from multiple biological tissues.

[0068] These biological tissues serve as samples for investigating the causal relationship between gene mutations and target states. Gene expression data from these biological tissues is obtained through gene sequencing, specifically the sequence of base pairs. Since the human body has over 20,000 genes, each biological tissue also has over 20,000 gene expression data points, with a one-to-one correspondence between genes and their expression data. This gene expression data is high-resolution digital expression profile information obtained through transcriptome sequencing technology. For example, the expression data of a particular gene refers to the amount of RNA it transcribes; a higher expression level indicates a potentially more active biological function corresponding to that gene.

[0069] 404. The terminal processes the gene expression data of multiple biological tissues to obtain reference biological process activity data of the multiple biological tissues to become the target state. All of the multiple biological tissues carry the target gene and are in the target state.

[0070] In some embodiments, the target state refers to cancer, in which case all of the multiple biological tissues are cancerous tissues; correspondingly, one biological tissue is a tissue or cell sample taken from cancerous tissue in a cancer patient. The target gene is a gene within the biological tissue that is selected to investigate the causal relationship between gene mutations and cancer. Reference biological process activity data are estimated biological process activity data, where biological process activity refers to the transcriptional level of a biological tissue at a given biological process, which is correlated with the expression levels of a set of genes associated with that biological process.

[0071] In one possible implementation, the terminal determines the correlation between multiple genes in the multiple biological tissues based on gene expression data of the multiple biological tissues. Based on the correlation between the multiple genes in the multiple biological tissues, the terminal identifies the core gene of each of the multiple biological tissues. The terminal then regresses the gene expression data of the core gene of the multiple biological tissues onto the average expression vector of the core gene of the multiple biological tissues to obtain reference biological process activity data for the multiple biological tissues to become the target state.

[0072] Among them, multiple genes in biological tissues refer to genes involved in the biological process that transforms biological tissues into the target state. In the following description, this biological process is referred to as the target biological process.

[0073] In this implementation, the terminal can determine the core gene in multiple biological tissues based on the correlation between multiple genes in multiple biological tissues, estimate biological processes based on the core gene, and obtain reference biological process activity data with high accuracy and efficiency.

[0074] To provide a clearer explanation of the above embodiments, the following description will be divided into several parts.

[0075] The first part involves the terminal determining the correlation between multiple genes in the multiple biological tissues based on gene expression data from these multiple biological tissues.

[0076] In one possible implementation, the terminal obtains multiple gene expression vectors corresponding to the multiple genes from the gene expression data of the multiple biological tissues. Based on the correlation between the multiple gene expression vectors, the terminal determines the correlation between the multiple genes in the multiple biological tissues.

[0077] In this context, each gene corresponds to a gene expression vector, which is an abstract representation of the corresponding gene.

[0078] For example, the terminal converts gene expression data from multiple biological tissues into a gene expression data matrix, where each row of the matrix represents the gene expression vector of a single gene. The terminal then determines the Pearson correlation coefficient and significance coefficient between every two gene expression vectors in the gene expression matrix, where the significance coefficient is also referred to as the significance level. In some embodiments, the Pearson correlation coefficients between every two gene expression vectors in the gene expression matrix constitute a correlation matrix for the multiple genes, and correspondingly, multiple significance coefficients constitute a significance matrix for the multiple genes. By using the correlation matrix and significance matrix, the Pearson correlation coefficient and significance coefficient between every two genes in the multiple genes can be quickly determined.

[0079] For example, these multiple biological tissues include P genes involved in the target biological process, the number of biological tissues is N, and the gene expression data matrix of N biological tissues is... in, This represents the set of gene expression vectors for the i-th biological tissue. Let U represent the set of gene expression vectors for the j-th gene, where 1 ≤ i ≤ N, 1 ≤ j ≤ P, and N, P, i, and j are all positive integers. For this gene expression data matrix U, the terminal determines the Pearson correlation coefficients among the P genes to obtain the correlation matrix. and significance matrix in, It is the Pearson correlation coefficient between the i-th gene and the j-th gene. That is the corresponding significance coefficient.

[0080] The second part involves the terminal identifying the core gene of the multiple biological tissues based on the correlation between multiple genes in these multiple biological tissues.

[0081] In one possible implementation, the terminal determines the global correlation between each gene and other genes within the plurality of genes based on the correlations among the plurality of genes in the plurality of biological tissues. Based on the global correlations between each gene and other genes, the terminal identifies the core genes of the plurality of biological tissues from among the plurality of genes.

[0082] Global correlation refers to the sum of correlations between a gene and other genes, while core genes refer to genes whose global correlation meets the correlation criteria among multiple genes.

[0083] For example, for any one of the multiple genes, the terminal fuses the correlation and corresponding significance between that gene and the other genes in the multiple genes to obtain the target correlation between that gene and other genes. The terminal then performs a weighted sum of the target correlations between that gene and other genes to obtain the global correlation between that gene and other genes. The terminal sorts the multiple genes in descending order of global correlation. The terminal identifies the top target number of genes among these multiple genes as the core genes of the multiple biological tissues.

[0084] For example, for any one of the multiple genes, the terminal multiplies the correlation between that gene and the other genes in the multiple genes by an indicator function using the following formula (1) to obtain the target correlation between that gene and other genes. The value of the indicator function is related to significance. The terminal performs a weighted summation of the target correlations between that gene and other genes to obtain the global correlation between that gene and other genes. The terminal sorts the multiple genes in descending order of global correlation. The terminal identifies the top target number of genes among the multiple genes as the core genes of the multiple biological tissues. In some embodiments, the terminal identifies the top 50% of genes with the highest global correlation as the core genes of the multiple biological tissues.

[0085]

[0086] in, For the global correlation of gene j, Here, α is the indicator function, and α is the significance threshold. The significance of the indicator function is to filter out the correlations with significance below the significance threshold and obtain the final global correlation.

[0087] The third part involves regressing the gene expression data of the core genes of the multiple biological tissues to obtain reference biological process activity data for the transformation of these multiple biological tissues into the target state.

[0088] In one possible implementation, the terminal determines a regression coefficient on the average expression vector of the core gene in the plurality of biological tissues, where the average expression vector is the average of the gene expression vectors of the core gene in the plurality of biological tissues. The terminal determines this regression coefficient as reference biological process activity data for the plurality of biological tissues to become the target state.

[0089] To provide a clearer explanation of the above implementation methods, the method for determining the average expression vector of the core gene in these multiple biological tissues will be described below.

[0090] In one possible implementation, for any one of the multiple core genes in a biological sample, the terminal determines the average representation value of that core gene, which is the average of the representation vectors of that core gene in multiple biological tissues. The terminal concatenates the average representation values ​​of the multiple core genes to obtain the average expression vector of the multiple core genes in the multiple biological tissues. For example, the terminal determines the average representation value of the core gene using the following formula (2) and the average expression vector using the following formula (3).

[0091]

[0092] in, This refers to the core gene j * The average value.

[0093]

[0094] Where K is the average representation vector, and [] is the floor function.

[0095] After introducing the method for determining the average expression vector of the core gene in multiple biological tissues, the method for determining the reference biological process activity data in the above embodiments will be explained below.

[0096] For example, the terminal determines the regression coefficient of the core gene on the average expression vector of the multiple biological tissues using the following formula (4).

[0097]

[0098] Among them, y i It is the regression coefficient of the i-th biological tissue, which is also the reference biological process activity data.

[0099] 406. The terminal encodes the somatic mutation data of the multiple biological tissues, the first type of confounding factor data of the multiple biological tissues, and the reference biological process activity data to obtain the second type of confounding factor data of the multiple biological tissues. The first type of confounding factor data and the second type of confounding factor data have different observability.

[0100] The somatic mutation data of biological tissues is used to indicate whether a target gene in a biological tissue has mutated. In some embodiments, when a target gene in a biological tissue has mutated, the somatic mutation data of that biological tissue is a first value. When the target gene in a biological tissue has not mutated, the somatic mutation data of that biological tissue is a second value, and the first value and the second value are different. The first type of confounding factor data is an abstract expression of the first type of confounding factor. In some embodiments, the first type of confounding factor refers to an observable confounding factor. Correspondingly, the second type of confounding factor data is an abstract expression of the second type of confounding factor. In some embodiments, the second type of confounding factor refers to an unobservable confounding factor. The first type of confounding factor is the somatic mutation data of genes other than the target gene in the biological tissue.

[0101] In one possible implementation, the terminal inputs somatic mutation data of the multiple biological tissues, first-type confounding factor data of the multiple biological tissues, and reference biological process activity data into an encoder. The terminal uses the encoder to encode the first-type confounding factor data and the reference biological process activity data of the multiple biological tissues, obtaining a first encoding vector for each biological tissue. The terminal then uses the encoder to perform secondary encoding on the first encoding vector based on the somatic mutation data of the multiple biological tissues, obtaining second-type confounding factor data for the multiple biological tissues.

[0102] To provide a clearer explanation of the above embodiments, the following description will be divided into two parts.

[0103] The first part involves the terminal encoding the first type of confounding factor data of the multiple biological tissues and the reference biological process activity data through the encoder, thereby obtaining the first encoding vector for each biological tissue.

[0104] In one possible implementation, for any one of the plurality of biological tissues, the terminal performs at least one full connection on the first type of confounding factor data of the biological tissue and the reference biological process activity data to obtain the first encoding vector of the biological tissue.

[0105] In one possible implementation, for any one of the plurality of biological tissues, the terminal performs at least one convolution on the first type of confounding factor data of the biological tissue and the reference biological process activity data to obtain the first encoding vector of the biological tissue.

[0106] It should be noted that the steps in this first part are the first stage of the encoder's processing, the purpose of which is to learn an abstract representation of the first type of confounding factor data and the reference biological process activity data. Let x be the first type of confounding factor data. i The reference biological process activity data are denoted as y. i The first encoding vector is g(x) i y i An abstract representation of ). The encoder includes a two-stage processing procedure aimed at estimating the posterior probability q(z|m, x, y) of the second type of confounding factor data, where z represents the second type of confounding factor data and m represents the somatic mutation data. In some embodiments, the distribution of the second type of confounding factor data is a multivariate Gaussian distribution, and the mean and variance of the second type of confounding factor data can be estimated through the two stages.

[0107] The second part involves the terminal using the encoder to perform secondary encoding on the first encoding vector based on the somatic mutation data of the multiple biological tissues, thereby obtaining the second type of confounding factor data of the multiple biological tissues.

[0108] In one possible implementation, for any one of the plurality of biological tissues, if the somatic mutation data of that biological tissue indicates that no gene mutation has occurred, the terminal encodes the first encoding vector of that biological tissue through the first neural network of the encoder to obtain the second type of confounding factor data of that biological tissue. If the somatic mutation data of that biological tissue indicates that a gene mutation has occurred, the terminal encodes the first encoding vector of that biological tissue through the second neural network of the encoder to obtain the second type of confounding factor data of that biological tissue.

[0109] The somatic mutation data includes a first value and a second value. The first value indicates that the target gene in the biological tissue has not mutated, and the second value indicates that the target gene in the biological tissue has mutated. In some embodiments, the first value is 0 and the second value is 1.

[0110] For example, for any one of the multiple biological tissues, if the somatic mutation data of that biological tissue is a first value, the terminal encodes the first encoding vector of that biological tissue through the first neural network of the encoder to obtain the second type of confounding factor data of that biological tissue. If the somatic mutation data of that biological tissue is a second value, the terminal encodes the first encoding vector of that biological tissue through the second neural network of the encoder to obtain the second type of confounding factor data of that biological tissue. For example, the terminal obtains the second type of confounding factor data through the following formula (5) or (6). The encoding using the first neural network and the second neural network is also the process of fitting.

[0111]

[0112]

[0113] Where f0() is the function corresponding to the first neural network, f1() is the function corresponding to the second neural network, and μ j ε represents the mean of the second type of confounding factor for biological tissue j. j z is the standard deviation of the second type of confounding factor of biological tissue j. i It is a type II confounding factor.

[0114] 408. The terminal decodes the second type of confounding factor data of the multiple biological tissues to obtain the target biological process activity data of the multiple biological tissues becoming the target state.

[0115] The target biological process activity data is biological process activity data regenerated based on the second type of confounding factor data, which eliminates the influence of the second type of confounding factor.

[0116] In one possible implementation, the terminal inputs the second type of confounding factor data of the multiple biological tissues into the generator. The terminal then uses the generator to generate data based on the second type of confounding factor to obtain the target biological process activity data of the multiple biological tissues.

[0117] In this embodiment, the encoder and generator belong to the same variational autoencoder (VAE). In the variational autoencoder, the generator is also called the decoder.

[0118] For example, the terminal uses the generator to decode the second type of confounding factors based on the following formula (7) to obtain the target biological process activity data.

[0119] p(x i y i )mi |z i )=p(x i |z i )p(m i |z i )p(y i |z i m i (7)

[0120] Where p(x) i |z i )=f x (z i ), p(m i |z i ) = Ber(elu(f m (z i ))), In the above formula (7), the prior distribution of z is determined to be a standard normal distribution in each dimension, i.e. elu() is an ELU layer used to capture non-linear representations, and Ber() represents the Bernoulli distribution, used to calculate somatic mutation data m. i The probability. Since the activity value of biological processes is continuous, y i The distribution is parameterized as a Gaussian distribution, with different mean values ​​corresponding to different somatic mutation data, and the variance is fixed at ε.

[0121] It should be noted that the variational autoencoder provided in this application embodiment is obtained by training by minimizing the KL divergence between the data and the reconstructed data, and the loss function of the training process is the following formula (8).

[0122]

[0123] Where L is the loss function.

[0124] During training, for a dataset containing N samples, 80% of the samples are used as the training set and 20% are used as the test set.

[0125] 410. Based on the target biological process activity data, the terminal determines the causal parameter between the target gene and the target state. The causal parameter is used to represent the probability that a mutation in the target gene will cause the biological tissue to be in the target state.

[0126] In one possible implementation, the target biological process activity data includes first biological process activity data and second biological process activity data. The first biological process activity data represents the biological process activity data of the biological tissue when the target gene has not mutated, and the second biological process activity data represents the biological process activity data of the biological tissue when the target gene has mutated. The terminal performs a weighted summation of the target differences to obtain a causal parameter between the target gene and the target state, whereby the target difference is the difference between the first and second biological process activity data.

[0127] For example, the terminal determines the causal parameters using the following formula (9).

[0128]

[0129] Here, ATE is a causal parameter. When ATE is positive, the mutation of the gene promotes the activity of the biological process; when ATE is 0 or negative, the mutation of the gene does not promote the activity of the biological process. i (m=0) represents the activity data of the first biological process, Y i (m=1) represents the activity data of the second biological process. In some implementations, the above formula (9) is also referred to as the Average Treatment Effect (ATE) formula. i (m=0) and Y i The method for determining (m=1) is that q(z) can be obtained through the encoder. i |m=0,x i y i ) and q(z) i |m=1,x i y i ), and thus we can obtain z i The representation of z. i Substituting into formula (7) above, we can obtain p(y) i |z i m i =0) and p(y i |z i m i =1).

[0130] The following will combine Figure 5 The technical solutions provided in the embodiments of this application will be described.

[0131] See Figure 5For any biological tissue among multiple biological tissues, the terminal determines the correlation between multiple genes in the biological tissue based on the gene expression data 501 of that biological tissue, obtaining a correlation matrix 502. Based on the gene expression data 501 and the correlation matrix 502, the terminal determines the core gene 503 of the biological tissue. Based on the core gene of the biological tissue, the terminal determines reference biological process activity data 504. The terminal encodes the somatic mutation data of the multiple biological tissues, the first type of confounding factor data of the multiple biological tissues, and the reference biological process activity data 504 to obtain the second type of confounding factor data z of the multiple biological tissues. Based on the second type of confounding factor data z, the terminal determines the target biological process activity data. Target somatic cell mutation data And the target first type of confounding factor data Finally, based on the activity data of the target biological process, causal parameters can be determined. Figure 5 The technical framework shown is called CEBP (Causal Effect of a Mutation on Cancer Biological Process) and is used to estimate the causal effect of gene mutations on cancer biological processes.

[0132] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0133] The technical solution provided in this application involves processing gene expression data from multiple biological tissues to obtain reference biological process activity data for the transformation of these tissues into the target state. This reference biological process activity data allows for the determination of the final target biological process activity data, where all biological tissues carry the target gene and are in the target state. Encoding somatic mutation data, first-type confounding factor data, and reference biological process activity data from multiple biological tissues yields second-type confounding factor data. The first and second types of confounding factor data have different observability. This process allows for the estimation of unobservable second-type confounding factor data. Decoding this second-type confounding factor data yields the target biological process activity data for the transformation of multiple biological tissues into the target state. Using the target biological process activity data, the causal parameter between the target gene and the target state can be determined. This determination of the causal parameter eliminates the confounding effects of the confounding factors, resulting in high accuracy.

[0134] Figure 6 This is a schematic diagram of a device for determining causal parameters provided in an embodiment of this application. See also... Figure 6The device includes: a reference biological process data acquisition module 601, an encoding module 602, a decoding module 603, and a causal parameter determination module 604.

[0135] The reference biological process data acquisition module 601 is used to process the gene expression data of multiple biological tissues to obtain reference biological process activity data of the multiple biological tissues into the target state, wherein the multiple biological tissues all carry the target gene and are in the target state.

[0136] Encoding module 602 is used to encode somatic mutation data of the multiple biological tissues, first type confounding factor data of the multiple biological tissues, and reference biological process activity data to obtain second type confounding factor data of the multiple biological tissues. The first type confounding factor data and the second type confounding factor data have different observability.

[0137] The decoding module 603 is used to decode the second type of confounding factor data of the multiple biological tissues to obtain the target biological process activity data of the multiple biological tissues becoming the target state.

[0138] The causal parameter determination module 604 is used to determine the causal parameter between the target gene and the target state based on the target biological process activity data. The causal parameter is used to represent the probability that a mutation of the target gene will cause the biological tissue to be in the target state.

[0139] In one possible implementation, the reference biological process data acquisition module 601 is used to determine the correlation between multiple genes in the multiple biological tissues based on gene expression data of the multiple biological tissues. Based on the correlation between the multiple genes in the multiple biological tissues, the core genes of the multiple biological tissues are identified from the multiple genes. The gene expression data of the core genes of the multiple biological tissues are regressed against the average expression vector of the core genes of the multiple biological tissues to obtain reference biological process activity data for the multiple biological tissues to become the target state.

[0140] In one possible implementation, the reference biological process data acquisition module 601 is used to acquire multiple gene expression vectors corresponding to the multiple genes from the gene expression data of the multiple biological tissues. Based on the correlation between the multiple gene expression vectors, the correlation between the multiple genes in the multiple biological tissues is determined.

[0141] In one possible implementation, the reference biological process data acquisition module 601 is used to determine the global correlation between each gene and other genes among the plurality of genes based on the correlation between multiple genes in the plurality of biological tissues. Based on the global correlation between each gene and other genes among the plurality of genes, the core genes of the plurality of biological tissues are identified from the plurality of genes.

[0142] In one possible implementation, the reference biological process data acquisition module 601 is used to, for any gene among the plurality of genes, fuse the correlation between that gene and the other genes among the plurality of genes with the corresponding significance to obtain the target correlation between that gene and other genes. Then, it performs a weighted summation of the target correlations between that gene and other genes to obtain the global correlation between that gene and other genes.

[0143] In one possible implementation, the reference biological process data acquisition module 601 is used to sort the multiple genes in descending order of global relevance. The top target number of genes among these multiple genes are identified as the core genes of the multiple biological tissues.

[0144] In one possible implementation, the reference biological process data acquisition module 601 is used to determine the regression coefficient of the expression vector of the core gene in the plurality of biological tissues on the average expression vector of the core gene in the plurality of biological tissues, wherein the average expression vector is the average value of the gene expression vector of the core gene in the plurality of biological tissues. This regression coefficient is determined as reference biological process activity data for the plurality of biological tissues to become the target state.

[0145] In one possible implementation, the encoding module 602 is used to input somatic mutation data of the plurality of biological tissues, first-type confounding factor data of the plurality of biological tissues, and reference biological process activity data into an encoder. The encoder encodes the first-type confounding factor data of the plurality of biological tissues and the reference biological process activity data to obtain a first encoding vector for each biological tissue. Based on the somatic mutation data of the plurality of biological tissues, the encoder performs secondary encoding on the first encoding vector to obtain second-type confounding factor data for the plurality of biological tissues.

[0146] In one possible implementation, the encoding module 602 is used to perform at least one full connection on any biological tissue among the plurality of biological tissues, using the first type of confounding factor data of the biological tissue and reference biological process activity data, to obtain the first encoding vector of the biological tissue.

[0147] In one possible implementation, the encoding module 602 is used to encode a first encoding vector of any biological tissue among the plurality of biological tissues, when the somatic mutation data of the biological tissue indicates that no gene mutation has occurred in the biological tissue, to obtain second type confounding factor data of the biological tissue. Conversely, when the somatic mutation data of the biological tissue indicates that a gene mutation has occurred in the biological tissue, the first encoding vector of the biological tissue is encoded by the second neural network of the encoder to obtain second type confounding factor data of the biological tissue.

[0148] In one possible implementation, the decoding module 603 is used to input the second type of confounding factor data of the plurality of biological tissues into the generator. The generator then generates data based on the second type of confounding factor to obtain the target biological process activity data of the plurality of biological tissues.

[0149] In one possible implementation, the target biological process activity data includes first biological process activity data and second biological process activity data. The first biological process activity data is the biological process activity data of the biological tissue when the target gene has not mutated, and the second biological process activity data is the biological process activity data of the biological tissue when the target gene has mutated. The causal parameter determination module 604 is used to perform a weighted summation of the target difference to obtain the causal parameter between the target gene and the target state. The target difference is the difference between the first biological process activity data and the second biological process activity data.

[0150] It should be noted that the causal parameter determination device provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the causal parameter determination device and the causal parameter determination method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0151] The technical solution provided in this application involves processing gene expression data from multiple biological tissues to obtain reference biological process activity data for the transformation of these tissues into the target state. This reference biological process activity data allows for the determination of the final target biological process activity data, where all biological tissues carry the target gene and are in the target state. Encoding somatic mutation data, first-type confounding factor data, and reference biological process activity data from multiple biological tissues yields second-type confounding factor data. The first and second types of confounding factor data have different observability. This process allows for the estimation of unobservable second-type confounding factor data. Decoding this second-type confounding factor data yields the target biological process activity data for the transformation of multiple biological tissues into the target state. Using the target biological process activity data, the causal parameter between the target gene and the target state can be determined. This determination of the causal parameter eliminates the confounding effects of the confounding factors, resulting in high accuracy.

[0152] This application provides a computer device for performing the above-described method. This computer device can be implemented as a terminal or a server. The structure of the terminal will be described below:

[0153] Figure 7 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. The terminal 700 can be a smartphone, tablet computer, laptop computer, or desktop computer. The terminal 700 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.

[0154] Typically, terminal 700 includes one or more processors 701 and one or more memories 702.

[0155] Processor 701 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 701 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 701 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 701 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 701 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0156] The memory 702 may include one or more computer-readable storage media, which may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 are used to store at least one computer program, which is executed by the processor 701 to implement the method for determining causal parameters provided in the method embodiments of this application.

[0157] In some embodiments, the terminal 700 may also optionally include a peripheral device interface 703 and at least one peripheral device. The processor 701, memory 702, and peripheral device interface 703 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 703 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 704, a display screen 705, a camera assembly 706, an audio circuit 707, and a power supply 708.

[0158] Peripheral device interface 703 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 701 and memory 702. In some embodiments, processor 701, memory 702 and peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 701, memory 702 and peripheral device interface 703 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0159] The radio frequency (RF) circuit 704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 704 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 704 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc.

[0160] Display screen 705 is used to display a user interface (UI). This UI may include graphics, text, icons, video, and any combination thereof. When display screen 705 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 701 for processing. In this case, display screen 705 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard.

[0161] The camera assembly 706 is used to capture images or videos. Optionally, the camera assembly 706 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal.

[0162] The audio circuit 707 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input to the processor 701 for processing, or input to the radio frequency circuit 704 to realize voice communication.

[0163] The power supply 708 is used to supply power to the various components in the terminal 700. The power supply 708 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery.

[0164] In some embodiments, the terminal 700 further includes one or more sensors 709. The one or more sensors 709 include, but are not limited to: an accelerometer 710, a gyroscope 711, a pressure sensor 712, an optical sensor 713, and a proximity sensor 714.

[0165] Accelerometer 710 can detect the magnitude of acceleration on the three coordinate axes of a coordinate system established with terminal 700.

[0166] The gyroscope sensor 711 can detect the orientation and rotation angle of the terminal 700. The gyroscope sensor 711 can work in conjunction with the accelerometer sensor 710 to collect the user's 3D movements on the terminal 700.

[0167] The pressure sensor 712 can be installed on the side bezel of the terminal 700 and / or on the lower layer of the display screen 705. When the pressure sensor 712 is installed on the side bezel of the terminal 700, it can detect the user's grip signal on the terminal 700, and the processor 701 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 712. When the pressure sensor 712 is installed on the lower layer of the display screen 705, the processor 701 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 705.

[0168] An optical sensor 713 is used to collect ambient light intensity. In one embodiment, the processor 701 can control the display brightness of the display screen 705 based on the ambient light intensity collected by the optical sensor 713.

[0169] The proximity sensor 714 is used to detect the distance between the user and the front of the terminal 700.

[0170] Those skilled in the art will understand that Figure 7 The structure shown does not constitute a limitation on terminal 700, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0171] The aforementioned computer equipment can also be implemented as a server. The structure of a server is described below:

[0172] Figure 8This is a schematic diagram of a server structure provided in an embodiment of this application. The server 800 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 801 and one or more memories 802. The one or more memories 802 store at least one computer program, which is loaded and executed by the one or more processors 801 to implement the methods provided in the various method embodiments described above. Of course, the server 800 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 800 may also include other components for implementing device functions, which will not be elaborated upon here.

[0173] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program that can be executed by a processor to perform the method for determining causal parameters in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0174] In an exemplary embodiment, a computer program product or computer program is also provided, which includes program code stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium and executes the program code, causing the computer device to perform the method for determining the causal parameters described above.

[0175] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.

[0176] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0177] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for determining causal parameters, characterized in that, The method includes: Gene expression data from multiple biological tissues are processed to obtain reference biological process activity data of the multiple biological tissues as they become the target state, wherein the multiple biological tissues all carry the target gene and are in the target state. Somatic mutation data of the multiple biological tissues, first type confounding factor data of the multiple biological tissues, and reference biological process activity data are encoded to obtain second type confounding factor data of the multiple biological tissues. The first type confounding factor data and the second type confounding factor data have different observability. Decode the second type of confounding factor data of the multiple biological tissues to obtain the target biological process activity data of the multiple biological tissues transforming into the target state; Based on the target biological process activity data, a causal parameter is determined between the target gene and the target state, wherein the causal parameter is used to represent the probability that a mutation in the target gene will cause the biological tissue to be in the target state.

2. The method according to claim 1, characterized in that, The process of processing gene expression data from multiple biological tissues to obtain reference biological process activity data for the transformation of these tissues into the target state includes: Based on the gene expression data of the multiple biological tissues, the correlation between multiple genes in the multiple biological tissues is determined; Based on the correlation between multiple genes in the multiple biological tissues, the core genes of the multiple biological tissues are identified from the multiple genes; By regressing the gene expression data of the core genes of the multiple biological tissues onto the average expression vector of the core genes of the multiple biological tissues, reference biological process activity data for the multiple biological tissues to become the target state are obtained.

3. The method according to claim 2, characterized in that, The determination of the correlation between multiple genes in the multiple biological tissues based on gene expression data of the multiple biological tissues includes: Obtain multiple gene expression vectors corresponding to the multiple genes from the gene expression data of the multiple biological tissues; Based on the correlation between the multiple gene expression vectors, the correlation between multiple genes in the multiple biological tissues is determined.

4. The method according to claim 2, characterized in that, The method of determining the core gene of the multiple biological tissues based on the correlation between multiple genes in the multiple biological tissues includes: Based on the correlations among multiple genes in the multiple biological tissues, the global correlations between each gene and other genes are determined. Based on the global correlation between each gene and other genes, the core genes of the multiple biological tissues are identified from the multiple genes.

5. The method according to claim 4, characterized in that, The determination of the global correlation between each gene and other genes based on the correlation among multiple genes in the multiple biological tissues includes: For any one of the plurality of genes, the correlation between the gene and other genes in the plurality of genes is fused with the corresponding significance to obtain the target correlation between the gene and other genes. The target correlation between any gene and other genes is weighted and summed to obtain the global correlation between any gene and other genes.

6. The method according to claim 4, characterized in that, The process of identifying the core genes of the multiple biological tissues based on the global correlation between each gene and other genes includes: The genes are sorted in descending order of global relevance; The first target number of genes among the multiple genes are identified as the core genes of the multiple biological tissues.

7. The method according to claim 2, characterized in that, The regression analysis of gene expression data of core genes in the plurality of biological tissues to obtain reference biological process activity data for the transformation of the plurality of biological tissues into the target state includes: Determine the regression coefficient of the expression vector of the core gene in the plurality of biological tissues on the average expression vector in the plurality of biological tissues, wherein the average expression vector is the average value of the gene expression vector of the core gene in the plurality of biological tissues; The regression coefficients are determined as reference biological process activity data for the transformation of the plurality of biological tissues into the target state.

8. The method according to claim 1, characterized in that, The step of encoding the somatic mutation data of the plurality of biological tissues, the first type of confounding factor data of the plurality of biological tissues, and the reference biological process activity data to obtain the second type of confounding factor data of the plurality of biological tissues includes: Somatic cell mutation data of the plurality of biological tissues, first-class confounding factor data of the plurality of biological tissues, and reference biological process activity data are input into the encoder; The encoder encodes the first type of confounding factor data of the plurality of biological tissues and the reference biological process activity data to obtain the first encoding vector of each of the biological tissues. The encoder performs secondary encoding on the first encoding vector based on the somatic mutation data of the multiple biological tissues to obtain the second type of confounding factor data of the multiple biological tissues.

9. The method according to claim 8, characterized in that, The step of encoding the first type of confounding factor data of the plurality of biological tissues and the reference biological process activity data through the encoder to obtain the first encoding vector of each of the biological tissues includes: For any biological tissue among the plurality of biological tissues, a first type of confounding factor data of the biological tissue and reference biological process activity data are fully connected at least once to obtain the first encoding vector of the biological tissue.

10. The method according to claim 8, characterized in that, The step of using the encoder to perform secondary encoding on the first encoding vector based on the somatic mutation data of the multiple biological tissues to obtain the second type of confounding factor data of the multiple biological tissues includes: For any biological tissue among the plurality of biological tissues, if the somatic mutation data of the biological tissue indicates that no gene mutation has occurred in the biological tissue, the first encoding vector of the biological tissue is encoded by the first neural network of the encoder to obtain the second type of confounding factor data of the biological tissue; When the somatic mutation data of the biological tissue indicates that the biological tissue has undergone gene mutation, the first encoding vector of the biological tissue is encoded by the second neural network of the encoder to obtain the second type of confounding factor data of the biological tissue.

11. The method according to claim 1, characterized in that, Decoding the second type of confounding factor data of the plurality of biological tissues to obtain target biological process activity data of the plurality of biological tissues transforming into the target state includes: Input the second type of confounding factor data of the multiple biological tissues into the generator; The generator generates data based on the second type of confounding factor to obtain target biological process activity data of the multiple biological tissues.

12. The method according to claim 1, characterized in that, The target biological process activity data includes first biological process activity data and second biological process activity data. The first biological process activity data is the biological process activity data of a biological tissue when the target gene has not mutated, and the second biological process activity data is the biological process activity data of a biological tissue when the target gene has mutated. Determining the causal parameters between the target gene and the target state based on the target biological process activity data includes: The target difference is weighted and summed to obtain the causal parameter between the target gene and the target state, where the target difference is the difference between the first biological process activity data and the second biological process activity data.

13. A device for determining causal parameters, characterized in that, The device includes: The reference biological process data acquisition module is used to process gene expression data of multiple biological tissues to obtain reference biological process activity data of the multiple biological tissues in the target state, wherein the multiple biological tissues carry the target gene and are in the target state. The encoding module is used to encode the somatic mutation data of the plurality of biological tissues, the first type of confounding factor data of the plurality of biological tissues, and the reference biological process activity data to obtain the second type of confounding factor data of the plurality of biological tissues. The first type of confounding factor data and the second type of confounding factor data have different observability. The decoding module is used to decode the second type of confounding factor data of the multiple biological tissues to obtain the target biological process activity data of the multiple biological tissues becoming the target state; The causal parameter determination module is used to determine the causal parameters between the target gene and the target state based on the target biological process activity data. The causal parameters are used to represent the probability that a mutation in the target gene will cause the biological tissue to be in the target state.

14. A computer device, characterized in that, The computer device includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, the computer program being loaded and executed by the one or more processors to implement the method for determining causal parameters as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the method for determining causal parameters as described in any one of claims 1 to 12.

16. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the method for determining causal parameters as described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Association mining method and device, equipment, medium and computer program product

    CN114283887A

  • Systems, methods, and processor-readable media for detecting disease causal variants

    US20190087534A1