Gene expression amount calculation method, electronic device, and storage medium
By constructing multiple models to consider microenvironmental factors and fitting gene expression levels, the problem of low accuracy in gene expression level calculation is solved, the calculation precision is improved, and biological research and medical diagnosis are supported.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
- Filing Date
- 2023-11-28
- Publication Date
- 2026-05-19
AI Technical Summary
In existing technologies, the calculation of gene expression levels fails to take into account the influence of the cell's microenvironment, resulting in low accuracy and affecting medical clinical diagnosis, drug efficacy assessment, and research on disease mechanisms.
By constructing multiple models that take into account the microenvironmental factors of the target sample cells, such as neighboring cell types, spatial distance, and ligand information, gene distribution is fitted using methods such as Poisson distribution and penalized regression to obtain gene expression levels.
It improves the accuracy of gene expression level calculation, provides reliable biological research data, helps to understand gene function and regulatory mechanisms, and supports disease research, drug response prediction, and biomarker discovery.
Smart Images

Figure CN120072050B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biomedical technology, specifically to a method for calculating gene expression levels, an electronic device, and a storage medium. Background Technology
[0002] In related technologies, the study of gene expression in cells typically describes cellular characteristics based on the role of ligand signaling molecules. This method assumes that gene expression within cells is independent of each other. Currently, there is no method that can account for the influence of the cellular microenvironment on gene expression, resulting in low accuracy in the calculation of gene expression levels.
[0003] If the accurate gene expression level cannot be determined, it will affect the analysis of gene expression, making it impossible to clarify which genes have changed expression, what the correlations are between genes, and how gene activity is affected under different conditions. This will inevitably lead to significant biases in medical clinical diagnosis, drug efficacy assessment, and disease pathogenesis research based on gene expression. Summary of the Invention
[0004] In view of the above, it is necessary to propose a gene expression level calculation method, electronic device and storage medium that can solve the problem of low accuracy in gene expression level calculation caused by not taking into account the influence of the cell's microenvironment on gene expression.
[0005] This application provides a method for calculating gene expression levels. The method includes: acquiring cell data from multiple sample cells, including target sample cells; constructing a first model based on first data in the cell data; solving the first model using a preset solution method to obtain a first gene expression level of the target sample cells; wherein the first data represents environmental factors in a first microenvironment in which the target sample cells are located; the first data includes the cell types of the nearest neighbor cells of the target sample cells; and the first model is used to fit the gene distribution of the target sample cells when affected by the first microenvironment. In one embodiment, constructing the first model based on the first data in the cell data includes: fitting the distribution of sample genes in the target sample cells using a preset distribution function based on the gene expression level of each sample gene in the target sample cells, the cell types of the nearest neighbor cells of the target sample cells, and the number of the nearest neighbor cells of the target sample cells.
[0006] In one embodiment, constructing the first model based on the first data in the cell data includes: fitting the distribution of sample genes of the target sample cells using a Poisson distribution, wherein the formula used includes: ,in, Indicates target sample cells Sample genes Gene expression levels, Indicates gene expression level Expectations; for sample genes ,Will Fit to Explanatory variables The combination of these formulas includes: ,in, This indicates the number of cell types belonging to the nearest neighbor cells of the target sample cell. ; and This represents the Poisson parameter.
[0007] In one embodiment, the solution method includes: solving for the Poisson parameter using the maximum likelihood method.
[0008] In one embodiment, the method further includes: applying a penalty regression to the solution method using a preset optimal regularization parameter, wherein the penalty regression includes ridge regression and lasso regression.
[0009] In one embodiment, the method further includes: determining, based on cross-validation, a regularization parameter that results in the highest determination coefficient score for the first model from a set of preset regularization parameters as the optimal regularization parameter.
[0010] In one embodiment, the method further includes performing a significance test on the Poisson parameter, comprising: estimating the Poisson parameter to obtain an estimated value of the Poisson parameter; determining the deviation between the estimated value of the Poisson parameter and the corresponding unknown true parameter, and approximating the deviation to a standard normal distribution based on the principle of asymptotic normality; calculating the z-value and p-value of a two-tailed test based on the deviation, and determining a q-value based on the p-value to control false positives of the Poisson parameter.
[0011] This application provides a method for calculating gene expression levels. The method includes: acquiring cell data of multiple sample cells, including target sample cells; constructing a second model based on second data in the cell data; solving the second model using a preset solution method to obtain the second gene expression level of the target sample cells; wherein the second data represents environmental factors in a second microenvironment in which the target sample cells are located; the second data includes the cell type to which the multiple sample cells belong and the spatial distance between the target sample cells and other sample cells; the second model is used to fit the gene distribution in the target sample cells when affected by the second microenvironment.
[0012] In one embodiment, constructing the second model based on the second data in the cell data includes: fitting the distribution of sample genes in the target sample cells using a preset distribution function based on the gene expression level of each sample gene in the target sample cells, the cell types to which the plurality of sample cells belong, the total number of cell types to which the plurality of sample cells belong, and the spatial distance between the target sample cells and other sample cells.
[0013] In one embodiment, the formula used to construct the second model based on second data from the cell data includes:
[0014] ,
[0015] in, Indicates target sample cells Sample genes Gene expression levels, Indicates gene expression level Expectation; error term , express variance This represents the total number of cells in the multiple sample samples. This represents the total number of cell types to which the multiple sample cells belong. Indicates target sample cells One-hot encoded vector, Indicates sample cells One-hot encoded vector, Indicates target sample cells With sample cells Distance weights, The function is used to... Transform a 3D matrix into 3D vector and This represents the Poisson parameter.
[0016] This application provides a method for calculating gene expression levels. The method includes: acquiring cell data from multiple sample cells, including target sample cells; constructing a third model based on third data in the cell data; solving the third model using a preset solution method to obtain the third gene expression level of the target sample cells; wherein the third data represents environmental factors in the third microenvironment in which the target sample cells are located; the third data includes the cell type of the multiple sample cells, the spatial distance between the target sample cells and other sample cells, and ligand information of the multiple sample cells; the third model is used to fit the gene distribution in the target sample cells when affected by the third microenvironment.
[0017] In one embodiment, constructing a third model based on third data in the cell data includes: fitting the distribution of sample genes in the target sample cells using a preset distribution function based on the gene expression level of each sample gene in the target sample cells, the cell types to which the plurality of sample cells belong, the total number of cell types to which the plurality of sample cells belong, the spatial distance between the target sample cells and other sample cells, and the ligand information of the plurality of sample cells.
[0018] In one embodiment, the formula used to construct the third model based on third data in the cell data includes:
[0019] ,
[0020] in, Indicates target sample cells Sample genes Gene expression levels, Indicates gene expression level Expectation; error term , express variance This represents the total number of cells in the multiple sample samples. This represents the total number of cell types to which the multiple sample cells belong. Indicates target sample cells One-hot encoded vector, Indicates sample cells One-hot encoded vector, Indicates target sample cells With sample cells Distance weights, The function is used to... Transform a 3D matrix into 3D vector Indicates target sample cells The Middle Gene expression levels of receptors in each ligand pair Indicates target sample cells The Middle Gene expression levels of ligands in each ligand pair and This represents the Poisson parameter.
[0021] This application provides a gene expression level calculation device, comprising: a data acquisition module for acquiring cell data of multiple sample cells, the multiple sample cells including target sample cells; a model construction module for constructing a first model based on first data in the cell data, solving the first model using a preset solution method, and obtaining a first gene expression level of the target sample cells, wherein the first data represents environmental factors in a first microenvironment in which the target sample cells are located; the first data includes the cell type of the nearest neighbor cells of the target sample cells; the first model is used to fit the gene distribution of the target sample cells under the influence of the first microenvironment; the model construction module is further configured to construct a second model based on second data in the cell data, solving the second model using the solution method, and obtaining a second gene expression level of the target sample cells, wherein the second data... The second data represents the environmental factors in the second microenvironment in which the target sample cells are located; the second data includes the cell types to which the plurality of sample cells belong and the spatial distance between the target sample cells and other sample cells; the second model is used to fit the gene distribution in the target sample cells when affected by the second microenvironment; the model building module is further used to construct a third model based on the third data in the cell data, solve the third model using the solution method and obtain the third gene expression level of the target sample cells, wherein the third data represents the environmental factors in the third microenvironment in which the target sample cells are located; the third data includes the cell types to which the plurality of sample cells belong, the spatial distance between the target sample cells and other sample cells, and the ligand information of the plurality of sample cells; the third model is used to fit the gene distribution in the target sample cells when affected by the third microenvironment.
[0022] An embodiment of this application provides an electronic device, which includes a processor and a memory, wherein the processor is used to implement the gene expression level calculation method when executing a computer program stored in the memory.
[0023] An embodiment of this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the gene expression level calculation method.
[0024] In summary, the gene expression level calculation method described in this application can take into account the influence of various factors (such as neighboring cells and ligand information) in the multiple microenvironments in which the target sample cells are located. By constructing different models to simulate the effects of various factors in the microenvironment on the gene expression level of the target sample cells, it can make full use of and mine spatial omics data, comprehensively and deeply realize the interaction analysis of spatial omics data, improve the accuracy of gene expression level calculation, and thus provide reliable data for biological research based on gene expression levels. Attached Figure Description
[0025] Figure 1 This is a structural diagram of an electronic device provided in an embodiment of this application.
[0026] Figure 2 This is a flowchart of a gene expression level calculation method provided in an embodiment of this application.
[0027] Figure 3 This is a flowchart of a significance test provided in an embodiment of this application.
[0028] Figure 4 This is a flowchart illustrating the process of modeling and gene expression analysis based on the microenvironment of the cell, as provided in one embodiment of this application.
[0029] Figure 5 This is an example diagram of gene expression levels provided in an embodiment of this application.
[0030] Figure 6 This is an example diagram of gene expression levels provided in another embodiment of this application.
[0031] Figure 7 This is an example diagram of gene expression levels provided in another embodiment of this application.
[0032] Figure 8 This is a structural diagram of a gene expression level calculation device provided in an embodiment of this application. Detailed Implementation
[0033] To better understand the above-mentioned objectives, features, and advantages of this application, the application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the specification of this application is for the purpose of describing an embodiment in one instance only and is not intended to be limiting of the application.
[0035] It should be noted that in this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and drawings of this application are used to distinguish similar objects, not to describe a specific order or sequence.
[0036] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner. Unless otherwise specified, the following embodiments and features described herein can be combined with each other.
[0037] In one embodiment, related technologies typically describe cellular characteristics based on the role of ligand signaling molecules when studying gene expression in cells. This approach assumes that gene expression within cells is independent of each other. Currently, there is no method that accounts for the influence of the cellular microenvironment on gene expression, resulting in low accuracy in the calculation of gene expression levels.
[0038] If the accurate gene expression level cannot be determined, it will affect the analysis of gene expression, making it impossible to clarify which genes have changed expression, what the correlations are between genes, and how gene activity is affected under different conditions. This will inevitably lead to significant biases in medical clinical diagnosis, drug efficacy assessment, and disease pathogenesis research based on gene expression.
[0039] To address the aforementioned issues, this application provides a method for calculating gene expression levels. The method involves acquiring cell data from multiple sample cells, including a target sample cell. A first model is constructed based on the cell types of the target sample cell's nearest neighbor cells within the cell data. The first model is then solved using a preset solution method to obtain the first gene expression level of the target sample cell influenced by its nearest neighbor cells in the microenvironment. A second model is constructed based on the cell types of the multiple sample cells and the spatial distance between the target sample cell and other sample cells. The second model is solved using the same solution method to obtain the second gene expression level of the target sample cell influenced by all surrounding cells in the microenvironment. Finally, a third model is constructed based on the cell types of the multiple sample cells, the spatial distance between the target sample cell and other sample cells, and the ligand information of the multiple sample cells. The third model is solved using the same solution method to obtain the third gene expression level of the target sample cell influenced by the ligand information of other cells in the microenvironment.
[0040] This application takes into account the influence of various factors (such as neighboring cells and ligand information) in the various microenvironments in which the target sample cells are located. By constructing different models to simulate the effect of various factors in the microenvironment on the gene expression level of the target sample cells, it can make full use of and mine spatial group data, and realize the interaction analysis of spatial group data in a comprehensive and in-depth manner, thereby improving the accuracy of gene expression level calculation and providing reliable data for biological research based on gene expression level.
[0041] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 10 can be a mobile phone, tablet computer, smart wearable device, augmented reality (AR) / virtual reality (VR) device, laptop computer, netbook, energy storage device, power distribution equipment, vehicle-mounted equipment, self-moving device, or other electronic devices. This embodiment of the application does not impose any restrictions on the specific type of electronic device.
[0042] like Figure 1 As shown, the electronic device 10 may include a communication module 101, a memory 102, a processor 103, an input / output (I / O) interface 104, and a bus 105. The processor 103 is coupled to the communication module 101, the memory 102, and the I / O interface 104 via the bus 105.
[0043] Communication module 101 may include a wired communication module and / or a wireless communication module. The wired communication module may provide one or more wired communication solutions such as Universal Serial Bus (USB) and Controller Area Network (CAN). The wireless communication module may provide one or more wireless communication solutions such as Wireless Fidelity (Wi-Fi), Bluetooth (BT), mobile communication networks, Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR).
[0044] Memory 102 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM). The RAM can be directly read and written by the processor 103, and can be used to store executable programs (such as machine instructions) of the operating system or other running programs, as well as user and application data. The RAM may include static random-access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc.
[0045] Non-volatile memory can also store executable programs and user and application data, and can be pre-loaded into random access memory for direct reading and writing by the processor 103. Non-volatile memory can include disk storage devices and flash memory.
[0046] Memory 102 is used to store one or more computer programs. The one or more computer programs are configured to be executed by processor 103. The one or more computer programs include multiple instructions that, when executed by processor 103, enable a gene expression level calculation method to be executed on electronic device 10.
[0047] In other embodiments, the electronic device 10 further includes an external memory interface for connecting to an external memory to expand the storage capacity of the electronic device 10.
[0048] Processor 103 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0049] The processor 103 provides computational and control capabilities. For example, the processor 103 is used to execute computer programs stored in the memory 102 to implement the gene expression level calculation method described above.
[0050] I / O interface 104 is used to provide a channel for user input or output. For example, I / O interface 104 can be used to connect various input and output devices, such as mouse, keyboard, touch device, display screen, etc., so that users can enter information or visualize information.
[0051] Bus 105 is used at least to provide a channel for communication between communication modules 101, memory 102, processor 103, and I / O interface 104 in electronic device 10.
[0052] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 10. In other embodiments of this application, the electronic device 10 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0053] Figure 2 This is a flowchart of a gene expression level calculation method provided in an embodiment of this application. The gene expression level calculation method is applied in electronic devices, for example... Figure 1 The electronic device 10 in the process includes the following steps. Depending on different needs, the order of the steps in the flowchart can be changed, and some steps can be omitted.
[0054] S21, acquire cell data from multiple sample cells.
[0055] In one embodiment, the plurality of sample cells can be selected according to the actual application. For example, the plurality of sample cells may be cells from a brain tissue slice taken the day after a salamander's brain injury. Specifically, the plurality of sample cells includes target sample cells. To analyze the gene expression levels of each cell in the plurality of sample cells, any one of the sample cells can be used as the target sample cell, thereby calculating and analyzing the gene expression levels of the target sample cell affected by the microenvironment.
[0056] In one embodiment, spatial transcriptomics technology based on microarray technology can be used to obtain cellular data of the multiple sample cells. Specifically, the spatial transcriptomics technology is a technique based on single-cell ribonucleic acid (RNA) sequencing, which can correlate the transcriptome information of a single cell with its spatial location in the tissue. For example, the spatial transcriptomics technology may include the following implementation process: (1) Tissue slicing: cutting tissue (or organ, embryo, etc.) into thin slices and fixing them on a glass slide; (2) RNA extraction: extracting RNA from each tissue slice; (3) Library preparation: converting the RNA of each single cell into complementary deoxyribonucleic acid (cDNA) and amplifying it through reverse transcription and library preparation; (4) Microarray loading: uniformly loading the sample into the microarray wells; (5) Image acquisition: taking images of each microarray point using a high-resolution microscope; (6) Data processing: combining the image data with the single-cell RNA sequencing data using specialized software to obtain the gene expression data of each single cell and its spatial location (cell coordinate) in the tissue or slice.
[0057] The gene expression level calculation method provided in this application includes the above steps (1)-(6). As can be seen from the above spatial transcriptomics technology, gene expression data and spatial location data of multiple sample cells in a single or multiple adjacent chips can be obtained as the cell data based on spatial transcriptomics technology.
[0058] The gene expression data may include the genes of the sample cells and their corresponding gene expression levels. Alternatively, a preset category of sample genes and their corresponding gene expression levels can be selected from the sample cells as the cell data. This preset category may be genes of interest selected based on the actual application, such as genes with spatial characteristics (e.g., genes with Moranis > 0.4) or genes with a high positive expression rate (e.g., genes with a positive rate greater than 5%). The spatial location data may include the two-dimensional or three-dimensional spatial coordinates of each sample cell.
[0059] In addition, the cell data also includes the cell type to which each sample cell belongs, wherein each sample cell belongs to a cell type, and a cell type may include multiple sample cells.
[0060] The cell data also includes receptor-ligand information from the multiple sample cells, which includes information on receptor-ligand pairs. Receptors are typically located on the cell membrane or in the cytoplasm, and receptors at different locations within the cell have different types and functions. Ligands are molecules produced within the cell that bind to receptors and trigger corresponding signal transduction. When a receptor and ligand bind, a receptor-ligand pair is formed, and there is a one-to-one correspondence between the receptor and ligand in the pair.
[0061] In one embodiment, although the cells in the plurality of sample cells are closer to the target sample cell than other unselected cells, the cell in the plurality of sample cells that has the greatest influence on the target sample cell is the nearest neighbor cell of the target sample cell. The method further includes: determining the nearest neighbor cell of the target sample cell in the plurality of sample cells.
[0062] Specifically, the k-Nearest Neighbor Algorithm can be used to determine the k nearest neighbor cells of the target sample cell by calculating the distance between each sample cell and the target sample cell. The number of nearest neighbor cells can be set according to actual needs; for example, the number of nearest neighbor cells can be proportional to the total number of sample cells. The methods used to determine the distance between each sample cell and the target sample cell include, but are not limited to, calculating the Euclidean distance or Manhattan distance between the spatial positions of each sample cell and the target sample cell.
[0063] In one embodiment, the aforementioned cellular data is used to obtain environmental factors that may affect gene expression levels in target sample cells. These environmental factors may include, but are not limited to, one or more of the following: the cell type of the target sample cell's nearest neighbor cells, the cell types to which the multiple sample cells belong, the spatial distance between the target sample cell and other sample cells, the ligand information of the multiple sample cells. By exploring the potential impact of these environmental factors on downstream gene expression in target sample cells, the influence of cell communication mechanisms on phenotypic biology and medicine research can be further explained.
[0064] S22, construct a first model based on the first data in the cell data, solve the first model using a preset solution method, and obtain the first gene expression level of the target sample cells.
[0065] In one embodiment, the first data includes the cell type of the nearest neighbor cells of the target sample cell, and the first data represents the environmental factors in the first microenvironment in which the target sample cell is located. The first model is used to fit the gene distribution in the target sample cell when it is affected by the first microenvironment. When constructing the first model for fitting the relationship between gene expression levels and the first microenvironment in which the target sample cell is located, a variety of distribution functions can be used to fit the gene distribution of the target sample cell. For example, the various distributions include, but are not limited to, Poisson distribution, Gaussian distribution, zero-inflated negative binomial distribution, etc.
[0066] In one embodiment, constructing the first model based on the first data in the cell data includes: fitting the distribution of sample genes in the target sample cells using a preset distribution function based on the gene expression level of each sample gene in the target sample cells, the cell type to which the nearest neighbor cells of the target sample cells belong, and the number of the nearest neighbor cells of the target sample cells belonging to the cell types.
[0067] Taking the Poisson distribution as an example, without considering the target sample cells... Under any influence of the microenvironment, the target sample cells Sample genes It follows the most basic Poisson distribution: That is, sample genes In target sample cells Gene expression levels Obey an expectation The Poisson distribution.
[0068] Based on the most basic Poisson distribution mentioned above, and considering the influence of the first microenvironment corresponding to the first data on the target sample genes... The influence of gene distribution and gene expression levels Expectations Things will change, therefore the process of building the first model includes expectations. The fitting and updating of the model, wherein the first model includes a generalized linear model based on the Poisson distribution.
[0069] The step of constructing the first model based on the first data in the cell data includes: fitting the distribution of sample genes of the target sample cells using a Poisson distribution, wherein the formula used includes: ,in, Indicates target sample cells Sample genes Gene expression levels, Indicates gene expression level Expectations; for sample genes ,Will Fit to Explanatory variables The combination of these formulas includes: ,in, This indicates the number of cell types belonging to the nearest neighbor cells of the target sample cell. ; and This represents the Poisson parameter.
[0070] In one embodiment, Fit to Explanatory variables At that time, explanatory variables This can include indicator variables, discrete variables, and continuous variables constructed based on the cell types of the nearest neighbor cells of the target sample cells. Among these, when the explanatory variable... When used as an indicator variable, it can include: when the nearest neighbor cell belongs to the cell type. hour, When the nearest neighbor cell does not belong to the cell type hour, When explanatory variables When the variable is discrete, it can include: Let =Nearest neighbor cells belong to cell type The quantity at which time. When the explanatory variable... When the variable is continuous, it can include: Let =Belongs to cell type Sample genes of the nearest neighbor cell The average gene expression level.
[0071] In one embodiment, due to gene expression levels Expectations of obedience It should be a non-negative value, so... The fit is as follows: .
[0072] In one embodiment, and The unknown Poisson parameter is represented by the first model, and the solution method includes solving for the Poisson parameter using the maximum likelihood method. Specifically, the formula used to solve for the Poisson parameter using the maximum likelihood method includes:
[0073] = ,
[0074] ,
[0075] in, Indicates by The vector formed by this. Where, The function is used to solve The minimum value.
[0076] In one embodiment, to avoid overfitting of the first model and difficulty in selecting feature variables, the method further includes: applying penalized regression to the solution method using a preset optimal regularization parameter. The penalized regression includes ridge regression (or L2 regularization) and lasso regression (or L1 regularization), and the formulas used include:
[0077] ,
[0078] in, This represents the preset weighting coefficient. The value range of is (0,1). This indicates the total number of cells in the sample. This represents the optimal regularization parameter. By constraining the model through penalized regression, the size or sparsity of the model parameters can be limited, making the model more concise and possessing stronger generalization ability.
[0079] In one embodiment, the method further includes: determining, based on cross-validation, a regularization parameter that results in the highest determination coefficient score for the first model from a set of preset regularization parameters as the optimal regularization parameter.
[0080] Specifically, based on cross-validation (e.g., 3-fold cross-validation), the regularization parameter that yields the highest pseudo-R2 score can be determined from a set of preset regularization parameters as the optimal regularization parameter. The formula used includes:
[0081] ,
[0082] in, This indicates the pseudo-R2 score. Indicates sample genes The actual gene expression level in cells; This represents the sample genes obtained by the first model solution when regularization parameters are included. Predicted gene expression levels; This represents the predicted gene expression level obtained by the first model without regularization parameters.
[0083] In one embodiment, the first model can be validated based on hypothesis testing, and the method further includes, for example, Figure 3 The steps for performing a significance test on the Poisson parameter are shown below:
[0084] S31, perform parameter estimation on the Poisson parameter to obtain the estimated value of the Poisson parameter.
[0085] In one embodiment, the least squares method can be used to... Perform parameter estimation to obtain The estimated value Specifically, the least squares method estimates the parameters by minimizing the squared error between the actual observed values (e.g., the actual gene expression level of the target sample gene) and the model predictions (e.g., the predicted gene expression level of the target sample gene calculated by the first model).
[0086] S32, determine the deviation between the estimated value of the Poisson parameter and the corresponding unknown true parameter, and approximate the deviation as a standard normal distribution according to the principle of asymptotic normality.
[0087] In one embodiment, an estimate of the Poisson parameter is determined. With unknown true parameters Deviation value between Based on the principle of asymptotic normality, the deviation value is... It approximates a standard normal distribution, that is: ,in, This represents the inverse Fisher information matrix, used to measure observable random variables. For unknown true parameters The amount of information.
[0088] S33, calculate the z-value and p-value of the two-tailed test based on the deviation, determine the q-value based on the p-value to control the false positive of the Poisson parameter, so that the false positive is lower than the preset false positive threshold.
[0089] In one embodiment, the z-value corresponding to the bias can be calculated using a standard normal distribution table or statistical software to obtain a two-tailed p-value, which measures the probability of the observed bias or more extreme cases occurring when the null hypothesis is true; the q-value is calculated using methods such as the Benjamini-Hochberg method, and the q-value corresponding to each p-value is determined according to the ranking of the p-values to control the false positives, for example, setting the false positive threshold to 0.05, that is, expecting 5% of the rejected false positive results to be incorrect.
[0090] In one embodiment, the method may further perform step S23 after step S21.
[0091] S23, construct a second model based on the second data in the cell data, solve the second model using the solution method, and obtain the expression level of the second gene in the target sample cells.
[0092] In one embodiment, the second data includes the cell types of the plurality of sample cells, the spatial distance between the target sample cell and other sample cells, and represents environmental factors in the second microenvironment in which the target sample cell is located; the second model is used to fit the gene distribution in the target sample cell when it is affected by the second microenvironment. Based on the first microenvironment, the second microenvironment extends the influence of nearest neighbor cells on the target sample cell to the influence of cell type and distance on all sample cells.
[0093] In one embodiment, constructing the second model based on the second data in the cell data includes: fitting the distribution of sample genes in the target sample cells using a preset distribution function based on the gene expression level of each sample gene in the target sample cells, the cell types to which the plurality of sample cells belong, the total number of cell types to which the plurality of sample cells belong, and the spatial distance between the target sample cells and other sample cells.
[0094] The formula used to construct the second model based on the second data from the cell data, building upon the first model, includes:
[0095] ,
[0096] in, Indicates target sample cells Sample genes Gene expression levels, Indicates gene expression level Expectation; error term , express variance This represents the total number of cells in the multiple sample samples. This represents the total number of cell types to which the multiple sample cells belong. Indicates target sample cells One-hot encoded vector, Indicates sample cells One-hot encoded vector, Indicates target sample cells With sample cells Distance weights, The function is used to... Transform a 3D matrix into 3D vector and Let represent the Poisson parameter. The error term represents the dispersion of the Poisson distribution due to unobserved random factors, which may include random effects related to a specific observation unit or other unknown factors. By adding the error term to the linear prediction model, the Poisson distribution characteristics of the data can be better captured, and the accuracy of the model can be improved.
[0097] In one embodiment, compared to the first model fitted based on the first data, The second model was constructed based on the second data. Among them, target sample cells One-hot encoded vector Represents a length of A vector is generated where only one element is 1, indicating that the cell belongs to the corresponding category, and all other elements are 0. Cells can be classified and labeled using one-hot encoded vectors for use in subsequent analysis and processing. Sample cells. One-hot encoded vector With one-hot encoded vector The construction method is similar and will not be described again. Furthermore, the target sample cells... With sample cells Distance weight Includes: when sample cells Target sample cells When the nearest neighbor cell, 1. Otherwise, .
[0098] In one embodiment, the second model can be considered as a concrete instantiation of the first model. Therefore, the solution method for the second model is the same as that for the first model, and the Poisson parameters of the second model can be solved using the solution method for the first model. Similarly, penalized regression, calculation of the optimal regularization parameter, significance testing, etc., can be performed on the second model. The specific steps can be referred to the relevant content of the first model, and will not be repeated here.
[0099] In one embodiment, the method may further perform step S24 after step S21.
[0100] S24, construct a third model based on the third data in the cell data, solve the third model using the solution method, and obtain the expression level of the third gene in the target sample cells.
[0101] In one embodiment, the third data includes the cell type of the plurality of sample cells, the spatial distance between the target sample cell and other sample cells, and the ligand information of the plurality of sample cells. The third data represents the environmental factors in the third microenvironment in which the target sample cell is located. The third model is used to fit the gene distribution in the target sample cell when it is affected by the third microenvironment. Based on the second microenvironment, the third microenvironment extends the influence of the target sample cell on a single sample cell to the influence of the ligand information of all sample cells.
[0102] In one embodiment, constructing a third model based on third data in the cell data includes: fitting the distribution of sample genes in the target sample cells using a preset distribution function based on the gene expression level of each sample gene in the target sample cells, the cell types to which the plurality of sample cells belong, the total number of cell types to which the plurality of sample cells belong, the spatial distance between the target sample cells and other sample cells, and the ligand information of the plurality of sample cells.
[0103] The formula used to construct the third model based on the third data from the cell data, building upon the second model, includes:
[0104] ,
[0105] in, Indicates target sample cells Sample genes Gene expression levels, Indicates gene expression level Expectation; error term , express variance This represents the total number of cells in the multiple sample samples. This represents the total number of cell types to which the multiple sample cells belong. Indicates target sample cells One-hot encoded vector, Indicates sample cells One-hot encoded vector, Indicates target sample cells With sample cells Distance weights, The function is used to... Transform a 3D matrix into 3D vector Indicates target sample cells The Middle Gene expression levels of receptors in each ligand pair Indicates target sample cells The Middle Gene expression levels of ligands in each ligand pair and This represents the Poisson parameter.
[0106] In one embodiment, compared to the first model fitted based on the first data, and the second model constructed based on the second data The third model is obtained based on the third data. Compared to the second model, the third model adds parameters related to ligand information. and .
[0107] In one embodiment, similar to the second model, the third model can be considered as a concrete instantiation of the first model. Therefore, the solution method for the third model is the same as that for the first model, and the Poisson parameters of the third model can be solved using the solution method for the first model. Similarly, penalized regression, calculation of the optimal regularization parameter, significance testing, etc., can be performed on the third model. The specific steps can be referred to the relevant content of the first model, and will not be repeated here.
[0108] In one embodiment, the method described above in this application can fit the effects of different microenvironments on cell gene expression, thereby improving the accuracy of calculating cell gene expression levels. Among them, the third gene expression level obtained by the third model can be used as the target gene expression level with the highest accuracy for the target sample cells because it integrates the most environmental factors.
[0109] Studying gene expression levels helps researchers understand gene activity levels under different conditions, thereby revealing gene function and regulatory mechanisms. Specific applications include, but are not limited to: understanding disease mechanisms: comparing gene expression differences between normal and diseased cells can identify key genes and pathways leading to disease, providing clues for disease research; predicting drug responses: analyzing gene expression levels can predict an individual's response to drug treatment, aiding in the selection and optimization of personalized drug therapy; discovering biomarkers: identifying gene expression changes associated with specific diseases or physiological states can serve as potential biomarkers for diagnosis, prognosis, and treatment monitoring; and revealing gene regulatory networks: gene expression data can be used to construct gene regulatory networks, further understanding the interactions and regulatory relationships between genes.
[0110] In one embodiment, taking the above-described method of understanding disease mechanisms based on gene expression as an example, the method may further include: comparing the gene expression levels of each gene in normal cells (e.g., normal breast cells) and disease cells (e.g., breast cancer cells); when it is determined that the gene expression level of any gene in the disease cells is significantly higher than that of the gene in normal cells and exceeds a preset gene expression level threshold, the gene is identified as a key gene involved in the proliferation and transformation of the disease cells; by identifying the intercellular signal transduction pathway involved by the key gene, the pathway is identified as an important pathway for the disease (e.g., breast cancer) corresponding to the disease cells.
[0111] In one embodiment, for example Figure 4 The diagram shown is an example flowchart of a process for modeling and analyzing gene expression based on the microenvironment of a cell, according to an embodiment of this application. The method provided in this embodiment can model the target sample cell based on the cell type of its nearest neighbor cells, and further combine cell type and ligand pair information for modeling. This allows for a comprehensive consideration of the influence of various microenvironments on downstream genes, achieving accurate analysis of gene expression.
[0112] In one embodiment, for example Figure 5 The diagram shown is an example of gene expression levels provided in an embodiment of this application. A third model can be used to calculate the gene expression levels of all genes when the target sample cell is a specific recipient cell, thus obtaining the target sample cell's... Poisson parameters ,according to OK Arrange the columns to get Figure 5 The horizontal or vertical axis in the diagram, where the horizontal axis represents the recipient cell. The vertical axis represents the recipient cell's neighboring cells (including nearest neighbor cells and other sample cells) (Sending Cell). Cell types. Figure 5 The values in the squares of the heatmap shown represent Furthermore, the number of genes whose gene expression differs significantly between ligand-recipient cell pairs is the gene expression level of the target sample cells.
[0113] In one embodiment, for example Figure 6The diagram shown is an example of gene expression levels provided in another embodiment of this application. Furthermore, gene expression studies can be conducted on target sample cells of a certain cell type based on the first model to obtain the influence of neighboring cell types on downstream gene expression of the target sample cells when the target sample cells of that cell type are used as recipient cell types or sender cell types. Figure 6 The gene expression levels shown are as follows.
[0114] In one embodiment, for example Figure 7 The diagram shown is an example of gene expression levels provided in another embodiment of this application. Furthermore, gene expression can also be studied based on a third model for a specific cell type pair composed of target sample cells of a certain cell type, to obtain the influence of the ligand pair in that specific cell type pair on the downstream gene expression of that specific cell type pair, i.e., as shown... Figure 7 The gene expression levels are shown in the figure. Among them, according to... Figure 7 It can be seen that the upregulation of integrins and growth-promoting axonogenic molecules (such as HRAS) is associated with significant cell type-specific (reaEGC-WSN) signaling.
[0115] Figure 8 This is a structural diagram of a gene expression level calculation device provided in an embodiment of this application.
[0116] In some embodiments, the gene expression level calculation device 70 may include a plurality of functional modules composed of computer program segments. The computer programs for each program segment in the gene expression level calculation device 70 may be stored in the memory of an electronic device and executed by at least one processor to perform (see details). Figure 2 (Description) Function for calculating gene expression levels.
[0117] In this embodiment, the gene expression level calculation device 70 can be divided into multiple functional modules according to its functions. These functional modules may include: a data acquisition module 701 and a model construction module 702. As used in this application, a module refers to a series of computer program segments that can be executed by at least one processor and perform a fixed function, stored in memory. In this embodiment, the functional implementation of each module in the gene expression level calculation device 70 can be found in the above description of the gene expression level calculation method, and will not be repeated here.
[0118] The data acquisition module 701 is used to acquire cell data of multiple sample cells, including target sample cells.
[0119] The model building module 702 is used to build a first model based on the first data in the cell data, solve the first model using a preset solution method and obtain the first gene expression level of the target sample cell, wherein the first data represents the environmental factors in the first microenvironment in which the target sample cell is located; the first data includes the cell type of the nearest neighbor cell of the target sample cell; the first model is used to fit the gene distribution in the target sample cell when it is affected by the first microenvironment.
[0120] The model building module 702 is further configured to build a second model based on the second data in the cell data, solve the second model using the solution method, and obtain the second gene expression level of the target sample cell, wherein the second data represents the environmental factors in the second microenvironment in which the target sample cell is located; the second data includes the cell type to which the plurality of sample cells belong and the spatial distance between the target sample cell and other sample cells; the second model is used to fit the gene distribution in the target sample cell when it is affected by the second microenvironment.
[0121] The model construction module 702 is further configured to construct a third model based on the third data in the cell data, solve the third model using the solution method, and obtain the third gene expression level of the target sample cell. The third data represents the environmental factors in the third microenvironment in which the target sample cell is located. The third data includes the cell type to which the plurality of sample cells belong, the spatial distance between the target sample cell and other sample cells, and the ligand information of the plurality of sample cells. The third model is used to fit the gene distribution in the target sample cell when it is affected by the third microenvironment.
[0122] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, and the method implemented when the program instructions are executed can refer to the methods in the above embodiments of this application.
[0123] The computer-readable storage medium can be the internal memory of the electronic device described in the above embodiments, such as the hard disk or memory of the electronic device. Alternatively, the computer-readable storage medium can be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device.
[0124] In some embodiments, the computer-readable storage medium may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application program required for at least one function, etc.; and the data storage area may store data created based on the use of the electronic device, etc.
[0125] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0126] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0127] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0128] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0129] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for calculating gene expression levels, characterized in that, The method includes: Acquire cell data from multiple sample cells, including target sample cells; Constructing a first model based on the first data in the cell data includes: fitting the distribution of sample genes in the target sample cells using a Poisson distribution, wherein the formula used includes: ,in, Indicates target sample cells Sample genes Gene expression levels, Indicates gene expression level Expectations; for sample genes ,Will Fit to Explanatory variables The combination of these formulas includes: ,in, This indicates the number of cell types belonging to the nearest neighbor cells of the target sample cell. ; and Represents the Poisson parameter; The first model is solved using a preset solution method to obtain the first gene expression level of the target sample cell, wherein the first data represents the environmental factors in the first microenvironment in which the target sample cell is located; the first data includes the cell type of the nearest neighbor cell of the target sample cell; the first model is used to fit the gene distribution of the target sample cell when it is affected by the first microenvironment.
2. The gene expression level calculation method according to claim 1, characterized in that, The solution method includes: using the maximum likelihood method to solve for the Poisson parameter.
3. The gene expression level calculation method according to claim 1, characterized in that, The method further includes: applying a penalty regression to the solution method using a preset optimal regularization parameter, wherein the penalty regression includes ridge regression and lasso regression.
4. The gene expression level calculation method according to claim 3, characterized in that, The method further includes: Based on cross-validation, the regularization parameter that gives the highest determination coefficient score to the first model is determined from a set of preset regularization parameters as the optimal regularization parameter.
5. The gene expression level calculation method according to claim 1, characterized in that, The method further includes performing a significance test on the Poisson parameter, including: The Poisson parameter is estimated to obtain the estimated value of the Poisson parameter; The deviation between the estimated value of the Poisson parameter and the corresponding unknown true parameter is determined, and the deviation is approximated as a standard normal distribution based on the principle of asymptotic normality. The z-value and p-value of the two-tailed test are calculated based on the deviation, and the q-value is determined based on the p-value to control the false positive of the Poisson parameter.
6. A method for calculating gene expression levels, characterized in that, The method includes: Acquire cell data from multiple sample cells, including target sample cells; A second model is constructed based on the second data from the cell data, using the following formulas: , in, Indicates target sample cells Sample genes Gene expression levels, Indicates gene expression level Expectation; error term , express variance This represents the total number of cells in the multiple sample samples. This represents the total number of cell types to which the multiple sample cells belong. Indicates target sample cells One-hot encoded vector, Indicates sample cells One-hot encoded vector, Indicates target sample cells With sample cells Distance weights, The function is used to... Transform a 3D matrix into 3D vector and Represents the Poisson parameter; The second model is solved using a preset solution method to obtain the second gene expression level of the target sample cell. The second data represents the environmental factors in the second microenvironment in which the target sample cell is located. The second data includes the cell type to which the plurality of sample cells belong and the spatial distance between the target sample cell and other sample cells. The second model is used to fit the gene distribution in the target sample cell when it is affected by the second microenvironment.
7. A method for calculating gene expression levels, characterized in that, The method includes: Acquire cell data from multiple sample cells, including target sample cells; A third model is constructed based on third data from the cell data, using the following formulas: , in, Indicates target sample cells Sample genes Gene expression levels, Indicates gene expression level Expectation; error term , express variance This represents the total number of cells in the multiple sample samples. This represents the total number of cell types to which the multiple sample cells belong. Indicates target sample cells One-hot encoded vector, Indicates sample cells One-hot encoded vector, Indicates target sample cells With sample cells Distance weights, The function is used to... Transform a 3D matrix into 3D vector Indicates target sample cells The Middle Gene expression levels of receptors in each ligand pair Indicates target sample cells The Middle Gene expression levels of ligands in each ligand pair and Represents the Poisson parameter; The third model is solved using a preset solution method to obtain the third gene expression level of the target sample cell. The third data represents the environmental factors in the third microenvironment in which the target sample cell is located. The third data includes the cell type to which the plurality of sample cells belong, the spatial distance between the target sample cell and other sample cells, and the ligand information of the plurality of sample cells. The third model is used to fit the gene distribution in the target sample cell when it is affected by the third microenvironment.
8. An electronic device, characterized in that, The electronic device includes a processor and a memory, the processor being configured to execute a computer program stored in the memory to implement the gene expression level calculation method as described in any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the gene expression level calculation method as described in any one of claims 1 to 7.