Gene expression quantity calculation method, electronic equipment and storage medium

By constructing multiple models to consider the influence of cellular microenvironment factors, the problem of low accuracy in the calculation of gene expression in the prior art is solved, and more accurate gene expression analysis and biological research data are achieved.

CN120072050AActive Publication Date: 2025-05-30SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311612802.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30
Estimated Expiration
2043-11-28

AI Technical Summary

Technical Problem

In the prior art, when calculating the cell gene expression level, the microenvironment influence of the cells is not considered, resulting in low calculation accuracy, affecting the accuracy of gene expression analysis and medical clinical diagnosis.

Method used

By obtaining cellular data from multiple sample cells, different models are constructed to fit the gene distribution of target sample cells in different microenvironments, including nearest neighbor cells, all peripheral cells and affected by ligand information. The Poisson distribution and other distribution functions are used for fitting, and the model parameters are solved by methods such as maximum likelihood method and penal regression.

Benefits of technology

It improves the calculation accuracy of gene expression, enables more comprehensive analysis of spatial group data, and provides reliable data to support biological research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072050A_ABST
    Figure CN120072050A_ABST
Patent Text Reader

Abstract

The invention relates to the field of biological medicine, and provides a gene expression quantity calculation method, electronic equipment and a storage medium, and the method comprises the following steps: obtaining cell data of a plurality of sample cells; constructing a first model according to first data in the cell data, and solving the first model by using a preset solving method, the first data including a cell type to which a nearest neighbor cell of the target sample cell belongs; constructing a second model according to the second data; a third model is constructed according to third data, the third model is solved through the solving method, the third gene expression quantity of the target sample cells is obtained, and the third data comprises the cell types of the multiple sample cells, the spatial distances between the target sample cells and other sample cells and ligand information of the multiple sample cells. By utilizing the method, the accurate analysis of cell gene expression can be realized in combination with the data of the microenvironment where the cells are located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical technologies, and particularly to a method for calculating gene expression levels, an electronic device, and a storage medium. Background Art

[0002] In related technologies, when studying the gene expression of cells, the characteristics of cells are usually described based on the action of ligand signaling molecules. This method assumes that the gene expressions in cells are independent of each other. Currently, there is no method that can take into account that the gene expression of cells is affected by the microenvironment in which the cells are located, resulting in relatively low accuracy in calculating the gene expression levels obtained currently.

[0003] If the accurate gene expression levels cannot be determined, it will affect the analysis of gene expression, making it impossible to clarify which genes have changed in expression, what the correlations are between genes, and how the activities of genes are affected under different conditions. Thus, it will inevitably lead to large deviations in research such as medical clinical diagnosis, drug efficacy judgment, and disease pathogenesis based on gene expression. Summary of the Invention

[0004] In view of the above, it is necessary to provide a method for calculating gene expression levels, an electronic device, and a storage medium, which can solve the problem of relatively low accuracy in calculating gene expression levels caused by the failure to consider that the gene expression of cells is affected by the microenvironment in which the cells are located.

[0005] An embodiment of this application provides a method for calculating gene expression levels. The method includes: obtaining cell data of multiple sample cells, where the multiple sample cells include target sample cells; constructing a first model based on first data in the cell data, and using a preset solution method to solve the first model and obtain a first gene expression level of the target sample cells, where the first data represents environmental factors in a first microenvironment in which the target sample cells are located; the first data includes the cell type to which the nearest neighbor cells of the target sample cells belong; the first model is used to fit the gene distribution in the target sample cells when affected by the first microenvironment. In one embodiment, constructing the first model based on the first data in the cell data includes: based on the gene expression levels of each sample gene in the target sample cells, the cell type to which the nearest neighbor cells of the target sample cells belong, and the number of the nearest neighbor cells of the target sample cells belonging to the cell type, using a preset distribution function to fit the distribution of the sample genes of the target sample cells.

[0006] In one embodiment, constructing the first model based on the first data in the cell data includes: using a Poisson distribution to fit the distribution of the sample genes of the target sample cells, and the formula used includes: y ij |λij ≈Poisson(λ ij ), where y ij represents the gene expression level of sample gene j in target sample cell i, and λ ij represents the expectation of the gene expression level y ij ; for sample gene j, λ ij is fitted as a combination of K explanatory variables x kj , and the formula used includes: where K represents the number of cell types to which the nearest neighbor cells of the target sample cells belong, k ∈ {1, 2,..., K}; β 0j and β kj represent Poisson parameters.

[0007] In one embodiment, the solving method includes: solving the Poisson parameter using the maximum likelihood method.

[0008] In one embodiment, the method further includes: performing penalty regression on the solving method using a preset optimal regularization parameter, and the penalty regression includes ridge regression and lasso regression.

[0009] In one embodiment, the method further includes: based on the cross-validation method, determining, from a preset plurality of regularization parameters, the regularization parameter that makes the coefficient of determination score of the first model the highest as the optimal regularization parameter.

[0010] In one embodiment, the method further includes performing a significance test on the Poisson parameter, including: performing parameter estimation on the Poisson parameter to obtain an estimated value of the Poisson parameter; determining the deviation between the estimated value of the Poisson parameter and the corresponding unknown true parameter, and approximating the deviation as a standard normal distribution according to the principle of asymptotic normality; calculating the z-value and p-value of the two-tailed test based on the deviation, and controlling the false positive of the Poisson parameter based on the p-value.

[0011] An embodiment of the present application provides a method for calculating gene expression levels, the method including: obtaining cell data of a plurality of sample cells, the plurality of sample cells including target sample cells; constructing a second model according to second data in the cell data, and using a preset solving method to solve the second model and obtain the second gene expression level of the target sample cells, where the second data represents environmental factors in a second microenvironment where the target sample cells are located; the second data includes the cell types to which the plurality of sample cells belong, and the spatial distance between the target sample cells and other sample cells; the second model is used to fit the gene distribution in the target sample cells when affected by the second microenvironment.

[0012] In one embodiment, constructing a second model based on the second data in the cell data includes: fitting the distribution of the sample genes of the target sample cell by using a preset distribution function based on the gene expression levels of each sample gene in the target sample cell, the cell types to which the multiple sample cells belong, the total number of the cell types to which the multiple sample cells belong, and the spatial distance between the target sample cell and other sample cells.

[0013] In one embodiment, the formula used for constructing a second model according to the second data in the cell data includes:

[0014]

[0015] where y ij represents the gene expression level of sample gene j in target sample cell i, λ ij represents the expectation of gene expression level y ij ; the error term represents ∈ ij variance, n represents the total number of the multiple sample cells, U represents the total number of the cell types to which the multiple sample cells belong, c iU represents the one-hot encoding vector of target sample cell i, c hU represents the one-hot encoding vector of sample cell h, ω ih represents the distance weight between target sample cell i and sample cell h, the vec() function is used to convert a U×U-dimensional matrix into a 1×U 2 -dimensional vector, β 0j and β g represent Poisson parameters.

[0016] An embodiment of the present application provides a method for calculating gene expression levels. The method includes: obtaining cell data of multiple sample cells, where the multiple sample cells include a target sample cell; constructing a third model according to the third data in the cell data, and using a preset solution method to solve the third model and obtain the third gene expression level of the target sample cell, where the third data represents environmental factors in a third microenvironment where the target sample cell is located; the third data includes the cell types to which the multiple sample cells belong, the spatial distance between the target sample cell and other sample cells, and the ligand information of the multiple sample cells; the third model is used to fit the gene distribution in the target sample cell when affected by the third microenvironment.

[0017] In one embodiment, constructing the third model based on the third data in the cell data includes: fitting the distribution of the sample genes of the target sample cells by using a preset distribution function based on the gene expression levels of each sample gene in the target sample cells, the cell types to which the multiple sample cells belong, the total number of cell types to which the multiple sample cells belong, the spatial distances between the target sample cells and other sample cells, and the ligand-receptor information of the multiple sample cells.

[0018] In one embodiment, the formula used for constructing the third model according to the third data in the cell data includes:

[0019]

[0020] where y ij represents the gene expression level of sample gene j in target sample cell i, λ ij represents the expectation of the gene expression level y ij ; the error term represents ∈ ij variance, n represents the total number of the multiple sample cells, U represents the total number of cell types to which the multiple sample cells belong, c iU represents the one-hot encoding vector of target sample cell i, c hU represents the one-hot encoding vector of sample cell h, ω ih represents the distance weight between target sample cell i and sample cell h, the vec() function is used to convert a U×U-dimensional matrix into a 1×U 2 -dimensional vector, R if represents the gene expression level of the receptor in the f-th ligand-receptor pair in target sample cell i, L if represents the gene expression level of the ligand in the f-th ligand-receptor pair in target sample cell i, β 0j and β fg represent Poisson parameters.

[0021] Embodiments of the present application provide a gene expression level calculation device, which includes: a data acquisition module for acquiring cell data of multiple sample cells, where the multiple sample cells include target sample cells; a model construction module for constructing a first model based on first data in the cell data, solving the first model using a preset solution method, and obtaining the first gene expression level of the target sample cells, where the first data represents environmental factors in a first microenvironment where the target sample cells are located; the first data includes the cell types to which the nearest neighbor cells of the target sample cells belong; the first model is used to fit the gene distribution in the target sample cells when affected by the first microenvironment; the model construction module is further configured to construct a second model based on second data in the cell data, solve the second model using the solution method, and obtain the second gene expression level of the target sample cells, where the second data represents environmental factors in a second microenvironment where the target sample cells are located; the second data includes the cell types to which the multiple sample cells belong, and the spatial distance between the target sample cells and other sample cells; the second model is used to fit the gene distribution in the target sample cells when affected by the second microenvironment; the model construction module is further configured to construct a third model based on third data in the cell data, solve the third model using the solution method, and obtain the third gene expression level of the target sample cells, where the third data represents environmental factors in a third microenvironment where the target sample cells are located; the third data includes the cell types to which the multiple sample cells belong, the spatial distance between the target sample cells and other sample cells, and the ligand information of the multiple sample cells; the third model is used to fit the gene distribution in the target sample cells when affected by the third microenvironment.

[0022] Embodiments of the present application provide an electronic device, which includes a processor and a memory. The processor is configured to implement the gene expression level calculation method when executing a computer program stored in the memory.

[0023] Embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, implements the gene expression level calculation method.

[0024] In summary, the gene expression level calculation method described in this application can take into account the influence of various factors (such as neighboring cell and ligand-receptor information) in multiple microenvironments where the target sample cells are located. By constructing different models to simulate the effects of various factors in the microenvironment on the gene expression level of the target sample cells, it can make full use of and explore spatial transcriptomic data, comprehensively and deeply achieve the interaction analysis of spatial transcriptomic data, improve the calculation accuracy of gene expression levels, and thus provide reliable data for biological research based on gene expression levels. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a structural diagram of an electronic device provided by an embodiment of this application.

[0026] Figure 2 is a flowchart of a gene expression level calculation method provided by an embodiment of this application.

[0027] Figure 3 is a flowchart of performing a significance test provided by an embodiment of this application.

[0028] Figure 4 is a flow diagram example of modeling and gene expression analysis based on the microenvironment where cells are located provided by an embodiment of this application.

[0029] Figure 5 is an example diagram of gene expression levels provided by an embodiment of this application.

[0030] Figure 6 is an example diagram of gene expression levels provided by another embodiment of this application.

[0031] Figure 7 is an example diagram of gene expression levels provided by yet another embodiment of this application.

[0032] Figure 8 is a structural diagram of a gene expression level calculation device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] In order to more clearly understand the above objects, features, and advantages of this application, the following describes this application in detail with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of this application and the features in the embodiments can be combined with each other.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the specification of this application herein are only for the purpose of describing an embodiment and are not intended to limit this application.

[0035] It should be noted that in this application, "at least one" means one or more, and "a plurality" means two or more than two. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and drawings of this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0036] In the embodiments of this application, words such as "exemplary" or "for example" are used to give examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0037] In one embodiment, in the related art, when studying the gene expression of cells, the cell characteristics are usually described based on the action of ligand signaling molecules. This method defaults that the gene expressions in cells are independent of each other. Currently, there is no method that can take into account that the gene expression of cells will be affected by the microenvironment in which the cells are located, resulting in a relatively low calculation accuracy of the currently obtained gene expression levels.

[0038] If the accurate gene expression level cannot be determined, it will affect the analysis of gene expression, and it will not be possible to clarify which genes have changed in expression, what the correlation is between genes, and how the activities of genes are affected under different conditions. Thus, it will inevitably lead to relatively large deviations in research such as medical clinical diagnosis, drug efficacy judgment, and disease occurrence mechanism based on gene expression.

[0039] To solve the above problems, an embodiment of the present application provides a method for calculating gene expression levels. By obtaining cell data of multiple sample cells including target sample cells, a first model is constructed based on the cell types to which the nearest neighbor cells of the target sample cells in the cell data belong, and a preset solution method is used to solve the first model to obtain a first gene expression level of the target sample cells affected by the nearest neighbor cells in the microenvironment where they are located; a second model is constructed based on the cell types to which the multiple sample cells belong, the spatial distances between the target sample cells and other sample cells, and the second gene expression level of the target sample cells affected by all surrounding cells in the microenvironment where they are located is obtained by using the solution method to solve the second model; a third model is constructed based on the cell types to which the multiple sample cells belong, the spatial distances between the target sample cells and other sample cells, and the ligand-receptor information of the multiple sample cells, and the third gene expression level of the target sample cells affected by the ligand-receptor information of other cells in the microenvironment where they are located is obtained by using the solution method to solve the third model.

[0040] The present application can take into account the influence of various factors (such as neighboring cells and ligand-receptor information) in multiple microenvironments where the target sample cells are located. By constructing different models to simulate the effects of various factors in the microenvironment on the gene expression levels of the target sample cells, the spatial transcriptome data can be fully utilized and mined, and the interaction analysis of the spatial transcriptome data can be comprehensively and deeply realized, improving the calculation accuracy of the gene expression levels and thus providing reliable data for biological research based on gene expression levels.

[0041] Figure 1 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 10 can be a mobile phone, a tablet computer, a smart wearable device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, a netbook, an energy storage device, a power distribution device, a vehicle-mounted device, a self-mobile device, and other electronic devices. The embodiment of the present application does not impose any restrictions on the specific type of the electronic device.

[0042] As Figure 1 shown, the electronic device 10 may include a communication module 101, a memory 102, a processor 103, an input / output (I / O) interface 104, and a bus 105. The processor 103 is respectively coupled to the communication interface 101, the memory 102, and the I / O interface 104 through the bus 105.

[0043] The communication module 101 may include a wired communication module and / or a wireless communication module. The wired communication module may provide one or more of the solutions for wired communication such as the universal serial bus (USB), Controller Area Network (CAN), etc. The wireless communication module may provide one or more of the solutions for wireless communication such as wireless fidelity (Wi-Fi), bluetooth (BT), mobile communication network, frequency modulation (FM), near field communication (NFC), infrared (IR), etc.

[0044] The memory 102 may include one or more random access memories (RAM) and one or more non-volatile memories (NVM). The random access memory can be directly read and written by the processor 103, can be used to store the operating system or the executable programs (such as machine instructions) of other running programs, and can also be used to store user and application data, etc. The random access memory may include static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc.

[0045] The non-volatile memory can also store executable programs and store user and application data, etc., and can be pre-loaded into the random access memory for the processor 110 to directly read and write. The non-volatile memory may include disk storage devices, flash memory.

[0046] The memory 102 is used to store one or more computer programs. The one or more computer programs are configured to be executed by the processor 103. The one or more computer programs include a plurality of instructions, and when the plurality of instructions are executed by the processor 103, a method for calculating gene expression levels executed on the electronic device 10 can be realized.

[0047] In other embodiments, the electronic device 10 further includes an external memory interface for connecting to an external memory to expand the storage capacity of the electronic device 10.

[0048] The processor 103 may include one or more processing units. For example, the processor 103 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0049] The processor 103 provides computing and control capabilities. For example, the processor 103 is used to execute the computer program stored in the memory 102 to implement the above gene expression level calculation method.

[0050] The I / O interface 104 is used to provide a channel for user input or output. For example, the I / O interface 104 can be used to connect various input and output devices, such as a mouse, a keyboard, a touch device, a display screen, etc., so that the user can input information or visualize the information.

[0051] The bus 105 is at least used to provide a communication channel between the communication module 101, the memory 102, the processor 103, and the I / O interface 104 in the electronic device 10.

[0052] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 10. In other embodiments of the present application, the electronic device 10 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0053] Figure 2 is a flowchart of a gene expression level calculation method provided by an embodiment of the present application. The gene expression level calculation method is applied to an electronic device, such as Figure 1 the electronic device 10 in, and specifically includes the following steps. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0054] S21, obtaining cell data of multiple sample cells.

[0055] In one embodiment, the plurality of sample cells can be selected according to actual applications. For example, the plurality of sample cells can be cells in a brain tissue section of a salamander on the second day after brain injury. Specifically, the plurality of sample cells include target sample cells. In order to analyze the gene expression levels of each cell in the plurality of sample cells, any sample cell in the plurality of sample cells can be used as the target sample cell, so as to calculate and analyze the gene expression levels affected by the microenvironment of the target sample cell.

[0056] In one embodiment, cell data of the plurality of sample cells can be obtained by using spatial transcriptomics technology based on chip technology. Specifically, the spatial transcriptomics technology is a technology based on single-cell ribonucleic acid (RNA) sequencing, which can correspond the transcriptome information of a single cell with its spatial position in the tissue. For example, the spatial transcriptomics technology may include the following implementation processes: (1) Tissue sectioning: cutting the tissue (or organ, embryo, etc.) into thin slices and fixing them on a glass slide; (2) RNA extraction: extracting RNA from each tissue section; (3) Library preparation: converting the RNA of each single cell into complementary DNA (cDNA) and amplifying it through reverse transcription and library preparation; (4) Chip loading: uniformly loading the sample into the chip wells; (5) Image acquisition: taking images of each chip spot using a high-resolution microscope; (6) Data processing: combining the image data with the single-cell RNA sequencing data using professional software to obtain the gene expression data of each single cell and its spatial position (cell coordinate) data in the tissue or section.

[0057] The preamble steps of the gene expression level calculation method provided by the embodiments of the present application include the above steps (1)-(6). As can be seen from the above spatial transcriptomics technology, based on the spatial transcriptomics technology, gene expression data and spatial position data of multiple sample cells in one or more adjacent chips can be obtained as the cell data.

[0058] Among them, the gene expression data may include genes of the sample cells and their corresponding gene expression levels. Preset types of sample genes and their corresponding gene expression levels can also be selected from the genes of the sample cells as the cell data. The preset types can be genes of interest selected according to actual applications. For example, genes with spatial characteristics (such as genes with morani > 0.4), genes with a relatively high positive rate of gene expression (such as genes with a positive rate greater than 5%), etc. The spatial position data may include two-dimensional or three-dimensional spatial position coordinates of each sample cell.

[0059] In addition, the cell data further includes the cell type to which each sample cell belongs, where each sample cell belongs to one cell type, and one cell type may include multiple sample cells.

[0060] The cell data further includes ligand-receptor information among the multiple sample cells, and the ligand-receptor information includes information on ligand-receptor pairs composed of a receptor and a ligand. Among them, receptors are usually present on the cell membrane or in the cytoplasm of cells, and receptors located at different positions in the cell have different types and functions; ligands include molecules produced inside the cell that bind to receptors and trigger corresponding signal transduction; when a receptor binds to a ligand, a ligand-receptor pair is formed between the receptor and the ligand, and there is a one-to-one correspondence between the receptor and the ligand in the ligand-receptor pair.

[0061] In one embodiment, although the cells among the multiple sample cells are closer to the target sample cell compared to other unselected cells, the cell that has the greatest impact on the target sample cell among the multiple sample cells is the nearest neighbor cell of the target sample cell, and the method further includes: determining the nearest neighbor cell of the target sample cell among the multiple sample cells.

[0062] Specifically, the k-nearest neighbor algorithm can be used to determine the k sample cells closest to the target sample cell as the nearest neighbor cells of the target sample cell by calculating the distance between each sample cell and the target sample cell. Among them, the number of nearest neighbor cells can be set according to actual needs. For example, the number of nearest neighbor cells can be made proportional to the total number of the multiple sample cells; the methods used to determine the distance between each sample cell and the target sample cell include but are not limited to: calculating the Euclidean distance or Manhattan distance of the spatial positions of each sample cell and the target sample cell.

[0063] In one embodiment, the above cell data is for obtaining environmental factors that may affect the gene expression level of the target sample cell. For example, the environmental factors may include but are not limited to one or more of the following: the cell type of the nearest neighbor cell of the target sample cell, the cell types to which the multiple sample cells belong, the spatial distance between the target sample cell and other sample cells, the cell types to which the multiple sample cells belong, the spatial distance between the target sample cell and other sample cells, the ligand-receptor information of the multiple sample cells. By exploring the potential impact of the above environmental factors on the downstream gene expression of the target sample cell, the impact of the cell communication mechanism on research such as phenotypic biology and medicine can be further explained.

[0064] S22. Construct a first model based on the first data in the cell data, and use a preset solution method to solve the first model and obtain the first gene expression level of the target sample cell.

[0065] In one embodiment, the first data includes the cell types to which the nearest neighbor cells of the target sample cell belong, and the first data represents the environmental factors in the first microenvironment where the target sample cell is located. The first model is used to fit the gene distribution in the target sample cell when affected by the first microenvironment. When constructing a first model for fitting the relationship between gene expression level and the first microenvironment where the target sample cell is located, the distribution functions of multiple distributions can be used to fit the gene distribution of the target sample cell. For example, the multiple distributions include, but are not limited to: Poisson distribution, Gaussian distribution, zero-inflated negative binomial distribution, etc.

[0066] In one embodiment, the constructing the first model according to the first data in the cell data includes: based on the gene expression level of each sample gene in the target sample cell, the cell types to which the nearest neighbor cells of the target sample cell belong, and the number of the cell types to which the nearest neighbor cells of the target sample cell belong, use a preset distribution function to fit the distribution of the sample genes of the target sample cell.

[0067] Taking the use of Poisson distribution as an example, without considering the influence of any microenvironment on the target sample cell i, the sample gene j in the target sample cell i follows the most basic Poisson distribution: That is, the gene expression level y of the sample gene j in the target sample cell i ij follows a Poisson distribution with an expectation of

[0068] Based on the above most basic Poisson distribution, considering the influence of the first microenvironment corresponding to the first data on the gene distribution in the target sample gene i, the expectation ij of the gene expression level y will change. Therefore, the process of constructing the first model includes fitting and updating the expectation , and the first model includes a generalized linear model based on Poisson distribution.

[0069] The constructing the first model according to the first data in the cell data includes: using Poisson distribution to fit the distribution of the sample genes of the target sample cell, and the formula used includes: y ij |λ ij ≈ Poisson(λ ij ), where y ij ​represents the gene expression level of sample gene j in target sample cell i, λ ij represents the expected value of gene expression level y ij ; for sample gene j, λ ij is fitted as a combination of K explanatory variables x kj . The formula used includes: where K represents the number of cell types to which the nearest neighbor cells of the target sample cell belong, k ∈ {1, 2, …, K}; β 0j and β kj represent Poisson parameters.

[0070] In one embodiment, when λ ij is fitted as a combination of K explanatory variables x kj , the explanatory variable x kj can include indicator variables, discrete variables, continuous variables, etc. constructed based on the cell types of the nearest neighbor cells of the target sample cell. Among them, when the explanatory variable x kj is an indicator variable, it can include: when the nearest neighbor cell belongs to cell type k, x kj = 1, and when the nearest neighbor cell does not belong to cell type k, x kj = 0. When the explanatory variable x kj is a discrete variable, it can include: let x kj = the number of times the nearest neighbor cell belongs to cell type k. When the explanatory variable x kj is a continuous variable, it can include: let x kj = the average value of the gene expression levels of sample gene j in the nearest neighbor cells belonging to cell type k.

[0071] In one embodiment, since the expected value λ ij that the gene expression level y ij obeys should be non - negative, λ ij is fitted as:

[0072] In one embodiment, β 0j and β kj represent unknown Poisson parameters. The solution method of the first model includes using the maximum likelihood method to solve the Poisson parameters. Specifically, when using the maximum likelihood method to solve the Poisson parameters, the formula used includes:

[0073]

[0074]

[0075] where β represents the vector composed of β kj . Among them, function is used to solve L(β 0j, the minimum value of β).

[0076] In one embodiment, to avoid overfitting of the first model and difficulty in selecting feature variables, the method further includes: performing penalized regression on the solution method using a preset optimal regularization parameter, where the penalized regression includes Ridge regression (Ridge regression or L2 regularization) and Lasso regression (Lasso regression or L1 regularization), and the formulas used include:

[0077]

[0078] where α represents a preset weight coefficient, the value range of α is (0, 1), n represents the total number of sample cells, and η represents the optimal regularization parameter. By constraining the model through penalized regression, the size or sparsity of the model parameters can be restricted, making the model more concise and having stronger generalization ability.

[0079] In one embodiment, the method further includes: based on the cross-validation method, determining the regularization parameter that maximizes the coefficient of determination score of the first model from a preset plurality of regularization parameters as the optimal regularization parameter.

[0080] Specifically, based on the cross-validation method (such as the 3-fold cross-validation method), the regularization parameter that maximizes the pseudo-R2 score can be determined from a preset plurality of regularization parameters as the optimal regularization parameter, and the formulas used include:

[0081]

[0082] where R 2 represents the pseudo-R2 score, L(β 0j , β) True represents the actual gene expression level of sample gene j in the cell; L(β 0j , β) Fit represents the predicted gene expression level of sample gene j obtained by solving the first model when including the regularization parameter; L(β 0j , β) Null represents the predicted gene expression level obtained by solving the first model without including the regularization parameter.

[0083] In one embodiment, the first model can be model-validated based on the method of hypothesis testing, and the method further includes the following steps for performing a significance test on the Poisson parameter as shown in Figure 3 :

[0084] S31, performing parameter estimation on the Poisson parameter to obtain an estimated value of the Poisson parameter.

[0085] In one embodiment, the least squares method can be used for parameter estimation of β kj to obtain the estimated value of β kj Specifically, the least squares method obtains the estimated value of the parameter by minimizing the squared error between the actual observed values (such as the actual gene expression level of the target sample gene) and the model predicted values (such as the predicted gene expression level of the target sample gene calculated by the first model). Specifically, the least squares method obtains the estimated value of the parameter by minimizing the squared error between the actual observed values (such as the actual gene expression level of the target sample gene) and the model predicted values (such as the predicted gene expression level of the target sample gene calculated by the first model).

[0086] S32. Determine the deviation between the estimated value of the Poisson parameter and the corresponding unknown true parameter, and approximate the deviation as a standard normal distribution according to the principle of asymptotic normality.

[0087] In one embodiment, determine the estimated value of the Poisson parameter and the deviation value from the unknown true parameter β kg According to the principle of asymptotic normality, approximate the deviation value as a standard normal distribution, that is: wherein, represents the inverse Fisher information matrix, which is used to measure the amount of information of the observable random variable y about the unknown true parameter β represents the inverse Fisher information matrix, which is used to measure the amount of information of the observable random variable y about the unknown true parameter β kg of the amount of information.

[0088] S33. Calculate the z-value and p-value of the two-tailed test according to the deviation, and determine the q-value based on the p-value to control the false positive of the Poisson parameter, so that the false positive is lower than the preset false positive threshold.

[0089] In one embodiment, a standard normal distribution table or statistical software can be used to calculate the z-value corresponding to the deviation and obtain the two-tailed p-value. The two-tailed p-value measures the probability of observing the deviation or more extreme cases when the null hypothesis is true; methods such as Benjamini-Hochberg are used to calculate the q-value, and the q-value corresponding to each p-value is determined according to the ranking of the p-values to control the false positive. For example, the false positive threshold is set to 0.05, that is, it is expected that 5% of the rejected false positive results are incorrect.

[0090] In one embodiment, the method may further execute step S23 after step S21.

[0091] S23. Construct a second model according to the second data in the cell data, and use the solution method to solve the second model and obtain the second gene expression level of the target sample cell.

[0092] In one embodiment, the second data includes the cell types to which the multiple sample cells belong, and the spatial distances between the target sample cell and other sample cells. The second data represents the environmental factors in the second microenvironment where the target sample cell is located. The second model is used to fit the gene distribution in the target sample cell when affected by the second microenvironment. Based on the first microenvironment, the second microenvironment extends the influence of the nearest neighbor cells on the target sample cell to the influence of the cell types and distances of all sample cells.

[0093] In one embodiment, constructing the second model according to the second data in the cell data includes: based on the gene expression level of each sample gene in the target sample cell, the cell types to which the multiple sample cells belong, the total number of the cell types to which the multiple sample cells belong, and the spatial distances between the target sample cell and other sample cells, using a preset distribution function to fit the distribution of the sample genes of the target sample cell.

[0094] Based on the first model, the formula used to construct the second model according to the second data in the cell data includes:

[0095]

[0096] where, y ij represents the gene expression level of sample gene j in target sample cell i, λ ij represents the expectation of the gene expression level y ij ; the error term represents ∈ ij variance, n represents the total number of the multiple sample cells, U represents the total number of the cell types to which the multiple sample cells belong, c iU represents the one-hot encoding vector of target sample cell i, c hU represents the one-hot encoding vector of sample cell h, ω ih represents the distance weight between target sample cell i and sample cell h, the vec() function is used to convert a U×U-dimensional matrix into a 1×U 2 -dimensional vector, β 0j and β g represent Poisson parameters. Among them, the error term is used to represent the dispersion of the Poisson distribution caused by unobserved random factors. The unobserved random factors may include random effects related to specific observation units or other unknown factors. By adding the error term to the linear prediction model, the Poisson distribution characteristics of the data can be better captured, and the accuracy of the model can be improved.

[0097] In one embodiment, compared with the fitting according to the first data in the first model In the second model, it is constructed based on the second data Among them, the one-hot encoded vector c of the target sample cell i iU represents a vector of length U, where only one element is 1, indicating that the cell belongs to the corresponding category, and the other elements are all 0. Through the one-hot encoded vector, cells can be classified and labeled for use in subsequent analysis and processing. The one-hot encoded vector c of the sample cell h hU has a similar construction method to the one-hot encoded vector c iU and will not be described repeatedly. In addition, the distance weight ω between the target sample cell i and the sample cell h ih includes: when the sample cell h is the nearest neighbor cell of the target sample cell i, ω ih = 1, otherwise, ω ih = 0.

[0098] In one embodiment, the second model can be regarded as a specific instantiated model of the first model. Therefore, the solution method of the second model is the same as that of the first model, and the Poisson parameter of the second model can be solved by using the solution method of the first model. Similarly, penalty regression, calculation of the optimal regularization parameter, significance test, etc. can be performed on the second model. The specific steps can refer to the relevant content of the first model and will not be described repeatedly.

[0099] In one embodiment, the method can further perform step S24 after step S21.

[0100] S24, construct a third model based on the third data in the cell data, and use the solution method to solve the third model and obtain the third gene expression level of the target sample cell.

[0101] In one embodiment, the third data includes the cell types to which the multiple sample cells belong, the spatial distances between the target sample cell and other sample cells, and the ligand-receptor information of the multiple sample cells. The third data represents the environmental factors in the third microenvironment where the target sample cell is located; the third model is used to fit the gene distribution in the target sample cell when affected by the third microenvironment. Based on the second microenvironment, the third microenvironment extends the influence of the target sample cell by a single sample cell to the influence of the ligand-receptor information in all sample cells.

[0102] In one embodiment, constructing the third model according to the third data in the cell data includes: based on the gene expression level of each sample gene in the target sample cell, the cell type to which the multiple sample cells belong, the total number of cell types to which the multiple sample cells belong, the spatial distance between the target sample cell and other sample cells, and the ligand-receptor information of the multiple sample cells, using a preset distribution function to fit the distribution of the sample genes of the target sample cell.

[0103] Based on the second model, the formula used to construct the third model according to the third data in the cell data includes:

[0104]

[0105] Where y ij represents the gene expression level of sample gene j in target sample cell i, and λ ij represents the expectation of the gene expression level y ij ; the error term represents ∈ ij variance, n represents the total number of the multiple sample cells, U represents the total number of cell types to which the multiple sample cells belong, c iU represents the one-hot encoding vector of target sample cell i, c hU represents the one-hot encoding vector of sample cell h, ω ih represents the distance weight between target sample cell i and sample cell h, the vec() function is used to convert a U×U-dimensional matrix into a 1×U 2 -dimensional vector, R if represents the gene expression level of the receptor in the fth ligand-receptor pair in target sample cell i, L if represents the gene expression level of the ligand in the fth ligand-receptor pair in target sample cell i, β 0j and β fg represent Poisson parameters.

[0106] In one embodiment, compared with the fitted according to the first data in the first model and the constructed according to the second data in the second model, the obtained according to the third data in the third model. Among them, compared with the second model, the third model adds parameters R if and L if related to ligand-receptor information.

[0107] In one embodiment, like the second model, the third model can be regarded as a specific instantiated model of the first model. Therefore, the solution method of the third model is the same as that of the first model, and the solution method of the first model can be used to solve the Poisson parameter of the third model. Similarly, penalty regression, calculation of the optimal regularization parameter, significance test, etc. can be performed on the third model. For the specific steps, reference can be made to the relevant content of the first model and will not be described repeatedly.

[0108] In one embodiment, the above method of the present application can fit to obtain the influence of different microenvironments on the gene expression of cells, so as to improve the accuracy of calculating the gene expression level of cells. Among them, since the third gene expression level obtained by the third model incorporates the most environmental factors, it can be used as the target gene expression level with the highest accuracy for the target sample cells.

[0109] Studying gene expression levels can help researchers understand the activity levels of genes under different conditions, thereby revealing gene functions and regulatory mechanisms. The specific uses can include, but are not limited to: understanding the disease occurrence mechanism: By comparing the gene expression differences between normal cells (such as normal breast cells) and diseased cells (such as breast cancer cells), key genes and pathways leading to the disease can be discovered, providing clues for disease research; predicting drug responses: By analyzing gene expression levels, the response of an individual to drug treatment can be predicted, which helps in the selection and optimization of personalized drug treatment; discovering biomarkers: Searching for gene expression changes related to specific diseases or physiological states can be used as potential biomarkers for diagnosis, prognosis, and treatment monitoring; revealing gene regulatory networks: Gene expression level data can be used to construct gene regulatory networks to further understand the interactions and regulatory relationships between genes.

[0110] In one embodiment, taking the example of understanding the disease occurrence mechanism based on gene expression as above, the method may further include: comparing the gene expression levels of each gene in normal cells (such as normal breast cells) and diseased cells (such as breast cancer cells). When it is determined that the gene expression level of any gene in the diseased cells is significantly higher than that of the same gene in the normal cells and exceeds a preset gene expression level threshold, it is determined that this gene is a key gene involved in the proliferation and transformation of the diseased cells; by determining the signal transduction pathway between cells involved in this key gene, it is determined that this pathway is an important pathway for the disease (such as breast cancer) corresponding to the diseased cells.

[0111] In one embodiment, for example Figure 4As shown, it is a flowchart example of modeling and gene expression analysis according to the microenvironment where cells are located provided by an embodiment of the present application. The method provided by the embodiment of the present application can model based on the cell types to which the nearest neighbor cells of the target sample cells belong, and further combine the cell types and ligand-receptor pair information for modeling, so as to comprehensively consider the influence of various microenvironments on downstream genes and achieve accurate analysis of gene expression.

[0112] In one embodiment, for example Figure 5 As shown, it is an example diagram of gene expression levels provided by an embodiment of the present application. Among them, a third model can be used to calculate the gene expression levels of all genes when the target sample cell is a certain receptor cell, and U×U Poisson parameters β of the target sample cell are obtained g , arranged in U rows and U columns to obtain Figure 5 the horizontal axis or vertical axis in, where the horizontal axis in English represents U cell types of the receiving cells, and the vertical axis in English represents U cell types of the neighboring cells (other sample cells including the nearest neighbor cells) of the receiving cells (sending cells). Figure 5 The value in the square of the heat map shown represents β g >0 and the number of genes with significant differences in gene expression in the ligand-receptor cell pair, that is, the gene expression level of the target sample cell.

[0113] In one embodiment, for example Figure 6 As shown, it is an example diagram of gene expression levels provided by another embodiment of the present application. Among them, the first model can also be used to study the gene expression of the target sample cells of a certain cell type, and obtain the influence of the cell types of the neighboring cells on the downstream gene expression of the target sample cells of this cell type when the target sample cells of this cell type are the receiving cell type or the sending cell type, that is, as Figure 6 shown in the gene expression levels.

[0114] In one embodiment, for example Figure 7 As shown, it is an example diagram of gene expression levels provided by yet another embodiment of the present application. Among them, the third model can also be used to study the gene expression of a specific cell type pair composed of the target sample cells of a certain cell type, and obtain the influence of the ligand-receptor pair in this specific cell type pair on the downstream gene expression of this specific cell type pair, that is, as Figure 7 shown in the gene expression levels. Among them, according to Figure 7 it can be seen that the upregulation of integrin and growth-promoting axonogenic molecules (such as HRAS) is related to significant cell type-specific (reaEGC-WSN) signals.

[0115] Figure 8 It is a structural diagram of a gene expression level calculation device provided by an embodiment of the present application.

[0116] In some embodiments, the gene expression level calculation device 7 may include a plurality of functional modules composed of computer program segments. The computer programs of each program segment in the gene expression level calculation device 7 may be stored in the memory of the electronic device and executed by at least one processor to perform the function of calculating the gene expression level (see details in Figure 2 description).

[0117] In this embodiment, the gene expression level calculation device 7 can be divided into a plurality of functional modules according to the functions it performs. The functional modules may include: a data acquisition module 701, a model construction module 702. A module referred to in the present application means a series of computer program segments that can be executed by at least one processor and can complete a fixed function, and is stored in the memory. In this embodiment, for the implementation manners of the functions of each module in the gene expression level calculation device 7, reference can be made to the limitations on the gene expression level calculation method above, and details will not be repeated here.

[0118] The data acquisition module 701 is used to acquire cell data of a plurality of sample cells, and the plurality of sample cells include target sample cells.

[0119] The model construction module 702 is used to construct a first model according to the first data in the cell data, solve the first model by using a preset solution method, and obtain the first gene expression level of the target sample cell, where the first data represents environmental factors in the first microenvironment where the target sample cell is located; the first data includes the cell type to which the nearest neighbor cells of the target sample cell belong; the first model is used to fit the gene distribution in the target sample cell when affected by the first microenvironment.

[0120] The model construction module 702 is further used to construct a second model according to the second data in the cell data, solve the second model by using the solution method, and obtain the second gene expression level of the target sample cell, where the second data represents environmental factors in the second microenvironment where the target sample cell is located; the second data includes the cell types to which the plurality of sample cells belong, and the spatial distance between the target sample cell and other sample cells; the second model is used to fit the gene distribution in the target sample cell when affected by the second microenvironment.

[0121] The model construction module 702 is further configured to construct a third model based on third data in the cell data, solve the third model by using the solving method, and obtain the third gene expression level of the target sample cell, where the third data represents environmental factors in a third microenvironment where the target sample cell is located; the third data includes the cell type to which the multiple sample cells belong, the spatial distance between the target sample cell and other sample cells, and the ligand-receptor information of the multiple sample cells; and the third model is used to fit the gene distribution in the target sample cell when affected by the third microenvironment.

[0122] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions, and the method implemented when the program instructions are executed may refer to the methods in the above various embodiments of the present application.

[0123] Among them, the computer-readable storage medium may be an internal memory of the electronic device described in the above embodiment, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0124] In some embodiments, the computer-readable storage medium may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the electronic device.

[0125] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0126] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0127] In the embodiments provided in the present application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0128] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0129] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for calculating gene expression levels, characterized in that, the method includes: obtaining cell data of multiple sample cells, where the multiple sample cells include target sample cells; constructing a first model based on first data in the cell data, and using a preset solution method to solve the first model and obtain the first gene expression level of the target sample cells, where the first data represents environmental factors in a first microenvironment where the target sample cells are located; the first data includes the cell types to which the nearest neighbor cells of the target sample cells belong; the first model is used to fit the gene distribution in the target sample cells when affected by the first microenvironment.

2. The method for calculating gene expression levels according to claim 1, characterized in that, the constructing of the first model based on the first data in the cell data includes: fitting the distribution of sample genes of the target sample cells by using a preset distribution function based on the gene expression levels of each sample gene in the target sample cells, the cell types to which the nearest neighbor cells of the target sample cells belong, and the number of the cell types to which the nearest neighbor cells of the target sample cells belong.

3. The method for calculating gene expression levels according to claim 1, characterized in that, the constructing of the first model based on the first data in the cell data includes: Fitting the distribution of the sample genes of the target sample cells using the Poisson distribution, the formulas used include: y ij |λ ij ≈Poisson(λ ij ), where y ij represents the gene expression level of sample gene j in target sample cell i, and λ ij represents the expectation of the gene expression level y ij ; for sample gene j, λ ij is fitted as a combination of K explanatory variables x kj , and the formulas used include: where K represents the number of cell types to which the nearest neighbor cells of the target sample cells belong, k ∈ {1, 2,..., K}; β 0j and β kj represent Poisson parameters.

4. The method for calculating gene expression levels according to claim 3, characterized in that, the solution method includes: using the maximum likelihood method to solve the Poisson parameter.

5. The method for calculating gene expression levels according to claim 3, characterized in that, the method further includes: performing penalty regression on the solution method by using a preset optimal regularization parameter, and the penalty regression includes ridge regression and lasso regression.

6. The method for calculating gene expression levels according to claim 5, characterized in that, the method further includes: based on the cross-validation method, determining the regularization parameter that makes the coefficient of determination score of the first model the highest from a preset plurality of regularization parameters as the optimal regularization parameter.

7. The method for calculating gene expression levels according to claim 3, characterized in that, the method further includes performing a significance test on the Poisson parameter, including: estimating the parameter of the Poisson parameter to obtain an estimated value of the Poisson parameter; determining the deviation between the estimated value of the Poisson parameter and the corresponding unknown true parameter, and approximating the deviation to a standard normal distribution according to the principle of asymptotic normality; calculating the z value and p value of the two-tailed test according to the deviation, and controlling the false positive of the Poisson parameter based on the p value to determine the q value.

8. A method for calculating gene expression levels, characterized in that, the method includes: obtaining cell data of multiple sample cells, where the multiple sample cells include target sample cells; Construct a second model based on the second data in the cell data, and use a preset solution method to solve the second model and obtain the second gene expression level of the target sample cell, where the second data represents environmental factors in the second microenvironment where the target sample cell is located; the second data includes the cell type to which the multiple sample cells belong, and the spatial distance between the target sample cell and other sample cells; the second model is used to fit the gene distribution in the target sample cell when affected by the second microenvironment.

9. The method for calculating gene expression level according to claim 8, wherein, the constructing of the second model based on the second data in the cell data includes: Based on the gene expression level of each sample gene in the target sample cell, the cell type to which the multiple sample cells belong, the total number of cell types to which the multiple sample cells belong, and the spatial distance between the target sample cell and other sample cells, use a preset distribution function to fit the distribution of the sample genes of the target sample cell.

10. The method for calculating gene expression level according to claim 8, wherein, the formula used for constructing the second model based on the second data in the cell data includes: Among them, y ij represents the gene expression level of sample gene j in target sample cell i, and λ ij represents the expectation of the gene expression level y ij ; the error term represents ∈ ij the variance of, n represents the total number of the multiple sample cells, U represents the total number of cell types to which the multiple sample cells belong, and c iU represents the one-hot encoding vector of target sample cell i, and c hU represents the one-hot encoding vector of sample cell h, ω ih represents the distance weight between target sample cell i and sample cell h, and the vec() function is used to convert a U×U-dimensional matrix into a 1×U 2 -dimensional vector, and β 0j and β g represent the Poisson parameters.

11. A method for calculating gene expression level, wherein, the method includes: Obtain the cell data of multiple sample cells, where the multiple sample cells include a target sample cell; Construct a third model based on the third data in the cell data, and use a preset solution method to solve the third model and obtain the third gene expression level of the target sample cell, where the third data represents environmental factors in the third microenvironment where the target sample cell is located; the third data includes the cell type to which the multiple sample cells belong, the spatial distance between the target sample cell and other sample cells, and the ligand information of the multiple sample cells; the third model is used to fit the gene distribution in the target sample cell when affected by the third microenvironment.

12. The method for calculating gene expression level according to claim 11, wherein, the constructing of the third model based on the third data in the cell data includes: Based on the gene expression level of each sample gene in the target sample cell, the cell type to which the multiple sample cells belong, the total number of cell types to which the multiple sample cells belong, the spatial distance between the target sample cell and other sample cells, and the ligand information of the multiple sample cells, use a preset distribution function to fit the distribution of the sample genes of the target sample cell.

13. The method for calculating gene expression level according to claim 11, wherein, the formula used for constructing the third model based on the third data in the cell data includes: Among them, y ij represents the gene expression level of sample gene j in target sample cell i, and λ ij represents the expectation of the gene expression level y ij ; the error term represents ∈ ij 's variance, n represents the total number of the multiple sample cells, U represents the total number of cell types to which the multiple sample cells belong, and c iU represents the one-hot encoding vector of target sample cell i, and c hU represents the one-hot encoding vector of sample cell h, ω ih represents the distance weight between target sample cell i and sample cell h, and the vec() function is used to convert a U×U-dimensional matrix into a 1×U 2 -dimensional vector, R if represents the gene expression level of the receptor in the f-th ligand-receptor pair in target sample cell i, and L if represents the gene expression level of the ligand in the f-th ligand-receptor pair in target sample cell i, and β 0j and β fg represent Poisson parameters.

14. An electronic device, wherein, the electronic device includes a processor and a memory, and the processor is used to implement the method for calculating gene expression level according to any one of claims 1 to 13 when executing the computer program stored in the memory.

15. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, it implements the gene expression level calculation method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Method for determining gene expression regulation mechanism based on unicellular transcriptome data

    CN111613268A

  • Pet toilet training device

    KR1020240039264A