Biological tissue microenvironment analysis method, electronic device and storage medium
By using spatial transcriptome data and sliding window technology in biological tissues, the tensor matrix is constructed and decomposed, the problem of insufficient accuracy in microenvironment analysis of biological tissues is solved, and more accurate cell relationship analysis and microenvironment structure revelation is achieved.
Patent Information
- Application Number
- PCT/CN2023/138271
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-19
AI Technical Summary
The analytical accuracy of biological tissue microenvironment in the prior art is insufficient, especially in the analysis of tumor microenvironment, and it is difficult to associate the cell phenotype with the microenvironment.
By obtaining spatial transcriptome data of biological tissues, using sliding windows to split the data into multiple window data, constructing an adjacency map of each sliding window, obtaining the first interaction matrix, and obtaining the analysis results of the biological tissue microenvironment through tensor matrix decomposition.
It improves the analytical accuracy of the biological tissue microenvironment, can more comprehensively understand and analyze the relationship between cells, and reveals the tissue structure of the biological tissue microenvironment.
Smart Images

Figure CN2023138271_19062025_PF_FP_ABST
Abstract
Description
Biological tissue microenvironment analysis method, electronic device and storage medium Technical Field
[0001] The present application relates to the field of biomedical technology, and in particular to a method for analyzing the microenvironment of biological tissue, an electronic device, and a storage medium. Background Art
[0002] Cell communication coordinates the activities of various cell types to form tissues, organs, and systems, and further complete various biological functions. Cell communication is also essential for complex body processes, and the body's homeostasis requires interactions between cells to maintain. In order to understand the biological functions of each cell type in the tissue, it is necessary to analyze the microenvironment through cell communication, such as the tumor microenvironment. In related technologies, cell communication analysis methods often focus on analyzing the interaction information between two cell types, and it is easy to ignore the correlation structure of different cell phenotypes, making it difficult to link the phenotype with the tumor microenvironment, resulting in insufficient accuracy in analyzing the tumor microenvironment.
[0003] Summary of the Invention
[0004] The embodiments of the present application disclose a method for analyzing the microenvironment of biological tissue, an electronic device, and a storage medium, which solve the technical problem of insufficient accuracy in analyzing the microenvironment of biological tissue in related technologies.
[0005] The present application provides a method for analyzing the microenvironment of a biological tissue, the method comprising: obtaining spatial transcriptome data of the biological tissue; determining window data corresponding to multiple sliding windows from the spatial transcriptome data through a sliding window, each sliding window comprising multiple detection areas; constructing an adjacency graph for each sliding window based on the adjacency relationship between the multiple detection areas; obtaining a first interaction matrix for each sliding window based on the adjacency graph; constructing a tensor matrix based on the first interaction matrix; and decomposing the tensor matrix to obtain an analysis result of the biological tissue microenvironment corresponding to the biological tissue.
[0006] In some embodiments of the present application, obtaining spatial transcriptome data of biological tissue includes: using reference data to annotate cell types of cells in initial spatial transcriptome data to obtain the spatial transcriptome data.
[0007] In some embodiments of the present application, constructing the adjacency graph of each sliding window based on the adjacency relationship between the multiple detection areas includes: determining the adjacency relationship of any detection area based on the spatial distance between the detection areas; taking the detection area corresponding to the spatial distance less than or equal to a preset radius threshold as the first adjacent detection area of any detection area; and constructing the adjacency graph of each sliding window based on all the detection areas in each sliding window and the first adjacent detection area corresponding to each detection area in all the detection areas.
[0008] In some embodiments of the present application, constructing the adjacency graph of each sliding window based on the adjacency relationship between the multiple detection areas includes: obtaining adjacent detection areas that have an adjacency relationship with any detection area; selecting a preset number of adjacent detection areas as second adjacent detection areas; and constructing the adjacency graph of each sliding window based on all detection areas in each sliding window and the second adjacent detection areas corresponding to each detection area in all detection areas.
[0009] In some embodiments of the present application, obtaining the first interaction matrix of each sliding window based on the adjacency graph includes: obtaining the interaction strength between each pair of ligand-receptor pairs in the detection areas of each pair of cell types based on the adjacency graph; and obtaining the first interaction matrix based on the interaction strength.
[0010] In some embodiments of the present application, the interaction strength between each pair of ligand receptor pairs in the detection areas of each pair of cell types is obtained based on the adjacency graph, including: obtaining the neighbor pairs of each pair of ligand receptor pairs in the detection areas of each pair of cell types, the neighbor pairs being the detection areas adjacent to the detection areas corresponding to each pair of cell types; calculating the sum of the expression amounts of the ligand receptor pairs of the neighbor pairs, and using the sum as the interaction strength.
[0011] In some embodiments of the present application, constructing a tensor matrix based on the first interaction matrix includes: simulating the interaction strength distribution of each ligand-receptor pair by randomly permuting the labels of the cell types of the spatial transcriptome data to obtain a random interaction strength distribution of the second interaction matrix of the spatial transcriptome data; obtaining a probability value of each ligand-receptor pair based on the interaction strength distribution of each ligand-receptor pair under actual conditions of the spatial transcriptome data and the random interaction strength distribution; and setting the matrix element values in the first interaction matrix that are greater than a preset probability value to a preset value based on the probability value.
[0012] In some embodiments of the present application, the decomposition of the tensor matrix to obtain the analytical results of the biological tissue microenvironment corresponding to the biological tissue includes: obtaining multiple decomposition results of the tensor matrix based on multiple preset ranks; calculating the reconstruction error of the decomposition result corresponding to each rank; obtaining the core matrix, the factor matrix of the sliding window, the factor matrix of the cell type pairs, and the factor matrix of the ligand receptor pairs based on the decomposition result corresponding to the minimum reconstruction error; and obtaining the analytical results based on the factor matrix of the sliding window.
[0013] The present application also provides an electronic device, which includes a processor and a memory, and the processor is used to implement the method for analyzing the biological tissue microenvironment when executing the computer program stored in the memory.
[0014] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for analyzing the biological tissue microenvironment is implemented.
[0015] In the analytical method of the biological tissue microenvironment of the present application, the spatial transcriptome data of the biological tissue is obtained, and a plurality of sliding windows containing window data are determined from the spatial transcriptome data through a sliding window, each sliding window containing a plurality of detection areas, and the spatial transcriptome data is split into multiple data for analysis in the form of a sliding window, which can improve the accuracy of data analysis to a certain extent. Based on the adjacency relationship between the multiple detection areas, an adjacency graph of each sliding window is constructed, and based on the adjacency graph, a first interaction matrix of each sliding window can be obtained, and a tensor matrix is constructed based on the first interaction matrix. The tensor matrix is decomposed to obtain the analytical result of the biological tissue microenvironment corresponding to the biological tissue, and the tensor analysis method is adopted to more comprehensively understand and analyze the relationship between cells, better reveal the organizational structure of the biological tissue microenvironment, and improve the accuracy of analyzing the biological tissue microenvironment. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] FIG1 is a schematic structural diagram of an electronic device provided in an embodiment of the present application.
[0017] FIG2 is a flow chart of a method for analyzing a biological tissue microenvironment provided in an embodiment of the present application.
[0018] FIG3 is a cell annotation diagram of single-cell mouse tumor-bearing data provided in an example of the present application.
[0019] FIG4 is a graph of reconstruction error provided in an embodiment of the present application.
[0020] FIG5 is a schematic diagram of a core matrix provided in an embodiment of the present application.
[0021] FIG6 is a schematic diagram of a factor matrix of a sliding window provided in an embodiment of the present application.
[0022] FIG7 is a schematic diagram of a factor matrix of cell type pairs provided in an embodiment of the present application.
[0023] FIG8 is a schematic diagram of a factor matrix of ligand-receptor pairs provided in an embodiment of the present application.
[0024] FIG9 is the analysis result provided by an embodiment of the present application.
[0025] FIG10 is a diagram showing modules in the analysis results provided in an embodiment of the present application.
[0026] FIG11 is a schematic diagram of cell communication in the tumor microenvironment module 0 provided in an embodiment of the present application.
[0027] FIG12 is a schematic diagram of cell communication in the tumor microenvironment module 2 provided in an embodiment of the present application.
[0028] FIG13 is a cell annotation diagram of liver cancer tumor boundary data provided in an embodiment of the present application.
[0029] FIG14 is a diagram showing each cell type in the cell annotation diagram provided in an embodiment of the present application.
[0030] FIG15 is a schematic diagram of reconstruction error provided by yet another embodiment of the present application.
[0031] FIG16 is an analysis result provided by another embodiment of the present application.
[0032] FIG17 is a module display diagram of the analysis results provided in another embodiment of the present application.
[0033] FIG18 is a schematic diagram of cell communication in a tumor microenvironment module 7 according to another embodiment of the present application.
[0034] FIG19 is a cell annotation diagram provided in yet another embodiment of the present application.
[0035] FIG20 is a display diagram of each cell type in a cell annotation diagram provided in another embodiment of the present application.
[0036] FIG21 is a graph of reconstruction error provided by yet another embodiment of the present application.
[0037] FIG22 is the analysis result provided by another embodiment of the present application.
[0038] FIG23 is a module display diagram of the analysis results provided in another embodiment of the present application.
[0039] FIG24 is a schematic diagram of cell communication in a tumor microenvironment module 0 according to another embodiment of the present application.
[0040] FIG25 is a schematic diagram of cell communication in a tumor microenvironment module 1 according to another embodiment of the present application. DETAILED DESCRIPTION
[0041] To facilitate understanding, some illustrations of concepts related to the embodiments of the present application are given for reference.
[0042] It should be noted that in this application, "at least one" means one or more, and "more than one" means two or more than two. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0043] Cell communication coordinates the activities of various cell types to form tissues, organs, and systems, and further complete various biological functions. Cell communication is also essential for complex body processes, and the body's homeostasis requires interactions between cells to maintain. In order to understand the biological functions of each cell type in the tissue, it is necessary to analyze the microenvironment through cell communication, such as the tumor microenvironment. In related technologies, cell communication analysis methods often focus on analyzing the interaction information between two cell types, and it is easy to ignore the correlation structure of different cell phenotypes, making it difficult to link the phenotype with the tumor microenvironment, resulting in insufficient accuracy in analyzing the tumor microenvironment.
[0044] In order to solve the technical problem of insufficient accuracy in analyzing the microenvironment of biological tissues in related technologies, the embodiments of the present application provide a method for analyzing the microenvironment of biological tissues, an electronic device and a storage medium. The structure of the electronic device of the method for analyzing the microenvironment of biological tissues of the present application is first described below.
[0045] Figure 1 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The method for analyzing the biological tissue microenvironment provided in an embodiment of the present application is applied to an electronic device 10, which can be a mobile phone, tablet computer, laptop computer, netbook, server, or other device. The embodiment of the present application does not impose any restrictions on the specific type of electronic device.
[0046] As shown in Figure 1, the electronic device 10 may include a communication module 101, a memory 102, a processor 103, an input / output (I / O) interface 104, and a bus 105. The processor 103 is coupled to the communication interface 101, the memory 102, and the I / O interface 104 via the bus 105.
[0047] The communication module 101 may include a wired communication module and / or a wireless communication module. The wired communication module may provide one or more wired communication solutions such as Universal Serial Bus (USB) and Controller Area Network (CAN). The wireless communication module may provide one or more wireless communication solutions such as Wireless Fidelity (Wi-Fi), Bluetooth (BT), mobile communication network, Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR).
[0048] The memory 102 may include one or more random access memories (RAMs) and one or more non-volatile memories (NVMs). The RAM can be directly read and written by the processor 103 and can be used to store executable programs (e.g., machine instructions) of the operating system or other running programs, as well as user and application data. The RAM may include static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc.
[0049] The non-volatile memory can also store executable programs and user and application data, etc., which can be pre-loaded into the random access memory for direct reading and writing by the processor 103. The non-volatile memory can include disk storage devices and flash memory.
[0050] The memory 102 is used to store one or more computer programs. The one or more computer programs are configured to be executed by the processor 103. The one or more computer programs include multiple instructions. When the multiple instructions are executed by the processor 103, the method for analyzing the biological tissue microenvironment can be implemented on the electronic device 10.
[0051] In other embodiments, the electronic device 10 further includes an external memory interface for connecting to an external memory to expand the storage capacity of the electronic device 10 .
[0052] The processor 103 may include one or more processing units. For example, the processor 103 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processor (NPU). Different processing units may be independent devices or integrated into one or more processors.
[0053] The processor 103 provides computing and control capabilities. For example, the processor 103 is used to execute the computer program stored in the memory 102 to implement the above-mentioned method for analyzing the biological tissue microenvironment.
[0054] The I / O interface 104 is used to provide a channel for user input or output. For example, the I / O interface 104 can be used to connect various input and output devices, such as a mouse, keyboard, touch device, display screen, etc., so that the user can enter information or visualize information.
[0055] The bus 105 is at least used to provide a channel for mutual communication among the communication module 101 , the memory 102 , the processor 103 , and the I / O interface 104 in the electronic device 10 .
[0056] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 10. In other embodiments of the present application, the electronic device 10 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0057] Please refer to Figure 2, which is a flow chart of a method for analyzing a biological tissue microenvironment according to an embodiment of the present application, which is applied to an electronic device (e.g., electronic device 10 of Figure 1). Depending on different requirements, the order of the steps in the flow chart can be changed, and some steps can be omitted.
[0058] Step S201: Acquire spatial transcriptome data of biological tissues.
[0059] In some embodiments of the present application, the biological tissue may be an animal tissue or a plant tissue, etc. For example, an animal tissue or a mouse brain tissue containing tumor cells. Stereo-seq technology is used to perform spatial transcriptome sequencing on the biological tissue to obtain spatial transcriptome data of the biological tissue. Among them, Stereo-seq technology is a method for spatial transcriptome sequencing. Stereo-seq technology combines spatial information and gene expression information. By associating RNA on tissue sections with spatial positions, the spatial distribution of gene expression can be analyzed at the single cell level, which helps to understand the developmental process of tissues and organs. The spatial transcriptome data generated by Stereo-seq technology contains spatial information and gene expression information.
[0060] In some embodiments of the present application, the spatial transcriptome data may also include annotations of cell types. The cell types of cells in the initial spatial transcriptome data are annotated using reference data, thereby obtaining spatial transcriptome data containing cell type annotations. The initial spatial transcriptome data is data without cell annotation, and the reference data may be data of known cell types.
[0061] In one example, Figure 3 is a cell annotation diagram of single-cell mouse tumor data provided in an embodiment of the present application. Take the cell annotation of single-cell mouse tumor data as an example for illustration. Obtain a publicly released data set or experimental data provided by relevant experiments as reference data, and the reference data is single-cell mouse tumor data and the cell type is known. The cell2location toolkit can be installed in an electronic device. Cell2location is a tool for parsing the cell type composition of spatial transcriptome data. Use the cell2location toolkit to train a cell expression model, which can learn the expression characteristics of cell types based on existing annotation information. Obtain the initial spatial transcriptome data to be annotated, and the initial spatial transcriptome data is obtained by sequencing mouse tumor data using Stereo-seq technology. The initial spatial transcriptome data is input into the cell expression model to generate spatial transcriptome data annotated with cell types. This spatial transcriptome data contains the cell type to which each cell belongs. As shown in Figure 3, each pixel is annotated with the corresponding cell type, including: B cells, CD4+ T cells, CD8+ T cells, DC, M1-like, M2-like, Monocytes, NK cells, Neutrophils, Non-immune cells, Teff, and Treg. To improve the accuracy of cell type annotation, the results generated by the cell expression model can also be verified, including but not limited to immunohistochemical staining and histological verification.
[0062] Step S202 : determining window data corresponding to a plurality of sliding windows from the spatial transcriptome data by sliding the window, where each sliding window includes a plurality of detection regions.
[0063] In some embodiments of the present application, the size of the sliding window can be set according to actual needs, and the present application is not limited thereto. The form of expression of the sliding window is not limited and can be a virtual window. For example, the size of the sliding window can be 100um*100um. The spatial transcriptome data is traversed by the sliding window, and the spatial transcriptome data is intercepted into window data one by one. Each sliding window contains part of the spatial transcriptome data, and each sliding window contains multiple detection areas. The detection area can be bin20, which represents a square close to 10um*10um, which is close to a cell. Assuming that the size of the sliding window is 100um*100um, the sliding window contains 200 detection areas bin20. Performing data analysis in bin20 can have a higher resolution. The detection area can also be one or more marked cells. In related technologies, each cell can be marked according to sequencing or staining data.
[0064] In some embodiments of the present application, the sliding method of the sliding window is not restricted until the entire piece of spatial transcriptome data is traversed through the sliding window to form sliding windows, each sliding window containing part of the spatial transcriptome data, and then the data in each sliding window is analyzed.
[0065] Step S203: constructing an adjacency graph for each sliding window based on the adjacency relationship between the multiple detection areas.
[0066] In some embodiments of the present application, each sliding window includes multiple detection areas, each of which has adjacent detection areas. Adjacent detection areas can be determined by a preset radius threshold or a preset number, but are not limited to this in practical applications. After each detection area has determined its corresponding adjacent detection area, an adjacency graph can be constructed based on each detection area and its corresponding adjacent detection area. The adjacency graph uses the detection areas as nodes, and constructs edges of the adjacency graph for detection areas with an adjacency relationship, representing the interaction between the two detection areas. The weight of the edge can represent the strength of the interaction.
[0067] In some embodiments of the present application, determining adjacent detection areas by setting a radius threshold includes: determining the adjacency relationship of any detection area based on the spatial distance between the detection areas, determining the detection area corresponding to a spatial distance less than or equal to a preset radius threshold as the first adjacent detection area of any detection area, and constructing an adjacency graph for each sliding window based on all detection areas within each sliding window and the first adjacent detection area corresponding to each detection area in all detection areas. By setting the radius threshold, data within the circular range formed by the radius threshold is divided into the first adjacent detection area.
[0068] In another embodiment of the present application, adjacent detection areas are determined by setting a preset number, including: obtaining adjacent detection areas that have an adjacent relationship with any detection area, and determining whether there is an adjacent relationship based on spatial distance. A preset number of adjacent detection areas are selected as second adjacent detection areas. For example, if there are 4 adjacent detection areas and the preset number is 3, then based on spatial distance, the 3 adjacent detection areas with the closest spatial distance are retained as second adjacent detection areas. Based on all detection areas in each sliding window and the second adjacent detection areas corresponding to each detection area in all detection areas, an adjacency graph for each sliding window is constructed, with each detection area as a node, and the connection between the nodes indicates that the two nodes are second adjacent detection areas to each other.
[0069] Step S204: obtaining a first interaction matrix of each sliding window based on the adjacency graph.
[0070] In some embodiments of the present application, adjacency graph provides the relationship information between detection areas, and the node of adjacency graph is detection area, and the size of this detection area is close to a cell.According to relationship information, each pair of ligand receptor pairs in the detection area between each pair of cell types can be obtained.Wherein, ligand (Ligand) refers to the molecule that can be combined with receptor (Receptor), usually signal molecule, receptor is the protein that can identify and combine specific signal molecules found in cell membrane or cell, ligand receptor pairs refer to the combination between a ligand and a receptor, such as, in nerve cells, the combination of neurotransmitter (for example, dopamine) and its receptor, the gene expression amount of ligand receptor pairs is the gene expression amount in detection area.Cell type refers to a group of cells with similar morphology, structure and function, such as, human immune cells, neurons and epithelial cells etc.Cell type pairs refer to the cell pair with interaction relationship between different cell types, and can also be called sender-receiver pairs according to interaction relationship.
[0071] In some embodiments of the present application, the interaction strength is the sum of the expression levels of the ligand receptor genes of the neighbor pairs of the cell type pair. The method for calculating the interaction strength is specifically: obtain the neighbor pairs of each pair of ligand receptor pairs in the detection area of each pair of cell types, and the neighbor pairs are the detection areas adjacent to the detection areas corresponding to each pair of cell types, calculate the sum of the expression levels of the ligand receptor pairs of the neighbor pairs, and obtain the interaction strength.
[0072] In one example, assume that based on the adjacency graph, it is determined that the detection region of cell type A has connecting lines with the detection regions of cell type B, cell type C, and cell type D. Then, to calculate the interaction strength between the detection region of cell type A and the detection region of cell type B (cell type pair), assume that the expression level of the ligand gene in all cells of cell type A is S1, and the expression level of the receptor gene in all cells of cell type B is S2. Then, the interaction strength S_AB of the ligand-receptor pair S1_S2 between the detection region of cell type A and the detection region of cell type B is S1+S2.
[0073] In some embodiments of the present application, for each sliding window, each (cell type pair) sender-receiver pair may have one or more ligand-receiver pairs. Therefore, each element in the first interaction matrix obtained based on the interaction strength is a numerical value representing the interaction strength of different ligand-receiver pairs between the sender-receiver pair. The rows of the first interaction matrix are sender-receiver pairs (i.e., cell type pairs), which are listed as ligand-receiver pairs.
[0074] Step S205: construct a tensor matrix based on the first interaction matrix.
[0075] In some embodiments of the present application, in order to screen out ligand-receptor pairs with more significant interaction strength, the first interaction matrix can be updated by calculating the second interaction matrix of the spatial transcriptome data. The second interaction matrix can be a full-slice interaction matrix, which represents the interaction matrix generated based on the spatial transcriptome data (all window data).
[0076] In some embodiments of the present application, a permutation check is performed on the spatial transcriptome data that has been annotated with cell types, and the permutation check is used to estimate the probability of the results observed in the given data appearing in a random background. The interaction strength distribution of each ligand receptor pair is simulated by randomly permuting the labels of the cell types of the spatial transcriptome data to obtain the random interaction strength distribution of the second interaction matrix. Based on the interaction strength distribution of each ligand receptor pair and the random interaction strength distribution under the actual situation of the spatial transcriptome data, the probability value of each ligand receptor pair is obtained, that is, the probability of the interaction strength of the first interaction matrix appearing under random circumstances. Based on the probability value, the matrix element value in the first interaction matrix that is greater than the preset probability value is set to a preset value. For example, the preset probability value can be 0.05, and the matrix element value with the probability value P>0.05 in the first interaction matrix is set to the preset value 0. The preset value can also be other integers, which are not limited by the present application.
[0077] In some embodiments of the present application, by updating the first interaction matrix of all sliding windows through the second interaction matrix, ligand-receptor pairs with significant interaction strength can be screened out, which can improve the efficiency of data processing to a certain extent.
[0078] In some embodiments of the present application, a tensor matrix is a data structure that can represent multiple dimensions. In the biomedical field, constructing a tensor matrix helps to process and analyze complex medical image data. The first interaction matrix after all sliding windows are updated is stacked to obtain a three-dimensional tensor matrix. Since when calculating the interaction intensity of the ligand receptor pairs between two cell types, the interaction intensity of the ligand receptor pairs between all cell types in a sliding window may be 0, therefore, in order to avoid the analysis of invalid data, in the process of stacking, the matrices that are all 0 in three dimensions are discarded. The dimensions of the tensor matrix are respectively sliding window, sender-receiver pair and ligand receptor pair.
[0079] Step S206 , decomposing the tensor matrix to obtain an analytical result of the biological tissue microenvironment corresponding to the biological tissue.
[0080] In some embodiments of the present application, the tensor matrix is decomposed, and the decomposition method may include but is not limited to: non-negative (Tucker) decomposition, decomposing the tensor matrix into the product form of multiple non-negative matrices, thereby extracting the main features and structures in the tensor matrix. In other embodiments of the present application, the tensor matrix can also be decomposed using CP decomposition (Canonical Polyadic Decomposition, CPD), singular value decomposition (SVD), alternating least squares (ALS) and non-negative least squares (NMF), etc. This application does not limit the decomposition method of the tensor matrix.
[0081] In order to obtain the optimal decomposition result, multiple rank numbers are set in advance to obtain multiple decomposition results of non-negative decomposition of the tensor matrix. Each rank number corresponds to a decomposition result, and the reconstruction error of the decomposition result corresponding to each rank number is calculated. The reconstruction error is the mean square difference between the tensor matrix before reconstruction and the tensor matrix after reconstruction. A curve is drawn for the reconstruction error corresponding to each rank number, so that the minimum reconstruction error can be determined based on the curve.
[0082] Figure 4 is a graph of the reconstruction error provided by an embodiment of the present application. As shown in Figure 4, the horizontal axis represents the sender-receiver pair (ct-ct) and the ligand-receiver pair (LR), the vertical axis represents the reconstruction error (reconstruction error), and TME represents the sliding window. Each line corresponds to the reconstruction error of a different rank number. The smaller the reconstruction error, the better the effect of Tucker decomposition. As shown in Figure 4, the curve with TME equal to 14 decreases slowly when the horizontal axis is equal to 14 to 18, indicating convergence. Therefore, it is determined that the Tucker decomposition of the curve with TME of 14 has the best effect, and the reconstruction error with TME of 14 is determined to be the smallest reconstruction error.
[0083] In some embodiments of the present application, based on the curve with TME equal to 14, the rank number is determined to be (14, 14, 14), and the decomposition result corresponding to the rank number is obtained to obtain the core matrix, the factor matrix of the sliding window, the factor matrix of the cell type pairs, and the factor matrix of the ligand-receptor pairs. Based on the factor matrix of the sliding window, the analytical result is obtained.
[0084] 5 to 9 are used for illustration. FIG5 is a schematic diagram of a core matrix provided in an embodiment of the present application. As shown in FIG5, each core matrix corresponds to a module of a tumor microenvironment. The core matrix's modules represent sender-receiver pairs, its columns represent ligand-receiver pairs, and its values represent the corresponding relationships between modules. FIG6 is a schematic diagram of a factor matrix for a sliding window provided in an embodiment of the present application. As shown in FIG6, each sliding window represents a small tumor microenvironment, its modules represent sliding windows, its columns represent modules, and its values represent the weights of different sliding windows within the module. FIG7 is a schematic diagram of a factor matrix for a cell type pair provided in an embodiment of the present application. As shown in FIG7, the factor matrix for a sender-receiver pair (cell type pair) represents sender-receiver pairs, its columns represent different modules, and its values represent the weights of the sender-receiver pairs within different modules. FIG8 is a schematic diagram of a factor matrix for a ligand-receiver pair provided in an embodiment of the present application. As shown in FIG8, the factor matrix for a ligand-receiver pair represents ligand-receiver pairs, its columns represent different modules, and its values represent the weights of the ligand-receiver pairs within different modules.
[0085] Figure 9 shows the parsing results provided by an embodiment of the present application. As shown in Figure 9, to determine different tumor microenvironment modules, the weights of different sliding windows are determined based on the factor matrix of the sliding window shown in Figure 6. Based on the weights of the different sliding windows, the modules to which the sliding windows belong can be determined, thereby obtaining the parsing results shown in Figure 9. Different modules in the parsing results represent different classes, and there are 14 classes in total, as shown in Figure 9.
[0086] Figure 10 is a diagram showing the modules in the analysis results provided in the examples of this application. As shown in Figure 10, the analysis results shown in Figure 9 are split into multiple sub-graphs for display, indicating the identification of different tumor microenvironments and the exploration of cellular communication in the tumor microenvironment.
[0087] In one example, the tumor microenvironment module 0 in Figure 10 is mainly enriched in cell pair module 0 and ligand receptor pair module 0, the cell pair module 0 is mainly related to non-immune cells, and the ligand receptor pair 0 is mainly cell interaction mediated by SPP1 and FN1, as shown in Figure 11. Figure 11 is a schematic diagram of the cell communication situation of the tumor microenvironment module 0 provided in an embodiment of the present application. As shown in Figure 11, the diagrams from left to right are the schematic diagram of the tumor microenvironment module 0 in Figure 10, the module schematic diagram of the first tumor microenvironment in Figure 5, the schematic diagram of non-immune cells, and the weight schematic diagram of the ligand receptor pair.
[0088] In another example, as shown in FIG10 , the tumor microenvironment module 2 mainly interacts with non-immune cells and neutrophils, and the main ligand receptors are chemokine-related ligand receptor pairs, suggesting that the tumor may recruit neutrophils during progression, as shown in FIG12 . FIG12 is a schematic diagram of the cell communication of the tumor microenvironment module 2 provided in an embodiment of the present application. As shown in FIG12 , the diagrams from left to right are, in sequence, a schematic diagram of the tumor microenvironment module 2 as shown in FIG10 , a schematic diagram of the module of the third tumor microenvironment as shown in FIG5 , a schematic diagram of non-immune cells and neutrophils (Neutrophils), and a weighted schematic diagram of the ligand receptor pairs.
[0089] In an embodiment of the present application, the spatial transcriptome data of biological tissue are obtained, multiple sliding windows comprising window data are obtained from the spatial transcriptome data by sliding window, each sliding window comprises multiple detection areas, by the form of sliding window, the spatial transcriptome data are split into multiple data for analysis, and the precision of data analysis can be improved to a certain extent. Based on the adjacency relationship between multiple detection areas, the adjacency graph of each sliding window is constructed, based on the adjacency graph, the first interaction matrix of each sliding window can be obtained, the first interaction matrix is updated using the second interaction matrix of spatial transcriptome data, and the ligand pairs comprising significant interaction intensity can be screened out. Based on the first interaction matrix after all sliding windows are updated, tensor matrix is constructed, tensor matrix is decomposed, and the analytical result of the biological tissue microenvironment corresponding to biological tissue is obtained, the analytical method of tensor is used, the relationship between and analytical cells can be more fully understood, the organizational structure of biological tissue microenvironment is better revealed, and the accuracy of analytical biological tissue microenvironment can be improved.
[0090] In other embodiments of the present application, in addition to validating the effectiveness of the present application on mouse tumor-bearing data, it can also be validated on liver cancer tumor boundary data. The liver cancer tumor boundary data to be processed and reference data are obtained, and cell annotation is performed using cell2location to obtain spatial transcriptome data corresponding to the annotated liver cancer tumor boundary data.
[0091] For illustration purposes, Figures 13 and 14 are provided. Figure 13 is a cell annotation diagram of liver cancer tumor boundary data provided in an embodiment of the present application. As shown in Figure 13 , the cell types include: B_cell, Cholangiocyte, DC, Endothelial, Fibroblast, Hepatocyte, Macrophage, Malignant, NK, and T_cell. Figure 14 is a diagram showing each cell type in the cell annotation diagram provided in an embodiment of the present application.
[0092] In some embodiments of the present application, an adjacency graph of each sliding window is constructed to construct a tensor matrix, and the reconstruction error at different ranks is calculated. The rank is obtained by constructing a curve graph of the reconstruction error, and non-negative decomposition is performed based on the rank to obtain the corresponding analytical results of the liver cancer tumor boundary data.
[0093] With reference to Figures 15 to 18, Figure 15 is a schematic diagram of the reconstruction error provided by another embodiment of the present application. As shown in Figure 15, the curve with TME equal to 14 tends to be flat when the horizontal coordinate is 15, that is, it converges, then the rank number is (14, 15, 15), and the tensor matrix is non-negatively decomposed according to the rank number (14, 15, 15), and the analytical result shown in Figure 16 is obtained, and Figure 16 is the analytical result provided by another embodiment of the present application. Figure 17 is a module display diagram in the analytical result provided by another embodiment of the present application. In the related art, as shown in Figures 16 and 17, it is usually necessary to split the input data and then generate the tumor microenvironment under different communication modes, and the detection area used is bin50. Compared to the related art, the present application can use the entire data as input, without disassembling the data, that is, the tumor microenvironment under different communication modules can be found, and the detection area used in the present application is bin20, which can improve the efficiency of the analysis and the accuracy of the results.
[0094] In one example, FIG18 is a schematic diagram of cell communication in a tumor microenvironment module 7 according to another embodiment of the present application. FIG18 shows, from left to right, the schematic diagram of the tumor microenvironment module 7 in FIG16 , a schematic diagram of the correspondence between cell type pairs and ligand-receptor pairs in the tumor microenvironment module 7, and a schematic diagram of the average communication strength of the first 15 ligand-receptor pairs for the first five cell type pairs in the tumor microenvironment module 7. As shown in FIG18 , the effectiveness of this embodiment in applying liver cancer tumor boundary data was determined.
[0095] In other embodiments of the present application, the validity of the present application can also be verified in other data types. For example, liver cancer MERFISH data, using public single-cell intrahepatic cholangiocarcinoma data through cell2location to annotate liver cancer MERIFSH data, obtain cell type annotation results, as shown in Figure 19, which is a cell annotation diagram provided in another embodiment of the present application, and Figure 20, which is a display diagram of each cell type in the cell annotation diagram provided in another embodiment of the present application. As shown in Figure 20, cell types include B cell, Endothelial, Fibroblast, Hepatocyte, Macrophage, NK, and T cell. Then calculate the reconstruction error under different ranks to form a curve graph.
[0096] Figure 21 is a graph of the reconstruction error provided by another embodiment of the present application. As shown in Figure 21, the reconstruction error corresponding to the liver cancer MERFISH data cannot converge with the increase of the rank number, but the reconstruction error only increases within a small range, indicating that different ranks have little effect on the decomposition results. The rank number can be selected as (8, 8, 8), and the tensor matrix is decomposed with the rank number of (8, 8, 8) to obtain the analytical result. Figure 22 is the analytical result provided by another embodiment of the present application, and Figure 23 is a module display diagram in the analytical result provided by another embodiment of the present application.
[0097] In one example, as shown in the tumor microenvironment module 0 in FIG23 , the tumor microenvironment module 0 mainly involves the interaction between liver cells through the cell adhesion-related ligand receptor pair FN1_IGA5_ITGB1, as shown in FIG24 . FIG24 is a schematic diagram of the cell communication of the tumor microenvironment module 0 provided in another embodiment of the present application. From left to right, the diagrams are the tumor microenvironment module 0 in FIG23 , the module schematic diagram of the tumor microenvironment, and the weight schematic diagram of the ligand receptor pair.
[0098] In another example, as shown in the tumor microenvironment module 1 in Figure 23, the tumor microenvironment module 1 contains the chemokine-related ligand receptor CXCL12_CXCR4 between fibroblasts and T cells, which plays a role in recruiting T cells. This is consistent with the histological characteristics of the tumor microenvironment module 1 being in the interstitial region and rich in immune cells, as shown in Figure 25. Figure 25 is a schematic diagram of the cell communication situation of the tumor microenvironment module 1 provided in another embodiment of the present application. From left to right, the diagrams are the tumor microenvironment module 1 in Figure 23, the module schematic diagram of the tumor microenvironment, and the weight schematic diagram of the ligand receptor pair.
[0099] In the examples of this application, the effectiveness of the method was verified on mouse tumor-bearing data, liver cancer tumor boundary data, and liver cancer MERFISH data. It can effectively discover the structure of the tumor microenvironment and provide important data support for tumor treatment.
[0100] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions. The method implemented when the program instructions are executed can refer to the methods in the above-mentioned embodiments of the present application.
[0101] The computer-readable storage medium may be an internal memory of the electronic device described in the above embodiment, such as a hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, etc. equipped on the electronic device.
[0102] In some embodiments, the computer-readable storage medium may include a program storage area and a data storage area, wherein the program storage area may store an operating system, applications required for at least one function, etc.; the data storage area may store data created according to the use of the electronic device, etc.
[0103] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0104] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0105] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0106] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0107] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for analyzing a biological tissue microenvironment, characterized in that, The method includes: Obtaining spatial transcriptome data of a biological tissue; Determining window data corresponding to a plurality of sliding windows from the spatial transcriptome data through a sliding window, each sliding window including a plurality of detection regions; Constructing an adjacency graph of each sliding window based on the adjacency relationship between the plurality of detection regions; Obtaining a first interaction matrix of each sliding window based on the adjacency graph; Constructing a tensor matrix based on the first interaction matrix; Decomposing the tensor matrix to obtain an analysis result of the biological tissue microenvironment corresponding to the biological tissue.
2. The method for analyzing a biological tissue microenvironment according to claim 1, characterized in that, The obtaining of the spatial transcriptome data of the biological tissue includes: Annotating the cell types of cells in the initial spatial transcriptome data by using reference data to obtain the spatial transcriptome data.
3. The method for analyzing a biological tissue microenvironment according to claim 1, characterized in that, The constructing of the adjacency graph of each sliding window based on the adjacency relationship between the plurality of detection regions includes: Determining the adjacency relationship of any detection region based on the spatial distance between detection regions; Taking the detection regions corresponding to the spatial distances less than or equal to a preset radius threshold as the first adjacent detection regions of the any detection region; Constructing the adjacency graph of each sliding window based on all detection regions within each sliding window and the first adjacent detection regions corresponding to each detection region among all detection regions.
4. The method for analyzing a biological tissue microenvironment according to claim 3, characterized in that, The constructing of the adjacency graph of each sliding window based on the adjacency relationship between the plurality of detection regions includes: Obtaining adjacent detection regions having an adjacency relationship with any detection region; Selecting a preset number of adjacent detection regions as second adjacent detection regions; Constructing the adjacency graph of each sliding window based on all detection regions within each sliding window and the second adjacent detection regions corresponding to each detection region among all detection regions.
5. The method for analyzing a biological tissue microenvironment according to claim 2, characterized in that, The obtaining of the first interaction matrix of each sliding window based on the adjacency graph includes: Obtaining the interaction intensity between the detection regions of each pair of cell types for each pair of ligand-receptor pairs based on the adjacency graph; Obtaining the first interaction matrix based on the interaction intensity.
6. The method for analyzing a biological tissue microenvironment according to claim 5, characterized in that, The obtaining of the interaction intensity between the detection regions of each pair of cell types for each pair of ligand-receptor pairs based on the adjacency graph includes: Obtaining the neighbor pairs of the detection regions of each pair of cell types for each pair of ligand-receptor pairs, the neighbor pairs being the detection regions adjacent to the detection regions corresponding to each pair of cell types; Calculating the sum value of the expression amounts of the ligand-receptor pairs of the neighbor pairs, and taking the sum value as the interaction intensity.
7. The method for analyzing a biological tissue microenvironment according to claim 1, characterized in that, The constructing of the tensor matrix based on the first interaction matrix includes: Simulating the interaction intensity distribution of each ligand-receptor pair by randomly permuting the labels of the cell types of the spatial transcriptome data to obtain the random interaction intensity distribution of the second interaction matrix of the spatial transcriptome data; Based on the actual interaction intensity distribution of each ligand-receptor pair in the spatial transcriptome data and the random interaction intensity distribution; Obtaining the probability value of each ligand-receptor pair; Based on the probability value, setting the matrix element values greater than a preset probability value in the first interaction matrix to a preset value to obtain an updated first interaction matrix; Constructing the tensor matrix based on the updated first interaction matrix.
8. The method for analyzing a biological tissue microenvironment according to claim 1, characterized in that,Decomposing the tensor matrix to obtain the analysis result of the biological tissue microenvironment corresponding to the biological tissue includes: Obtaining a plurality of decomposition results of decomposing the tensor matrix based on a plurality of preset ranks; Calculating the reconstruction error of the decomposition result corresponding to each rank; Based on the decomposition result corresponding to the minimum reconstruction error, obtaining a core matrix, a factor matrix of the sliding window, a factor matrix of cell type pairs, and a factor matrix of ligand-receptor pairs; Based on the factor matrix of the sliding window, obtaining the analysis result.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory, and the processor is configured to execute a computer program stored in the memory to implement the method for analyzing the biological tissue microenvironment according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the method for analyzing the biological tissue microenvironment according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Method for analyzing periodontitis immune microenvironment through space transcriptome technology
CN112185465A
Cell data annotation method, device, equipment and medium
CN115116549A
Spatial transcription data organization region analysis method, system and equipment and storage medium
CN116646005A
Feature learning method for identifying spatial transcriptome spatial region and cell type
CN116741273A
Spatial transcriptome data clustering method and device based on comparative learning and medium
CN117153260A