A method, device and system for identifying a high gene module score region
By dividing spatial transcriptomics data into unit grids and calculating the integrated gene set score index, the problem of existing technologies being unable to integrate gene set co-expression and spatial location information is solved, enabling accurate identification and localization of functionally active regions and improving the ability to analyze complex diseases and biological development processes.
Patent Information
- Application Number
- CN202511596425.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-11-04
Smart Images

Figure CN121054101B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of spatial transcriptomics technology, and in particular to a method, apparatus and system for identifying regions with high gene module scores. Background Technology
[0002] In biomedical research, understanding the molecular mechanisms within tissues and organs is crucial for revealing disease patterns and the mysteries of life. The functions of an organism are achieved through the collaboration of different cells in specific spatial locations and the synergistic expression of multiple genes. Therefore, gene expression analysis while preserving tissue spatial information is essential. Spatial transcriptomics (ST) technology, by combining high-throughput sequencing with tissue imaging, enables the simultaneous acquisition of gene expression profiles and spatial location information in situ within tissues. This provides a powerful tool for exploring gene function and cellular heterogeneity from a spatial perspective, demonstrating enormous application potential in fields such as oncology, neuroscience, and developmental biology.
[0003] Existing analytical methods for spatial transcriptomics data mainly focus on two levels. One type of method aims to identify genes with significantly different expression levels in different tissue regions, such as using traditional differential gene analysis tools (e.g., DESeq2, edgeR) to find marker genes associated with specific pathological structures. The other type of method explores gene sets with similar expression patterns, such as using tools like weighted gene co-expression network analysis (WGCNA) to identify gene modules that may be functionally related and synergistic. These techniques can effectively screen core genes or gene communities from massive amounts of data and provide a preliminary display of their spatial distribution, offering clues for understanding tissue function.
[0004] However, traditional techniques have significant limitations in analyzing gene synergistic effects and their spatial characteristics. On the one hand, differential gene analysis methods essentially perform statistical analysis on individual genes, neglecting inter-gene interactions, thus failing to reveal gene communities that significantly influence biological functions through synergistic effects. On the other hand, while methods such as WGCNA can identify functionally related gene sets, they typically cannot further quantify the enrichment and activity levels of these gene sets across tissue spatial dimensions. Although these methods can preliminarily determine which genes are synergistically acting, they struggle to precisely pinpoint the specific tissue region where synergistic effects are most active.
[0005] In summary, a key technological bottleneck exists in the current field of spatial transcriptomics data analysis: the lack of analytical tools capable of effectively integrating synergistic expression and spatial location information of gene sets, and quantifying the spatial effects of gene synergy. This technological deficiency prevents the precise identification and localization of functionally highly active regions formed by the synergistic effects of specific gene modules. In studying complex diseases or biological development processes exhibiting high spatial heterogeneity, the inability to accurately capture the dynamic functional characteristics of different tissue regions limits our understanding of key pathophysiological processes. Summary of the Invention
[0006] In a first aspect, the present invention provides a method for identifying regions with high gene module scores, comprising:
[0007] Spatial transcriptome data containing gene expression information and spatial coordinate information of multiple cells is obtained. The cells are divided into multiple cell grids according to the spatial coordinate information of the cells. The corresponding grid spatial coordinates are determined for each cell grid, and the cell grids are numbered to form an ordered arrangement of cell grids.
[0008] Based on a preset gene set, calculate the module score value for each cell, and determine the set of high-scoring cells based on the module score values;
[0009] For each of the said unit grids, an integrated gene set score index is calculated based on the grid space coordinates of the unit grid, the cells included in the unit grid, and the corresponding module score values;
[0010] Based on the integrated gene set score index of the unit grid, the unit grid is divided into regions, and high gene module score regions are identified.
[0011] In an optional implementation, the step of dividing the cell into multiple cell grids based on the cell's spatial coordinate information, and determining the corresponding grid spatial coordinates for each cell grid, includes:
[0012] Based on the spatial coordinate information of all cells, determine the maximum and minimum values of the x-axis and the maximum and minimum values of the y-axis in the spatial transcriptome data.
[0013] According to the preset meshing quantity parameter, the coordinate ranges of the x-axis and the y-axis are divided into equal parts to form multiple unit meshes;
[0014] Each cell is assigned to the corresponding cell grid based on the spatial coordinate information;
[0015] Calculate the geometric centroid of the spatial coordinates of all cells included in each unit grid, and use the geometric centroid as the grid spatial coordinates of that unit grid.
[0016] In an optional implementation, numbering the plurality of cell grids to form an ordered cell grid includes: sequentially numbering all the cell grids according to their spatial location; wherein the numbering proceeds along columns, with odd-numbered columns numbered consecutively from top to bottom, and even-numbered columns numbered consecutively from bottom to top; and / or,
[0017] The magnitude of the x-axis direction in each of the aforementioned cell grids is: ;
[0018] The magnitude of the y-axis direction is: ;
[0019] Among them, L x L represents the magnitude of the x-axis direction in the cell grid. y The x represents the magnitude of the y-axis direction in the cell grid; max Represents the maximum value on the x-axis; x min Represents the minimum value on the x-axis; y max Represents the maximum value on the y-axis; y min The minimum value of the y-axis is represented by N; N represents the number of meshing parameters; the total number of all cell meshes after equal division is N×N.
[0020] In an optional implementation, the step of calculating a module score value for each cell based on a preset gene set, and determining a set of high-scoring cells based on the module score values, includes:
[0021] The AddModuleScore function is used to calculate the module score value of the preset gene set in each cell;
[0022] The module scores are sorted from largest to smallest, and the cells before the preset ranking after sorting are extracted as the high-scoring cell set; and the mean and standard deviation of the module scores of the high-scoring cell set are calculated.
[0023] In an optional implementation, the calculation expression for the integrated gene set scoring index is:
[0024] ;
[0025] Among them, IGSI w The score index represents the integrated gene set.
[0026] In the expression, of The Euclidean distance represents the spatial coordinates of cell c to the grid spatial coordinates of its cell cell w. The attenuation coefficient representing the grid;
[0027] In the expression, S c When calculating the integrated gene set score index of a cell grid, the module score value of cell c in that cell grid is used; S ci This represents the module score value of any cell ci in any cell i when all cell grids are summed up.
[0028] In the expression, Cell C represents the set of high-scoring cells. The probability within the interval;
[0029] In the expression, of Indicates the cell density sensitivity coefficient; f represents the local cell density; f represents the overall or average cell density within a single grid cell w.
[0030] In the expression, The correlation weight value between cell c within cell w and cell w represents the correlation weight value between cell c and cell w; where, The gene expression vector representing cell c; The average expression profile of cells in the unit grid w represents the average expression profile of cells in the unit grid w.
[0031] In an optional implementation, in the expression, ;in, Let represent the distance between any cell l and any cell j in the cell grid w, and ; and / or,
[0032] In the expression, The calculation formula is: ;in, Let (x, y) represent the spatial coordinates of cell c, K represent the number of cells, and K represent the cell that is spatially closest to cell c. The bandwidth parameter representing the control space attenuation is calculated using the following formula: The distance between cell c and any cell j in the unit grid w; and / or,
[0033] In the expression, the formula for calculating f is: Where KW represents the number of cells in the unit grid w. ; Let represent the distance between any cell l and any cell j in the cell grid w, and ; and / or,
[0034] In the expression, The calculation formula is: ; and / or,
[0035] In the expression, The calculation formula is: ; and / or,
[0036] In the expression, The calculation formula is:
[0037] ;
[0038] The average score of all cell modules in the middle; The standard deviation of the score values for all cell modules.
[0039] In an optional implementation, the step of dividing the unit grid into regions based on the integrated gene set scoring index of the unit grid and determining the high gene module scoring regions therein includes:
[0040] The ordered unit grids are divided according to the integrated gene set scoring index to form multiple regions to be analyzed with different means;
[0041] From the multiple regions to be analyzed, the region that simultaneously meets all preset conditions is determined as the high gene module scoring region;
[0042] The preset conditions include:
[0043] A. The mean of the integrated gene set scoring index of the region to be analyzed is within the preset ranking range in all regions;
[0044] B. The region to be analyzed contains at least one cell from the set of high-scoring cells.
[0045] In a second aspect, the present invention provides a device for identifying high gene module scoring regions, comprising:
[0046] The segmentation module is used to acquire spatial transcriptome data containing gene expression information and spatial coordinate information of multiple cells, divide the cells into multiple cell grids according to the spatial coordinate information of the cells, and determine the corresponding grid spatial coordinates for each cell grid.
[0047] The calculation module is used to calculate the module score value of each cell according to a preset gene set, and to determine the set of high-scoring cells based on the module score value;
[0048] The calculation module is also used to calculate the integrated gene set score index for each unit grid based on the grid space coordinates of the unit grid, the cells included in the unit grid, and the corresponding module score values;
[0049] The screening module is used to divide the unit grid into regions based on the integrated gene set score index of the unit grid, and to determine the high gene module score regions therein.
[0050] Thirdly, the present invention provides a computer device including a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the method for identifying high gene module scoring regions as described in any of the foregoing embodiments.
[0051] Fourthly, the present invention provides a computer storage medium storing a computer program, which, when executed on a processor, implements the method for identifying high gene module scoring regions according to any one of the foregoing embodiments.
[0052] This application provides a method, apparatus, and system for identifying high gene module scoring regions. The method for identifying high gene module scoring regions calculates module scoring values for a preset gene set, enabling the analysis of multiple genes as a functional whole. This overcomes the limitation of focusing only on changes in the expression of a single gene, which fails to reveal the interactions and synergistic regulatory patterns between genes, allowing the analysis to better reflect complex biological processes involving multiple genes.
[0053] Furthermore, by calculating an integrated gene set score index for each cell grid, the spatial quantification of gene set synergistic effects was achieved. This index combines the module score value of the gene set with the spatial coordinate information of the cell grid, assigning a clear and quantifiable activity score to each specific spatial location. This solves the problem that existing technologies can identify gene sets but cannot measure their enrichment degree in key regions of tissues, effectively filling the gap in the spatial distribution characteristics analysis of gene synergistic effects.
[0054] Finally, based on the obtained integrated gene set scoring index, all unit grids are divided into regions, enabling precise identification and localization of high gene module scoring areas. In this way, key tissue regions with active specific gene modules can be directly identified, allowing for a more comprehensive capture and analysis of the dynamic changes in different regions during complex diseases or biological development. This provides a more accurate perspective for understanding the physiological functional layout and pathological mechanism evolution of tissues. Attached Figure Description
[0055] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and therefore should not be considered as a limitation on the scope of protection of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0056] Figure 1 This is a schematic diagram of the hardware operating environment involved in an embodiment of the high gene module scoring region identification method of the present invention;
[0057] Figure 2 This is a flowchart illustrating Embodiment 1 of the method for identifying high gene module scoring regions of the present invention;
[0058] Figure 3 This is a detailed flowchart of step S100 in Embodiment 2 of the method for identifying high gene module scoring regions of the present invention;
[0059] Figure 4 This is a schematic diagram of the division, arrangement, and numbering of the unit grid in Embodiment 2 of the method for identifying high gene module scoring regions of the present invention;
[0060] Figure 5 This is a detailed flowchart of step S200 in embodiment 3 of the method for identifying high gene module scoring regions of the present invention;
[0061] Figure 6 This is a detailed flowchart of step S400 in embodiment 4 of the method for identifying high gene module scoring regions of the present invention;
[0062] Figure 7 This is a schematic diagram showing the identification results of the high gene module scoring region identification method of the present invention in Example 5, and the location of the tissue lymphoid structure in the H&E staining image of the literature (PMID:35149721);
[0063] Figure 8 This is a schematic diagram of the HGMSR location in aging-related regions in different tissues in Example 5 of the method for identifying high gene module score regions of the present invention (PMID: 39500323).
[0064] Figure 9 This is a schematic diagram of the module connection of the high gene module scoring region identification device of the present invention. Detailed Implementation
[0065] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0066] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0067] In the following, the terms “comprising,” “having,” and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as excluding, firstly, the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations thereof.
[0068] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0069] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.
[0070] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0071] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0072] like Figure 1 The diagram shown is a structural schematic of the hardware operating environment of the terminal involved in an embodiment of the present invention.
[0073] The high genotype module scoring region identification system of this invention can be a PC, or a mobile terminal device such as a smartphone, tablet, or laptop. This high genotype module scoring region identification system may include: a processor 1001, such as a CPU; a network interface 1004; a user interface 1003; a memory 1005; and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen, an input unit such as a keyboard, or a remote control; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory or a stable memory, such as a disk storage device. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001. Optionally, the high genotype module scoring region identification system may also include RF (Radio Frequency) circuitry, audio circuitry, a Wi-Fi module, etc. In addition, the identification system for the high gene module scoring region can also be configured with other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, which will not be elaborated here.
[0074] Those skilled in the art will understand that Figure 1 The identification system for high gene module scoring regions shown is not intended to limit it and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. Figure 1 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a data interface control program, a network connection program, and a program for identifying high gene module scoring regions.
[0075] In summary, the method provided by this invention analyzes multiple genes as a functional whole, overcoming the limitations of traditional single-gene analysis in revealing synergistic regulation. Its innovation lies in calculating an integrated gene set score index for each spatial location, thereby achieving spatial quantification of gene set synergistic effects and solving the problem of existing technologies being unable to locate gene set enrichment regions. Ultimately, this method can accurately identify and locate key tissue regions where specific gene modules are functionally active, providing a precise tool for analyzing the dynamic changes in complex diseases and biological development.
[0076] Example 1
[0077] Reference Figure 2 This embodiment provides a method for identifying regions with high gene module scores, including:
[0078] Step S100: Obtain spatial transcriptome data containing gene expression information and spatial coordinate information of multiple cells; divide the cells into multiple unit grids according to the spatial coordinate information of the cells; determine the corresponding grid spatial coordinates for each unit grid; and number the unit grids to form an ordered arrangement of unit grids.
[0079] The purpose of this step is to discretize the continuous tissue space to facilitate computation. First, we receive raw biological data containing thousands of cells, each with two key pieces of information: which genes it expresses (gene expression information) and its precise location on the tissue slice ((x,y) spatial coordinates).
[0080] Next, the entire organizational space is processed, like drawing a grid on a map, dividing it into small, regular cell grids. Then, based on the (x,y) coordinates of each cell, it is determined which grid it falls into.
[0081] Finally, to ensure that each grid cell also has a representative location, a single coordinate point, known as the "grid space coordinate," is calculated for each cell.
[0082] After processing, the original, unordered cell point cloud data is transformed into a structured mesh diagram. This includes the correspondence between each cell and its corresponding cell within the same mesh, as well as a representative spatial coordinate for each individual cell.
[0083] In the step "and numbering the unit grids to form an ordered unit grid," after dividing the cells into grids and determining the coordinates of the grids, each individual unit grid is assigned a unique sequence number representing its order. Its core purpose is not merely to give the grids a "name," but to transform these grids arranged in two-dimensional space into a one-dimensional list or sequence with a clear order, i.e., an "ordered unit grid," through a specific numbering rule.
[0084] The processing involves traversing all the cell grids according to a pre-defined path rule and assigning them sequential numbers. This process establishes a mapping from two-dimensional space to a one-dimensional sequence, laying the foundation for subsequent steps that require sequential data processing.
[0085] The greatest advantage of this step is that it transforms a complex two-dimensional spatial analysis problem into a simpler, more algorithmically manageable one-dimensional sequence analysis problem. By forming an "ordered arrangement," subsequent algorithms (such as region segmentation algorithms) can linearly scan and compute along this sequence.
[0086] More importantly, the advantages become even more pronounced if a specific numbering rule that preserves spatial proximity is adopted. This numbering ensures that grids that are physically adjacent to each other are also substantially adjacent in the resulting one-dimensional sequence. This is crucial for accurately identifying spatially connected functional areas, effectively avoiding the problem of incorrectly classifying spatially scattered but similarly rated grids as the same region.
[0087] This step simplifies complex spatial problems into the analysis of a finite number of grids, greatly reducing computational complexity and providing a basic framework for the subsequent integration of spatial and genetic information.
[0088] Step S200: Calculate the module score value of each cell according to the preset gene set, and determine the high-scoring cell set based on the module score value.
[0089] This step shifts the focus from "space" to "gene function".
[0090] The aforementioned “preset gene set” refers to a group of genes known to work together to perform a certain biological function (e.g., genes related to immune responses).
[0091] This step calculates the overall activity level of this gene set within each cell, resulting in a comprehensive "module score," rather than analyzing individual genes. A higher score indicates greater activity of that gene set within that cell. After calculating scores for all cells, a threshold is set to filter out the cells with the highest scores, forming a special "high-scoring cell set." This ensures that each cell is matched with a module score, and a "high-scoring cell set" comprised of the cells with the highest scores.
[0092] This approach allows for a holistic assessment of the functional status of a group of genes, reflecting complex biological processes more effectively than analyzing individual genes. Furthermore, by screening high-scoring cells, it can initially identify the most functionally active cell populations within a tissue.
[0093] Step S300: For each unit grid, calculate the integrated gene set score index based on the grid space coordinates of the unit grid, the cells included in the unit grid, and the corresponding module score values.
[0094] This step is the core of the entire method, as it integrates the spatial and gene function information obtained in the first two steps. This step calculates a comprehensive score called the "Integrated Gene Set Score Index (IGSI)" for each cell grid. This score is not simply an average of the module scores of cells within the grid, but rather a weighted calculation based on multiple factors. According to the description, the calculation is based on the grid's own spatial coordinates, as well as the attributes of each cell within the grid (such as the module score).
[0095] This gives each grid cell a unique, quantified IGSI value. The higher the value, the stronger the functional synergy of the target gene set within the spatial microenvironment of that grid cell.
[0096] This index is a highly condensed information indicator that simultaneously reflects the activity level of gene sets and their spatial distribution characteristics. In this way, the spatial quantification of gene set synergistic effects is achieved, solving the problem that existing technologies cannot measure the specific enrichment degree of gene sets in tissues.
[0097] While this step is described functionally, a concrete implementation involves a comprehensive weighted formula. For example, a formula could be constructed to calculate the IGSI value, where the contribution of each cell within the grid is obtained by multiplying the following components:
[0098] (1) Significance of cell scores: The higher the module score of a cell, the greater its contribution.
[0099] (2) Spatial weight: The closer a cell is to the center point of its grid (i.e., grid spatial coordinates), the greater its contribution. This reflects the calculation "based on grid spatial coordinates".
[0100] (3) Cell density weight: The denser the cells in the local area where the cell is located, the greater the contribution.
[0101] (4) Functional similarity weight: The more similar the gene expression pattern of a cell is to the average pattern of its grid, the greater its contribution.
[0102] (5) Add up the contribution values of all cells in the grid and then perform global normalization to obtain the IGSI value of the grid.
[0103] Step S400: Based on the integrated gene set score index of the unit grid, the unit grid is divided into regions, and high gene module score regions are determined.
[0104] This step aims to numerically identify biologically significant functional regions.
[0105] First, analyze the IGSI score map covering the entire tissue obtained in the previous step. Cellular cells with similar and spatially adjacent IGSI values are grouped into larger "regions." After region division, based on a series of preset criteria, the regions that best meet the standards are selected and ultimately identified as "high genotype module score regions." The final output is one or more high genotype module score regions clearly identified on the tissue spatial map. These regions represent the spatial locations where the preset gene sets are most functionally active and exhibit the strongest synergistic effects.
[0106] This step transforms discrete grid scores into functional regions with clear boundaries and biological significance, enabling precise identification and localization of target regions and providing clear guidance for subsequent biological research.
[0107] Example 2
[0108] Reference Figure 3 This embodiment provides a method for identifying regions with high gene module scores. Based on the aforementioned embodiment, in step S100, the cells are divided into multiple unit grids according to their spatial coordinate information, and the corresponding grid spatial coordinates are determined for each unit grid, including:
[0109] Step S110: Based on the spatial coordinate information of all cells, determine the maximum and minimum values of the x-axis and the y-axis in the spatial transcriptome data.
[0110] The purpose of this step is to define the spatial extent of the entire tissue sample. It does this by processing the location information of all cells in the dataset to find the boundaries of the entire cell population. Specifically, the process involves iterating through the (x, y) coordinates of all cells and finding the maximum value (x, y) among all x-coordinates. max ) and minimum value (x) min ), and the maximum value among all y coordinates (y max ) and minimum value (y) min ).
[0111] After performing this step, you will get the four values mentioned above. These four values together define a minimum rectangular region that can completely surround all cells in the tissue sample.
[0112] This step provides a clear and standardized analysis range for subsequent gridding processing, ensuring that the analysis range can cover all cells, avoiding data omissions, and laying the foundation for establishing a unified grid system.
[0113] Step S120: According to the preset meshing quantity parameter, the coordinate ranges of the x-axis and the y-axis are divided into equal parts to form multiple unit meshes.
[0114] This step involves drawing grid lines within the rectangular area defined in the previous step, thereby creating the basic spatial units for analysis.
[0115] Specifically, using a pre-defined "mesh quantity parameter" (e.g., denoted as N), the lengths of the x-axis and y-axis are each divided into N equal segments. These dividing lines crisscross, forming a mesh system composed of multiple identical-sized unit meshes.
[0116] This results in a regular grid system covering the entire organizational region, composed of multiple (e.g., N×N) unit grids. By dividing the space into equal parts, the continuous organizational space is discretized into standard, uniform analysis units, which greatly simplifies subsequent spatial quantization calculations and makes the analysis results from different locations comparable.
[0117] Furthermore, in step S100, the plurality of cell grids are numbered to form an ordered arrangement of cell grids, including:
[0118] All the unit grids are numbered sequentially according to their spatial location; wherein, the numbering is carried out along columns, with odd-numbered columns numbered continuously from top to bottom, and even-numbered columns numbered continuously from bottom to top.
[0119] This step defines specific rules for assigning unique sequence numbers to all cell grids. After the grid is formed, to facilitate subsequent algorithmic processing, these two-dimensional grids need to be converted into a one-dimensional ordered list. This step describes not a conventional left-to-right, row-by-row numbering method, but a special "serpentine" or "Z-shaped" scanning path.
[0120] The specific processing can be carried out strictly in the order of the "columns", and the numbering direction can be determined according to the parity of the column:
[0121] (1) Starting from the first column (odd column), number the grids from the top and move down, assigning consecutive numbers to each grid in the column.
[0122] (2) After completing the first column, move to the bottom grid of the second column (even-numbered column) so that its number is immediately following the last number of the first column, and then move up to assign consecutive numbers to each grid of the column in turn.
[0123] (3) After completing the second column, move to the top of the third column (odd column) and repeat the downward numbering process.
[0124] (4) Continue alternating in this way until all columns of the grid are numbered.
[0125] Each cell grid is assigned a unique, sequential numerical designation.
[0126] Through the above numbering process, a one-dimensional, ordered sequence of cell grids is finally obtained. The special feature of this sequence is that its sorting method preserves, to the greatest extent possible, the local proximity information in the original two-dimensional space. For example, refer to... Figure 4 The specific arrangement of the medium-sized cell grid is N=5, forming a matrix of 5×5=25 cell grids.
[0127] Specifically, subsequent region segmentation algorithms (such as CBS) treat all grids as a one-dimensional sequence. If conventional row-by-row numbering is used, the grids at the end of the first row and the beginning of the second row are physically adjacent, but their numbers differ significantly. This "jump" in numbering can be misjudged by the segmentation algorithm, affecting the accuracy of region division. The "serpentine" numbering rule in this step ensures that adjacent grids in physical space (such as the bottom of the first and second columns) are also numbered consecutively, reducing this artificially created "jump." This allows subsequent segmentation algorithms to more accurately identify truly contiguous functional regions in physical space, improving the overall accuracy of the method. Furthermore, other algorithms acceptable in this identification method can be used instead of the CBS algorithm.
[0128] Furthermore, the magnitude of the x-axis direction in each of the said cell grids is (Formula 1): ;
[0129] The magnitude of the y-axis direction is (Formula 2): ;
[0130] Among them, L x L represents the magnitude of the x-axis direction in the cell grid. y The x represents the magnitude of the y-axis direction in the cell grid; max Represents the maximum value on the x-axis; x min Represents the minimum value on the x-axis; y max Represents the maximum value on the y-axis; y min The minimum value of the y-axis is represented by N; N represents the number of meshing parameters; the total number of all cell meshes after equal division is N×N.
[0131] For example, if an organization's x-axis range is 0 to 100 and its y-axis range is 0 to 200, and the set meshing quantity parameter N is 10, then the x-axis will be divided into 10 segments (each segment is 10 in length), and the y-axis will also be divided into 10 segments (each segment is 20 in length), ultimately forming 100 10×10 unit meshes of size 10×20.
[0132] Step S130: Each cell is assigned to the corresponding unit grid based on the spatial coordinate information.
[0133] This step associates the biological entities (cells) with the spatial analysis units (cell grids) created in the previous step. The process involves determining which cell grid each cell in the dataset belongs to based on its own (x, y) spatial coordinates, and then "placing" it into that grid.
[0134] This results in a mapping relationship where each cell is uniquely assigned to a cell grid. At this point, each cell grid contains a set of cells that fall within its spatial range.
[0135] This step transforms the data from unstructured to structured, binding discrete cell points to a regular grid system, which is a necessary prerequisite for subsequent grid-based statistical and scoring calculations.
[0136] Step S140: Calculate the geometric centroid of the spatial coordinates of all cells included in each unit grid, and use the geometric centroid as the grid spatial coordinates of that unit grid.
[0137] This step aims to calculate a single coordinate point that represents the "center position" of each cell in a grid cell.
[0138] The aforementioned "geometric centroid" can be the arithmetic mean of the coordinates of all points. The process is as follows: for each cell grid, the x-coordinates of all cells contained within it are summed and averaged to obtain the x-coordinate of the centroid; similarly, the y-coordinates of all cells are summed and averaged to obtain the y-coordinate of the centroid.
[0139] This gives each cell a unique (x, y) coordinate pair, which is the "grid space coordinate" of that cell.
[0140] This step condenses the spatial information of multiple cells within a region into a representative point, greatly simplifying subsequent calculations that require measuring distances between grids or between cells. This coordinate point becomes the spatial "avatar" of the grid, enabling the quantification of complex inter-regional relationships.
[0141] For example, if a cell grid contains K cells with coordinates (x1, y1), (x2, y2), ..., (xK, yK), then the formula for calculating the geometric centroid (i.e., the grid spatial coordinates) (Gx, Gy) of that grid is:
[0142] (1) Formula 3: ;
[0143] (2) Formula 4: ;
[0144] If a single grid cell contains 3 cells with coordinates (1, 5), (2, 7), and (6, 3), then the grid space coordinates of this cell are x = (1+2+6) / 3 = 3 and y = (5+7+3) / 3 = 5, which is (3, 5).
[0145] Example 3
[0146] Reference Figure 5 This embodiment provides a method for identifying high gene module score regions. Based on the aforementioned embodiment, step S200, which calculates the module score value of each cell according to a preset gene set and determines a high-scoring cell set based on the module score values, includes:
[0147] Step S210: The AddModuleScore function is used to calculate the module score value of the preset gene set in each cell.
[0148] This step aims to quantify the overall activity of a specific set of genes for each cell. Instead of analyzing individual genes, this step evaluates a set of functionally related "preset gene sets" as a whole (i.e., a module).
[0149] The process can be to calculate the score of a set of genes (G in total) in each cell using the AddModuleScore function of the Seurat package in R.
[0150] This function takes a preset list of genes and gene expression data for each cell as input, and then calculates a single, comprehensive "module score" for each cell through an internal algorithm.
[0151] This assigns a numerical value, or module score, to each cell in the dataset. The higher the score, the stronger the overall expression level or biological activity of the preset gene set in that cell.
[0152] This step condenses the complex expression information of multiple genes into a single quantitative indicator, simplifying the analysis process. It can more accurately reflect the biological functional state resulting from the synergistic effect of multiple genes, rather than looking at individual genes in isolation.
[0153] Step S220: Sort the module score values from largest to smallest, and extract the cells before the preset ranking after sorting as the high-scoring cell set; and calculate the mean and standard deviation of the module score values of the high-scoring cell set.
[0154] The purpose of this step is to identify and isolate the cell populations with the most potent genetic modules from all cells. The process consists of two steps:
[0155] First, arrange all cells from highest to lowest according to the module score calculated in the previous step.
[0156] Secondly, based on a "preset ranking" as the filtering criterion, the top-ranked cells are extracted from this sorted list. This "preset ranking" can be a percentage or a specific quantity.
[0157] In some implementations, the default ranking is the top 5%.
[0158] This results in a special subset of cells, known as the "high-scoring cell set." This set contains the cells in the entire tissue where the target gene set is most active.
[0159] This step, through sorting and filtering, precisely identifies functionally "hotspot" cells, providing a high-quality, highly active cell population sample for subsequent analysis, making the analysis more targeted.
[0160] In practical implementation, this corresponds to standard programming sorting operations. For example, the "preset ranking" can be flexibly set. If set to "top 5%", in a dataset containing 2000 cells, this step will select the top 100 cells by module score. If set to "top 50", then regardless of the total number of cells, only the top 50 cells will be selected.
[0161] In this step, the module scores can be sorted from highest to lowest, and the top 5% of cells can be extracted to form a high-scoring cell set. top Next, calculate the Cell. top The mean μ of the ratings top and standard deviation δ top .
[0162] Example 4
[0163] This embodiment provides a method for identifying high gene module scoring regions. Based on the aforementioned embodiment, the calculation expression for the integrated gene set scoring index (Formula 5) is as follows:
[0164] ;
[0165] Among them, IGSI w The score index represents the integrated gene set.
[0166] In the above expression, of The Euclidean distance represents the spatial coordinates of cell c to the grid spatial coordinates of its cell cell w. The attenuation coefficient represents the mesh size; in this expression, ;in, Let represent the distance between any cell l and any cell j in the cell grid w, and This calculation expression can preserve the characteristic that cells that are physically closer to each other contribute more to the grid, which is the spatial weight value of the cells.
[0167] In the above expression, S c When calculating the integrated gene set score index of a cell grid, the module score value of cell c in that cell grid is used; S ci This represents the module score value of any cell ci in any cell i when all cell grids are summed.
[0168] In the above expression, Cell C represents the set of high-scoring cells. The probability within the interval. In this expression, The calculation formula is:
[0169] ;
[0170] in, Cells representing high-scoring cells top The average score of all cell modules in the middle; Cells representing high-scoring cells top The standard deviation of the score values for all cell modules.
[0171] In the above expression, of Indicates the cell density sensitivity coefficient; represents the local cell density; f represents the overall or average cell density within a single grid cell w. For these two variables, the specific calculation methods are as follows:
[0172] (1) The calculation formula is: ;in, Let (x, y) represent the spatial coordinates of cell c, and K represent the number of cells, indicating the cell that is spatially closest to cell c; here, K can take the value 10.
[0173] The bandwidth parameter representing the control space attenuation is calculated using the following formula: , where d cj Let be the distance between cell c and any cell j in the cell grid w.
[0174] (2) The formula for calculating f is: Where KW represents the number of cells in the unit grid w. ; Let represent the distance between any cell l and any cell j in the cell grid w, and .
[0175] It should be noted that, generally speaking, the denser the cells, the greater the opportunity for physical contact between cells and neighboring cells, resulting in a stronger concentration gradient of signaling molecules. Therefore, this calculation can provide some feedback on the spatial cell density of tissue distribution.
[0176] In the above expression, The correlation weight value between cell c within cell w and cell w represents the correlation weight value between cell c and cell w; where, The gene expression vector representing cell c; The average expression profile of cells in the unit grid w represents the average expression profile of cells in the unit grid w.
[0177] In the expression, The calculation formula is: ; The calculation formula can be , and These can represent norms.
[0178] Example 5
[0179] Reference Figure 6 This embodiment provides a method for identifying high gene module score regions. Based on the aforementioned embodiment, step S400, which involves dividing the unit grid into regions according to the integrated gene set score index of the unit grid and determining the high gene module score regions therein, includes:
[0180] Step S410: Divide all the ordered unit grids according to the integrated gene set scoring index to form multiple regions to be analyzed with different means.
[0181] In this step, the previously obtained Integrated Gene Set Score Index (IGSI) map, which represents the entire tissue, is "cut" along a predetermined path to identify functionally similar regions.
[0182] The core here is "segmentation according to their arrangement order," which indicates that the segmentation operation is not performed arbitrarily in two-dimensional space, but rather acts on an ordered sequence that has already transformed the two-dimensional grid into one dimension. This "arrangement order" can be determined by the unique "snake-like" (S-shaped) numbering rule in the previous steps.
[0183] The process involves first arranging all cell grids into a one-dimensional list according to their "S-shaped" numbering order, where each element is the IGSI value of the corresponding grid. The algorithm then scans this ordered list, identifying "mutation points" based on changes in the IGSI values. When the algorithm detects a segment of continuous grids where the mean IGSI value differs significantly from the next segment, it performs a "segmentation" at that point.
[0184] After processing, multiple "regions to be analyzed" are obtained. Each region consists of one or more continuous cell grids from the sequence. Therefore, these regions also have high continuity and clustering in the original two-dimensional space. At the same time, the IGSI value within each region is relatively stable and has an average IGSI value different from that of its neighboring regions.
[0185] The greatest advantage of this step lies in its clever use of "arrangement order" to preserve local spatial features. Because the original ordered arrangement ensures that physically adjacent grids are also basically adjacent in the sequence, the segmentation algorithm can more accurately identify functional regions that are truly connected in the tissue, rather than some spatially scattered but similarly scored grid points, thus greatly improving the accuracy and biological significance of region identification.
[0186] In this embodiment, the segmentation process can be implemented using the Circular Binary Segmentation (CBS) algorithm. For example, suppose there is a 3x3 grid, numbered according to an ordered "S-shaped" rule, with the sequence 1-2-3-4-5-6-7-8-9. Their IGSI value sequence might be as follows: [0.1, 0.15, 0.8, 0.9, 0.85, 0.75, 0.2, 0.1, 0.25]. When processing this sequence, the CBS algorithm will find that the grids numbered 3 to 6 (IGSI values 0.8, 0.9, 0.85, 0.75 respectively) form a clear high-resolution plateau. Therefore, it might segment this sequence into three "regions to be analyzed":
[0187] (1) Region 1 to be analyzed: composed of grids 1 and 2 (low partition).
[0188] (2) Region 2 to be analyzed: composed of grids 3, 4, 5 and 6 (high partition).
[0189] (3) Region 3 to be analyzed: composed of grids 7, 8 and 9 (low partition).
[0190] Since the segmentation is based on an arrangement order that preserves spatial proximity, these three regions are also continuous block-shaped regions in two-dimensional space.
[0191] Step S420: From the multiple regions to be analyzed, determine the region that simultaneously meets all preset conditions as the high gene module score region.
[0192] The preset conditions include:
[0193] A. The mean of the integrated gene set scoring index of the region to be analyzed is within the preset ranking range in all regions;
[0194] B. The region to be analyzed contains at least one cell from the set of high-scoring cells.
[0195] This step aims to screen out truly high-functional-activity regions from all segmented "regions to be analyzed." A dual criterion is used: a region must not only have a sufficiently high score, but it must also contain the most functionally active "seed cells." The process involves checking each region simultaneously for both conditions A and B. Only regions that meet both conditions are ultimately identified as "high-gene module score regions." The final output is one or more precisely identified and located "high-gene module score regions."
[0196] This dual filtering mechanism ensures the reliability of the identification results. Condition A guarantees the statistical significance of the region (high overall score), while condition B guarantees its biological authenticity (containing confirmed highly active cells), effectively eliminating false positive results that may be caused by noise or calculation bias.
[0197] In some implementations, the preset ranking range is the top 40%, that is, the IGSI values are ranked from largest to smallest in the top 40% of regions.
[0198] This step involves HGMSR identification and determination. HGMSR stands for High Gene Module Score Region. The specific process involves checking whether each region to be analyzed meets both conditions A and B. Only regions that meet both conditions are ultimately identified as "High Gene Module Score Regions." The final output will then be one or more precisely identified and located "High Gene Module Score Regions."
[0199] This dual filtering mechanism ensures the reliability of the identification results. Condition A guarantees the statistical significance of the region (high overall score), while condition B guarantees its biological authenticity (containing confirmed highly active cells), effectively eliminating false positive results that may be caused by noise or calculation bias.
[0200] For condition A: "Within the preset ranking range" means that a threshold can be set. For example, if 10 regions are obtained after segmentation, the "preset ranking range" can be set to "top 40%". Then, the algorithm will calculate the average IGSI value of each region and only keep the regions with the highest average value in the top 4 for the next step of filtering.
[0201] For condition B: This is an existence check. The algorithm will traverse all cell grids within a given region to be analyzed, and then check whether there is at least one cell from the set of high-scoring cells among all the cells contained in these grids.
[0202] Example 5
[0203] To better illustrate the methods used in the foregoing embodiments and to verify their feasibility, this embodiment analyzes specific external data.
[0204] 1. Experimental Method:
[0205] (1) Data acquisition:
[0206] 1) Dataset ID:STDS0000247 in the BGI Spatiotemporal Omics Database (STOmics DB) is a public dataset from a Cell academic journal article (PMID:39500323). Multiple stereo-seq sequencing data samples were downloaded and tested.
[0207] 2) The experiment used publicly available data from one of the 10x gemomics numbers GSE169749, with the sample name GSM5213483. It was derived from an article in Nature Communications, with NCBI accession number PMID:35149721.
[0208] (2) Data Analysis:
[0209] After converting the downloaded data file into a Seurat object file and obtaining the data results (spatial transcriptome data), the high gene module score region identification method described in the previous embodiment was used for analysis.
[0210] 2. Experimental Results and Analysis:
[0211] refer to Figure 7 and Figure 8 Based on the results obtained from this solution, the distribution characteristics and patterns of HGMSR (high gene module score region) in different platforms and tissues can be clearly identified.
[0212] refer to Figure 7The results clearly pinpointed the lymphoid structures in the tissue (A). Specifically, the pink dotted areas represent high gene module score regions (HGMSR), while the gray dotted areas represent non-high gene module score regions (non-HGMSR) outside of HGMSR. This matches the location of the lymphoid structures on the H&E staining map (B) given in the literature PMID:35149721.
[0213] refer to Figure 8 This method was applied to exploratory studies of aging patterns in various mouse tissues. As shown in the figure, this method successfully identified aging-related regions (HGMSRs indicated by the pink area in the figure) in various tissues, including the lung, heart, liver, lymph nodes, intestine, hippocampus, spinal cord, spleen, and testis. The results clearly show that distinct spatial distribution regions of aging-related patterns exist in different tissues. More importantly, the aging patterns exhibit significant spatial consistency across multiple biological replicates within the same tissue (e.g., three liver tissue samples or two heart tissue samples in the figure). These findings highly match the aging-related regions reported in the original literature (PMID: 39500323), fully demonstrating the high feasibility, reliability, and significant academic application value of this method in identifying the spatial characteristics of specific biological processes.
[0214] The above results demonstrate that the method for identifying high gene module scoring regions provided in the preceding embodiments offers an efficient, accurate, and widely applicable approach capable of identifying the spatial distribution characteristics of gene modules associated with specific biological processes or disease states across platforms and tissues. This method not only contributes to a deeper understanding of the functional layout of gene modules in tissues but also provides new tools and perspectives for disease mechanism research and clinical applications.
[0215] refer to Figure 9 This application also provides a device for identifying high gene module scoring regions, comprising:
[0216] The segmentation module 10 is used to acquire spatial transcriptome data containing gene expression information and spatial coordinate information of multiple cells, divide the cells into multiple unit grids according to the spatial coordinate information of the cells, and determine the corresponding grid spatial coordinates for each unit grid.
[0217] The calculation module 20 is used to calculate the module score value of each cell according to a preset gene set, and to determine the set of high-scoring cells according to the module score value;
[0218] The calculation module 20 is further configured to calculate the integrated gene set score index for each unit grid based on the grid space coordinates of the unit grid, the cells included in the unit grid, and the corresponding module score values;
[0219] The screening module 30 is used to divide the unit grid into regions based on the integrated gene set score index of the unit grid, and determine the high gene module score regions therein.
[0220] It is understood that the device in this embodiment corresponds to the high gene module scoring region identification method in the above embodiments, and the options in the above embodiments are also applicable to this embodiment, so they will not be described again here.
[0221] This application also provides a computer device, exemplary of which includes a processor and a memory, wherein the memory stores a computer program, and the processor, by running the computer program, causes the computer device to perform the functions of the various modules in the above-described method for identifying high genotype module scoring regions or the above-described device for identifying high genotype module scoring regions.
[0222] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0223] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving execution instructions.
[0224] This application also provides a computer storage medium for storing the computer program used in the aforementioned computer device. The computer storage medium can be a readable storage medium, a non-volatile storage medium, or a volatile storage medium. For example, the computer storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0225] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, in alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0226] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0227] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0228] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for identifying a high genomic module score region, characterized in that, The method comprises the following steps: acquiring spatial transcriptome data containing gene expression information and spatial coordinate information of a plurality of cells, dividing the cells into a plurality of unit grids according to the spatial coordinate information of the cells, determining corresponding grid spatial coordinates for each unit grid, and numbering the unit grids to form an orderly arranged unit grid; calculating a module score value of each cell according to a preset gene set, and determining a high-score cell set according to the module score value; for each unit grid, calculating an integrated gene set score index based on the grid spatial coordinates of the unit grid and the cells included in the unit grid and the corresponding module score value; the calculation expression of the integrated gene set score index is: ; wherein IGSI w represents the integrated gene set score index; In the expression, of represents the Euclidean distance from the spatial coordinates of a cell c to the grid spatial coordinates of the cell grid w in which it is located; represents the attenuation coefficient of the grid; S represents the module score value of the cell c in the unit grid when the integrated gene set score index of the unit grid is calculated c S represents the module score value of the cell c in the unit grid when the integrated gene set score index of the unit grid is calculated ci S represents the module score value of the cell c in the unit grid when the integrated gene set score index of the unit grid is calculated In the expression, represents the probability that a cell c belongs to the high-scored cell set within the interval. In the expression, of represents a cell density sensitivity coefficient; represents a local cell density; f represents an overall or average cell density within a unit grid w; represents the correlation weight value of cell c within unit grid w and unit grid w; wherein, represents the correlation weight value of cell c within unit grid w and unit grid w; wherein, represents the gene expression vector of cell c; represents the average expression profile of cells in unit grid w; dividing the unit grids into regions according to the integrated gene set score index of the unit grids and determining a high gene module score region therein.
2. The method of claim 1, wherein the high genomic module score region is identified by the steps of: calculating a genomic module score for each genomic module in the genomic region; and identifying the high genomic module score region as the genomic module with the highest genomic module score. The step of dividing the cells into a plurality of unit grids according to the spatial coordinate information of the cells and determining corresponding grid spatial coordinates for each unit grid comprises the following steps: determining the maximum and minimum values of the x-axis and the maximum and minimum values of the y-axis in the space of the spatial transcriptome data according to the spatial coordinate information of all cells; dividing the coordinate ranges of the x-axis and the y-axis into equal parts respectively according to a preset grid number parameter to form a plurality of unit grids; assigning each cell to the corresponding unit grid based on the spatial coordinate information; calculating the geometric center of the spatial coordinates of all cells included in each unit grid, and taking the geometric center as the grid spatial coordinates of the unit grid.
3. The method for identifying high gene module scoring regions as described in claim 2, characterized in that, The step of numbering the plurality of unit grids to form an orderly arranged unit grid comprises the following steps: sequentially numbering all unit grids according to their spatial positions; wherein the numbering is along the columns, and the odd columns are numbered continuously from top to bottom, and the even columns are numbered continuously from bottom to top.
4. The method for identifying high gene module scoring regions as described in claim 2, characterized in that, The size of the x-axis direction in each of the unit grids is: ; The size in the y-axis direction is: ; wherein, L x represents the size of the x-axis direction in the unit grid; L y represents the size of the y-axis direction in the unit grid; x max represents the maximum value of the x-axis; x min represents the minimum value of the x-axis; y max represents the maximum value of the y-axis; y min represents the minimum value of the y-axis; N represents the grid number parameter; the total number of all unit grids after equal division is N x N.
5. The method of claim 1, wherein the high gene module score region is identified by the steps of: calculating a gene module score for each gene module; and identifying a high gene module score region as a region in which the gene module score is higher than a predetermined threshold value. The step of calculating a module score value of each cell according to a preset gene set and determining a high-score cell set according to the module score value comprises the following steps: calculating the module score value of the preset gene set in each cell by using the AddModuleScore function; sorting the module score values from large to small, and extracting the cells before the preset ranking in the sorted order as the high-score cell set; and calculating the mean and standard deviation of the module score values of the high-score cell set.
6. The method of claim 1, wherein the high gene module score region is identified by the steps of: calculating a gene module score for each gene module; and identifying a high gene module score region as a region in which the gene module score is higher than a predetermined threshold value. In the expression, ; wherein, represents the distance between any cell l and any cell j in the unit grid w, and .
7. The method of claim 1, wherein the high gene module score region is identified by the steps of: calculating a gene module score for each gene module; and identifying a high gene module score region as a region in which the gene module score is higher than a predetermined threshold value. The expression is The calculation formula is: ; wherein, represents the spatial coordinates (x, y) of the cell c, K represents the cell number, and represents the cell closest in space to the cell c; represents a bandwidth parameter for controlling spatial attenuation, and the calculation formula is , wherein d cj is the distance between the cell c and any cell j in the unit grid w.
8. The method for identifying high gene module scoring regions as described in claim 1, characterized in that, In the expression, the calculation formula of f is: wherein KW represents the number of cells in the unit grid w, ; represents the distance between any cell I and any cell j in the unit grid w, and .
9. The method for identifying high gene module scoring regions as described in claim 1, characterized in that, In the expression, The calculation formula is: .
10. The method for identifying high gene module scoring regions as described in claim 1, characterized in that, In the expression, The calculation formula of ; In the expression, The calculation formula is: ; Cell representing the high scoring cell set Cell top the mean of all cell module score values in Cell Cell representing the high scoring cell set Cell top the standard deviation of all cell module score values in Cell 11. The method for identifying high gene module scoring regions as described in claim 1, characterized in that, The step of dividing the unit grids into regions according to the integrated gene set score index of the unit grids and determining a high gene module score region therein comprises the following steps: segmenting all ordered unit grids according to the integrated gene set score index to form a plurality of analysis regions with different means; determining a region that meets all preset conditions at the same time from the plurality of analysis regions as the high gene module score region; wherein the preset conditions include: A. The mean of the integrated gene set score index of the analysis region is within the preset ranking range among all regions. B. The cells in the at least one high-score cell set in the region to be analyzed.
12. A device for identifying high gene module scoring regions, characterized in that, Comprise: A segmentation module configured to obtain spatial transcriptome data comprising gene expression information and spatial coordinate information of a plurality of cells, divide the cells into a plurality of unit grids according to the spatial coordinate information of the cells, and determine corresponding grid spatial coordinates for each of the unit grids; A calculation module configured to calculate a module score value of each of the cells according to a preset gene set, and determine a high-score cell set according to the module score value; The calculation module is further configured to, for each of the unit grids, calculate an integrated gene set score index based on the grid spatial coordinates of the unit grid, and the cells included in the unit grid and the corresponding module score values; and the calculation expression of the integrated gene set score index is: ; wherein IGSI w represents the integrated gene set score index; In the expression, of represents the Euclidean distance of the spatial coordinates of a cell c to the grid spatial coordinates of the cell grid w in which it is located; represents the attenuation coefficient of the grid; S represents the module score value of the cell c in the unit grid when the integrated gene set score index of the unit grid is calculated c S represents the module score value of the cell c in the unit grid when the integrated gene set score index of the unit grid is calculated ci S represents the module score value of the cell c in the unit grid when the integrated gene set score index of the unit grid is calculated In the expression, represents the probability that a cell c belongs to the high-scored cell set within the interval In the expression, of represents a cell density sensitivity coefficient; represents a local cell density; f represents an overall or average cell density within a unit grid w; represents the correlation weight value of cell c within the unit grid w and the unit grid w; wherein, represents the correlation weight value of cell c within the unit grid w and the unit grid w; wherein, represents the gene expression vector of cell c; represents the average expression profile of cells in the unit grid w; A screening module configured to divide the unit grids into regions according to the integrated gene set score index of the unit grids, and determine a high gene module score region therefrom.
13. A computer device, comprising: The computer device comprises a processor and a memory, the memory stores a computer program, and the processor is configured to execute the computer program to implement the high gene module score region identification method in any one of claims 1-11.
14. A computer storage medium, characterized in that The computer program is stored in the memory and executed on the processor to implement the high gene module score region identification method in any one of claims 1-11.
Citation Information
Patent Citations
Spatial domain identification method in spatial transcriptomics based on deep graph learning
CN117708628A
Gene module analysis method, device and equipment and storage medium
CN120877881A