Gene expression analysis method and device, electronic device and storage medium

By clustering and gradient-direction division of the spatial transcriptome data of the target slices, the problem that gene expression analysis in existing technologies cannot capture changes within cluster regions is solved. Gene expression analysis in the biological gradient direction is realized, and spatial gradients and directional changes are quantified.

CN120356521BActive Publication Date: 2025-09-26BGI RESEARCH HANGZHOU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510847964.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-26
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Existing technologies have difficulty capturing the spatial continuity and local variation characteristics within cluster regions in gene expression analysis, especially in scenarios where regional boundaries are fuzzy or gene expression gradients are obvious, making it impossible to perform detailed spatial gradient change analysis.

Method used

By clustering the spatial transcriptome data of the target slices, a regional distribution map is generated, the target analysis area is determined and divided into sub-regions based on the biological gradient direction, the regional expression level of the target gene in each sub-region is obtained, and gene expression analysis is performed to describe the changing trend of the gene in the biological gradient direction.

Benefits of technology

It realizes gene expression analysis based on specific directions within the same biological functional area or irregular target analysis area, breaking through the limitations of traditional methods and being able to quantify the spatial gradients and directional changes of gene expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356521B_ABST
    Figure CN120356521B_ABST
Patent Text Reader

Abstract

The present disclosure provides a gene expression analysis method and device, electronic device and storage medium. The gene expression analysis method includes: clustering multiple capture sites according to the spatial transcriptome data of the target slice to obtain biological functional areas; generating a regional distribution map according to the spatial position information of the biological functional areas and each capture site; receiving a region selection operation to determine the target analysis area; determining the biological gradient direction, and dividing the target analysis area into sub-areas according to the biological gradient direction; obtaining the regional expression amount of the target gene in each sub-area according to the spatial transcriptome data; performing gene expression analysis according to the biological gradient direction and the regional expression amount corresponding to each sub-area; gene expression analysis is used to describe the changing trend of the gene expression of the target gene along the biological gradient direction in the target analysis area. The embodiment of the present application can realize gene expression analysis based on the gradient direction within the target analysis area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of bioinformatics, and in particular to a gene expression analysis method and device, an electronic device, and a storage medium. Background Art

[0002] Spatial transcriptomics is a technology that can simultaneously achieve gene expression analysis and cell spatial localization in tissue sections.

[0003] Currently, clustering algorithms based on similarity in gene expression patterns can be used to divide spatial points in tissue sections into multiple discrete cluster regions. However, this method can only compare gene expression differences between different cluster regions, resulting in certain limitations in gene expression analysis. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a gene expression analysis method and device, electronic equipment and storage medium, aiming to realize gene expression analysis based on gradient direction within the target analysis area.

[0005] To achieve the above objectives, the first aspect of the embodiments of the present application provides a gene expression analysis method, the method comprising:

[0006] Clustering multiple capture sites based on the spatial transcriptome data of the target slices to obtain multiple biological functional regions;

[0007] generating a regional distribution map according to the spatial position information of the plurality of biological functional regions and each of the capture sites;

[0008] receiving a region selection operation to determine a target analysis region, wherein the region selection operation indicates selecting all or part of the first biological function region and / or all or part of the second biological function region in the region distribution map;

[0009] Determining a biological gradient direction, and dividing the target analysis area into a plurality of sub-areas according to the biological gradient direction;

[0010] Obtaining the regional expression level of the target gene in each of the sub-regions according to the spatial transcriptome data;

[0011] Gene expression analysis is performed according to the biological gradient direction and the regional expression amount corresponding to each sub-region; the gene expression analysis is used to describe the change trend of gene expression of the target gene along the biological gradient direction in the target analysis region.

[0012] In some embodiments, each of the capture sites includes gene expression information; and obtaining the regional expression level of the target gene in each of the sub-regions based on the spatial transcriptome data includes:

[0013] Aligning the image coordinate system corresponding to the regional distribution map with the spatial coordinate system corresponding to the spatial transcriptome data to establish a positional correspondence between pixel points and capture sites;

[0014] For each image block of the sub-region in the region distribution map, construct a first data set according to pixel coordinates of each pixel point in the image block and the region label of the sub-region;

[0015] Determining, based on the positional correspondence between the pixel points and the capture sites and in combination with the first data set, a region label corresponding to the capture site;

[0016] constructing a second data set based on the gene expression information, spatial location information, and the region labels corresponding to the capture sites;

[0017] Based on the second data set, obtaining multidimensional data of the target gene at the corresponding capture site; the multidimensional data includes gene expression information, spatial location information and the region label;

[0018] For each of the sub-regions, the regional expression level of the target gene in the sub-region is counted based on the multi-dimensional data of the target gene.

[0019] In some embodiments, a method for establishing a positional correspondence between a pixel point and a capture site includes:

[0020] For each pixel point in the area distribution map, scaling the pixel coordinates of the pixel point according to the image resolution of the area distribution map;

[0021] Determining a coordinate offset according to the image coordinate system of the regional distribution map and the spatial coordinate system corresponding to the spatial transcriptome data, and correcting the pixel coordinates after the scaling process according to the coordinate offset;

[0022] The spatial position information of the capture site and the corrected pixel coordinates are used to perform distance calculation, and a position correspondence relationship between the capture site and the corresponding pixel point is established based on the distance calculation result.

[0023] In some embodiments, performing gene expression analysis based on the biological gradient direction and the regional expression level corresponding to each sub-region includes:

[0024] Sorting the coordinates corresponding to the target gene in each sub-region according to the biological gradient direction to obtain a coordinate sequence;

[0025] A gradient curve is generated according to the coordinate sequence and the regional expression amount of the target gene in the sub-region, and the gradient curve is used to represent the change trend of the gene expression of the target gene in the biological gradient direction.

[0026] In some embodiments, dividing the target analysis region into a plurality of sub-regions according to the biological gradient direction includes:

[0027] Segmenting the target analysis region according to the biological gradient direction and a preset number of segments to obtain a plurality of segmented contours;

[0028] Extracting a plurality of minimum closed areas according to the plurality of segmented contours, and performing a boundary check on each of the minimum closed areas;

[0029] When the boundary check result indicates that the corresponding minimum closed area is within the target analysis area, the minimum closed area is used as the sub-area.

[0030] In some embodiments, the method for determining the direction of the biological gradient includes any of the following:

[0031] Determining the biological gradient direction according to the expression trend of the target gene;

[0032] Alternatively, a first area of ​​the target tissue in the target analysis area and a second area of ​​the non-target tissue in the target analysis area are determined, and a biological gradient direction is determined based on the first area and the second area, where the biological gradient direction represents a direction from the first area to the second area.

[0033] To achieve the above objectives, a second aspect of the embodiments of the present application provides a gene expression analysis device, comprising:

[0034] A site clustering unit is used to cluster multiple capture sites based on the spatial transcriptome data of the target slice to obtain multiple biological functional regions;

[0035] an image generating unit, configured to generate a regional distribution map based on the spatial position information of the plurality of biological functional regions and each of the capture sites;

[0036] an analysis region determining unit, configured to receive a region selection operation and determine a target analysis region, wherein the region selection operation indicates selecting all or part of the first biological function region and / or all or part of the second biological function region in the region distribution map;

[0037] a region division unit, configured to determine a biological gradient direction and divide the target analysis region into a plurality of sub-regions according to the biological gradient direction;

[0038] an expression quantity acquisition unit, configured to acquire the regional expression quantity of the target gene in each of the sub-regions based on the spatial transcriptome data;

[0039] A gene expression analysis unit is used to perform gene expression analysis based on the biological gradient direction and the regional expression amount corresponding to each sub-region; the gene expression analysis is used to describe the change trend of gene expression of the target gene along the biological gradient direction in the target analysis region.

[0040] To achieve the above-mentioned purpose, a third aspect of an embodiment of the present application proposes an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.

[0041] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.

[0042] To achieve the above-mentioned purpose, the fifth aspect of the embodiments of the present application proposes a computer program product, which includes a computer program. The computer program is read and executed by a processor of a computer device, so that the computer device executes the method described in the first aspect.

[0043] The gene expression analysis method and device, electronic device and storage medium proposed in the embodiment of the present application can obtain any regular or irregular target analysis area by selecting all or part of the first biological function area in the regional distribution map, and / or selecting all or part of the second biological function area in the regional distribution map. In addition, by dividing the target analysis area in the biological gradient direction and obtaining the regional expression amount of the target gene in each divided sub-area, it is possible to analyze the expression of the target gene based on the biological gradient direction in the target analysis area. It can be seen from this that, compared to the method in which the related art can only compare the gene expression differences between different biological function areas, the embodiment of the present application is based on the selection of the target analysis area, and can realize gene expression analysis based on a specific direction (i.e., the biological gradient direction) within the same biological function area (or within any irregular target analysis area). BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The accompanying drawings are used to provide a further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation to the technical solution of the present disclosure.

[0045] Figure 1 is a flow chart of the gene expression analysis method provided in the embodiments of the present application;

[0046] Figure 2A This is a schematic diagram of extracting the overall outer contour of a space group chip provided in an embodiment of the present application;

[0047] Figure 2B is a schematic diagram of the regional distribution map provided in the embodiment of the present application;

[0048] Figure 2C This is a schematic diagram of color filling of various biological function areas in the regional distribution diagram provided in an embodiment of the present application;

[0049] Figure 3A 3 is a schematic diagram of selecting the target analysis area 301 provided in an embodiment of the present application;

[0050] Figure 3B 3 is a schematic diagram of selecting the target analysis area 301 provided in an embodiment of the present application;

[0051] Figure 4A 3 is a schematic diagram of the gradient direction corresponding to the target analysis area 301 provided in an embodiment of the present application;

[0052] Figure 4B 3 is a schematic diagram of segmenting the target analysis region 301 based on the gradient direction provided in an embodiment of the present application;

[0053] Figure 5A 3 is a schematic diagram of the gradient direction corresponding to the target analysis area 302 provided in an embodiment of the present application;

[0054] Figure 5B 3 is a schematic diagram of segmenting the target analysis region 302 based on the gradient direction provided in an embodiment of the present application;

[0055] Figure 6 yes Figure 1 Flowchart of an embodiment of step S105 in FIG.

[0056] Figure 7A 3 is a schematic diagram of color filling of each sub-area in the target analysis area 301 provided in an embodiment of the present application;

[0057] Figure 7B 3 is a schematic diagram of color filling of each sub-area in the target analysis area 302 provided in an embodiment of the present application;

[0058] Figure 8A This is a visualization diagram of the gene expression analysis of the target analysis region 301 provided in the embodiment of the present application;

[0059] Figure 8BThis is a visualization diagram of the gene expression analysis of the target analysis region 302 provided in the embodiment of the present application;

[0060] Figure 9 is a schematic diagram of a gene expression analysis device provided in an embodiment of the present application;

[0061] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not intended to limit the present disclosure.

[0063] Before further explaining the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations:

[0064] Spatial transcriptomics (ST) is a technology that analyzes RNA spatially and is used to analyze all RNA in a single tissue section. Each slide used for library construction by 10x Genomics Spatial Transcriptomics has four capture areas, each measuring 6.5x6.5mm. Each capture area contains 5,000 barcoded spots, each with a diameter of 55μm and a center-to-center distance of 100μm. Each spot includes multiple capture probes that can bind to RNA, each with unique spatial barcodes to mark the spatial location of the captured RNA. Through on-machine sequencing, the sequence of each RNA transcript can be mapped back to its original location in the tissue section, providing data support for downstream tasks. Therefore, ST sequencing technology is a technology that can simultaneously obtain RNA expression and RNA spatial location information in a single experiment.

[0065] Spatial transcriptomics technology is rapidly developing in the biomedical field. Unlike traditional transcriptome sequencing (such as single-cell RNA sequencing), spatial transcriptomics can precisely analyze the spatial distribution of gene expression within tissue sections. Combining histological and molecular biological information, spatial transcriptomics can be applied to tumor microenvironment research, neurological analysis, organ development mechanism analysis, and disease case studies. In oncology research, spatial transcriptomics can analyze the interactions between different cell types within the tumor core, margins, and microenvironment. In neuroscience, spatial transcriptomics can reveal the spatial network characteristics of neurons and their surrounding cells. In developmental biology and pathology research, spatial transcriptomics can precisely analyze molecular changes during organ development and the spatial patterns of disease development.

[0066] Mainstream spatial transcriptomics technologies share the following common characteristics: they can generate gene expression data for each spatial location while preserving spatial information. These technologies enable the creation of high-resolution spatial gene expression maps of the research subject. However, due to the complexity of tissues and cells, analytical methods in these technologies primarily focus on differential analysis between spatial domains of unsupervised clustering, neglecting the continuity and local variations of gene expression within spatial domains. This limitation not only makes it difficult to capture spatial gradients of gene expression (such as the gradual transition from the tumor core to the periphery) but also restricts the analysis of biological directional changes (such as the spatiotemporal dynamics of gene expression during development). In particular, for complex pathological tissues, such as areas of inflammation or tumor infiltration, gene expression often exhibits local gradients or fuzzy boundaries, making it difficult for existing methods to precisely quantify these key features.

[0067] In addition, traditional unsupervised clustering methods may not be able to accurately identify and cluster lesion areas that require focused analysis, because clustering results usually rely on the similarity of gene expression patterns, which may lead to low clustering accuracy for lesion areas with complex boundaries or gradual properties. Traditional clustering methods have certain limitations when dealing with dynamic gene expression changes with blurred boundaries or across regions, and cannot flexibly extract any spatial region of interest for detailed analysis. Therefore, relying on traditional clustering technology is difficult to meet the needs of accurate extraction and analysis of specific areas, especially lesion areas or other areas with special biological significance.

[0068] In summary, the gene expression analysis methods in related technologies have the following problems:

[0069] 1. Clustering algorithms based on similarity in gene expression patterns in related technologies divide spatial locations into discrete cluster regions. This method is often used to compare gene expression differences between cluster regions, but it struggles to capture spatially continuous gene expression gradients within a cluster region. This is especially true in scenarios with blurred region boundaries or pronounced gene expression gradients (such as the infiltrated transition zone from the core to the edge of a tumor). This makes it impossible to extract interesting regions for analysis of spatial gradient expression changes.

[0070] 2. When focusing on expression characteristics at spatial locations, related technologies often employ global clustering approaches. However, they lack specialized quantification methods for continuous expression changes within a region, local gradients, and directional changes (such as expression patterns along a specific axis or radius). This limits in-depth research on certain biological phenomena, such as cell differentiation, signal diffusion, and dynamic changes in the tumor microenvironment.

[0071] Based on this, the embodiments of the present application provide a gene expression analysis method and apparatus, an electronic device, and a storage medium, so as to enable gene expression analysis in a gradient direction within a target analysis region.

[0072] The gene expression analysis method provided in the embodiment of the present application relates to the field of bioinformatics. The gene expression analysis method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the gene expression analysis method, etc., but is not limited to the above forms.

[0073] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0074] The gene expression analysis method provided in the examples of this application is described below.

[0075] Reference Figure 1 In some embodiments, the gene expression analysis method provided in the embodiments of the present application includes but is not limited to steps S101 to S106.

[0076] Step S101, clustering multiple capture sites according to the spatial transcriptome data of the target slice to obtain multiple biological function regions;

[0077] Step S102, generating a regional distribution map based on the spatial location information of the multiple biological functional regions and each capture site;

[0078] Step S103, receiving a region selection operation and determining a target analysis region, wherein the region selection operation indicates selecting the entire region or a portion of the first biological function region in the region distribution map, and / or selecting the entire region or a portion of the second biological function region;

[0079] Step S104, determining the biological gradient direction, and dividing the target analysis area into multiple sub-areas according to the biological gradient direction;

[0080] Step S105, obtaining the regional expression level of the target gene in each sub-region based on the spatial transcriptome data;

[0081] Step S106 , performing gene expression analysis based on the biological gradient direction and the regional expression amount corresponding to each sub-region. The gene expression analysis is used to describe the change trend of gene expression of the target gene along the biological gradient direction in the target analysis region.

[0082] In the steps S101 to S106 shown in the embodiment of the present application, by selecting all or part of the first biological function area in the regional distribution map, and / or, selecting all or part of the second biological function area in the regional distribution map, any regular or irregular target analysis area can be obtained. In addition, by dividing the target analysis area in the biological gradient direction and obtaining the regional expression amount of the target gene in each divided sub-area, it is possible to analyze the expression of the target gene based on the biological gradient direction in the target analysis area. It can be seen from this that, compared to the method in which the related art can only compare the gene expression differences between different biological function areas, the embodiment of the present application is based on the selection of the target analysis area, and can realize gene expression analysis based on a specific direction (i.e., the biological gradient direction) within the same biological function area (or within any irregular target analysis area).

[0083] In step S101 of some embodiments, the target slice may refer to a slice of a tissue or sample to be subjected to gene expression analysis. Spatial transcriptome data integrates the gene expression information and spatial location information of the target slice. The gene expression information may include the expression level of the gene in each capture site (i.e., spot or cell), and the spatial location information may include the actual coordinates of each capture site in the target slice, i.e., each capture site may include spatial location information and gene expression values. By clustering the spatial transcriptome data using a clustering algorithm (such as a Bayesian unsupervised clustering algorithm), the multiple capture sites corresponding to the target slice can be divided into different spatial domains, resulting in multiple biological functional regions.

[0084] In step S102 of some embodiments, for each biological functional region, the boundary of the biological functional region can be drawn based on the coordinate information (i.e., spatial position information) of multiple capture sites in the biological functional region. In this way, a regional distribution map can be generated based on the boundary lines of multiple biological functional regions. For example, Figure 2A The figure shows the outline of the target slice obtained by the chip based on the spatial transcriptome technology. Based on the coordinate information of each biological functional area and the capture site, the outline of each biological functional area (i.e., the boundary line) can be determined, so that the following can be obtained: Figure 2B It is understood that it is also possible to Figure 2B The regional distribution maps shown are contoured to improve the accuracy of each biological function region.

[0085] In step S103 of some embodiments, a region selection operation may be performed on the region distribution map to determine a region of interest (i.e., target analysis region, ROI). It is understood that the region selection operation may refer to an operation of selecting the entire region of at least one biological function region in the region distribution map, i.e., it may refer to selecting the entire region of the first biological function region and / or selecting the entire region of the second biological function region. For example, when the region selection operation represents the operation of selecting the entire region of the first biological function region, Figure 2B When two biological function areas are selected from the regional distribution map shown in FIG, the following can be obtained: Figure 3A The target analysis area 301 is shown. In addition, the region selection operation can also refer to the operation of selecting a portion of at least one biological function area (i.e., it can refer to the selection of a portion of the first biological function area and / or the selection of a portion of the second biological function area). For example, the region selection operation can refer to the operation of selecting a region on the region distribution map based on an interactive tool. In this case, the following can be obtained: Figure 3B Target analysis region 302 is shown. In other words, taking region A as the first biological functional region and region B as the second biological functional region, region selection operations include, but are not limited to, the following operations: 1. Selecting a portion of region A; 2. Selecting the entire region of region A; 3. Selecting a portion of region B; 4. Selecting the entire region of region B; 5. Selecting a portion of region A + selecting a portion of region B; 6. Selecting the entire region of region A + selecting a portion of region B. It will be understood that although the above examples use the first and second biological functional regions, region selection operations are not limited to selecting only two biological functional regions. For example, region selection operations may also refer to performing related operations on the first, second, and third biological functional regions, and this is not specifically limited.

[0086] It is understandable that if Figure 2C As shown in , different areas in the regional distribution map can be filled with colors to distinguish different biological functional areas. For example, when the target slice is a slice corresponding to an AD hippocampal brain sample, the biological functional areas may include the CA1 area, CA2 area, CA3 area, CA4 area, DG area (anterior part of the hippocampus, serrated structure) and GC area (glial cell area), as shown in Figure 2C As shown, different colors can be used to distinguish the above areas. Figure 3A The region selection operation shown can refer to the entire area of ​​CA1 region, Figure 2C The entire area of ​​the middle area 201 is selected. Figure 3B The region selection operation shown may refer to a selection operation on a portion of the region between the CA1 region and the DG region.

[0087] In step S104 of some embodiments, the biological gradient direction may refer to a preset direction for gene expression analysis and having biological significance. The target analysis area may be divided into multiple sub-areas according to the biological gradient direction. For example, Figure 4A As shown, the biological gradient direction may refer to the direction from the boundary region between CA1 and CA2 to the region 201, so that the target analysis region 301 may be divided into the following Figure 4B Alternatively, Figure 5A As shown, the biological gradient direction may refer to the direction from the CA1 region to the DG region, so the target analysis area 302 may be divided into the following Figure 5B Multiple sub-areas shown.

[0088] In some embodiments, in steps S105 and S106, after obtaining multiple subregions, the gene expression value of the target gene at the capture site corresponding to the subregion can be determined for each subregion based on the spatial transcriptome data. The regional expression level of the target gene in the subregion can be obtained by integrating the gene expression values ​​of the corresponding sites (e.g., taking the mean value). The regional expression levels of the target gene in each subregion are sorted based on the biological gradient direction, thereby enabling analysis of the expression change trend of the target gene based on the biological gradient direction in the target analysis region.

[0089] In some embodiments, the method for determining the direction of the biological gradient includes any of the following:

[0090] Determine the direction of the biological gradient based on the expression trend of the target gene;

[0091] Alternatively, a first region of the target tissue in the target analysis region and a second region of the non-target tissue in the target analysis region are determined, and a biological gradient direction is determined based on the first region and the second region, where the biological gradient direction represents a direction from the first region to the second region.

[0092] In embodiments of the present application, the direction of the biological gradient can be determined based on biological hypotheses, spatial organizational structures, specific directions, tissue types, and the like. A biological hypothesis can refer to determining the direction of the biological gradient based on the expression trend of a specific gene (e.g., a target gene). Taking the determination of the direction of the biological gradient in a tumor microenvironment as an example, the biological hypothesis can refer to a gradient variation in gene expression between the tumor center and surrounding normal tissue. For example, hypoxia-related genes (e.g., HIF1A and VEGFA) are more highly expressed in the tumor core and less expressed in the edge regions. Thus, the direction of the biological gradient can be determined to be the direction extending outward from the tumor center (e.g., toward the edge regions). In this case, analyzing the expression of the target gene based on the biological gradient direction allows for the study of how the gene changes with distance.

[0093] Spatial organization structure can refer to determining the direction of the biological gradient based on the specific structure contained in the target slice. For example, the target slice may include a lesion area and healthy tissue, so the direction of the biological gradient can refer to the direction from the lesion area to the healthy tissue.

[0094] The specific direction may refer to a horizontal gradient (ie, X-axis) direction, a vertical gradient (ie, Y-axis) direction, or other set directions (eg, oblique directions, etc.).

[0095] Tissue type can refer to the type of target tissue to be studied. For example, when the target tissue is a lesion, non-lesion areas (such as healthy tissue) can be used as non-target tissue. In this way, the area within the target analysis region where the target tissue is located (i.e., the first area) can be used as the starting point, and the area within the target analysis region where the non-target tissue is located (i.e., the second area) can be used as the endpoint. Thus, the direction of the biological gradient can be derived based on the determined starting and endpoints. However, it is understood that in addition to using healthy tissue as non-target tissue, tissues with a strong correlation with the lesion area can also be used as non-target tissue based on analysis requirements, thereby constructing biological gradient directions in other directions.

[0096] The method for determining the direction of a biological gradient in a variety of ways in the embodiments of the present application can improve the flexibility of determining the direction of a biological gradient, thereby improving the flexibility of gene expression analysis based on the direction of the biological gradient.

[0097] In some embodiments, step S104 of "dividing the target analysis region into multiple sub-regions according to the biological gradient direction" may include the following sub-steps:

[0098] The target analysis area is segmented according to the biological gradient direction and the preset number of segments to obtain multiple segmented contours;

[0099] Extract multiple minimum closed areas based on multiple segmented contours, and perform boundary inspection on each minimum closed area;

[0100] When the boundary check result indicates that the corresponding minimum closed area is within the target analysis area, the minimum closed area is used as a sub-area.

[0101] In an embodiment of the present application, the number of segments may refer to the number of sub-regions into which the target analysis region is expected to be divided. The specific value of the number of segments may be adaptively set according to actual needs. The target analysis region is segmented according to the biological gradient direction and the number of segments to generate multiple segmentation contour lines (i.e., segmentation contours) in the target analysis region. Multiple minimum closed regions are extracted from the target analysis region with the segmentation contour lines, and a boundary check is performed on each minimum closed region. The boundary check may refer to checking whether all regions of the minimum closed region are within the target analysis region, and checking whether the minimum closed region intersects with other regions outside the target analysis region (i.e., external regions) or exceeds the chip boundary.

[0102] The minimum enclosed region for which the boundary check results meet the requirements (i.e., the boundary check results indicate that all regions are within the target analysis region and do not intersect with the outside or exceed the chip boundary) is considered a subregion. For minimum enclosed regions whose boundary check results do not meet the above requirements, the outline of the minimum enclosed region can be modified to meet the boundary check requirements. It is understood that the determined subregions can be marked to facilitate subsequent analysis.

[0103] The embodiment of the present application improves the accuracy of determining the sub-region contour line by performing boundary detection on the minimum closed area and determining the sub-region based on the boundary detection result, thereby improving the accuracy of determining the gene expression corresponding to the sub-region.

[0104] It is understood that the sizes of the multiple subregions determined based on the above embodiments can be the same or different. For example, embodiments of the present application can equally divide the target analysis region along the biological gradient based on the number of segments to obtain multiple subregions of equal size along the biological gradient. The following describes a method for segmenting a region where multiple subregions have different sizes along the biological gradient.

[0105] Reference Figure 6 In some embodiments, step S105 may include but is not limited to the following sub-steps:

[0106] Step S601, aligning the image coordinate system corresponding to the regional distribution map with the spatial coordinate system corresponding to the spatial transcriptome data, and establishing a positional correspondence between the pixel points and the capture sites;

[0107] Step S602 , for each image block of the sub-region in the region distribution map, construct a first data set based on the pixel coordinates of each pixel point in the image block and the region label of the sub-region;

[0108] Step S603, based on the position correspondence between the pixel points and the capture sites, combined with the first data set, determining the region label corresponding to the capture site;

[0109] Step S604, constructing a second data set based on the gene expression information, spatial location information, and region labels corresponding to the capture sites;

[0110] Step S605: Based on the second data set, obtain multidimensional data of the target gene at the corresponding capture site; the multidimensional data includes gene expression information, spatial location information, and region labels;

[0111] Step S606 : For each sub-region, the regional expression level of the target gene in the sub-region is counted based on the multi-dimensional data of the target gene.

[0112] In step S601 of some embodiments, the image coordinate system of the regional distribution map can be aligned with the spatial coordinate system corresponding to the spatial transcriptome data to establish a positional correspondence between pixels and capture sites. Thus, for each capture site in the spatial coordinate system, a corresponding pixel (i.e., target pixel) can be found in the regional distribution map. This positional correspondence can be used to describe the positional association between the capture site in the spatial coordinate system and the target pixel in the image coordinate system.

[0113] In some embodiments, the method for establishing the position correspondence between the pixel points and the capture sites in step S601 may include, but is not limited to, the following sub-steps:

[0114] For each pixel point in the area distribution map, scaling the pixel coordinates of the pixel point according to the image resolution of the area distribution map;

[0115] Determine the coordinate offset according to the image coordinate system of the regional distribution map and the spatial coordinate system corresponding to the spatial transcriptome data, and correct the pixel coordinates after scaling according to the coordinate offset;

[0116] The distance between the spatial position information of the capture site and the corrected pixel coordinates is calculated, and the position correspondence between the capture site and the corresponding pixel site is established based on the distance calculation result.

[0117] In an embodiment of the present application, the image coordinate system can be aligned with the spatial coordinate system based on scaling and offset operations. Specifically, the image resolution of the regional distribution map (i.e., the physical size per pixel) can be used as a scaling factor, and the pixel coordinates of each pixel in the regional distribution map can be scaled based on the scaling factor. Then, the offset can be determined according to the image coordinate system and the spatial coordinate system, such as performing a translation correction on the position of the upper left corner of the regional distribution map in the spatial coordinate system of the chip corresponding to the spatial transcriptome, so that the offset can be determined (i.e., the offset is the physical distance of the origin of the image coordinate system relative to the spatial coordinate system). The pixel coordinates after scaling are corrected based on the determined offset. It can be understood that when the regional distribution map is subjected to operations such as rotation compared to the spatial coordinate system, it is also necessary to use a rotation matrix or the like to perform additional transformations on the pixel coordinates. After correcting the pixel coordinates of each pixel, for each capture site in the spatial transcriptome data, the distance (e.g., Euclidean distance or other distance) between the capture site and each pixel can be calculated based on the capture site's spatial location information and the pixel's corrected pixel coordinates. The pixel with the smallest distance is then used as the target pixel for that capture site. That is, for each pixel, the pixel closest to that site is used as the target pixel for that capture site. In this way, a positional correspondence can be established between the spatial coordinates of the capture site in the spatial coordinate system and the pixel coordinates of the corresponding target pixel in the image coordinate system.

[0118] The embodiment of the present application uses a method for determining the target pixel point corresponding to each capture site based on distance, and then establishing a position correspondence relationship, which can ensure that each capture site can find the corresponding target pixel point, that is, a corresponding position correspondence relationship can be established, thereby effectively realizing the assignment of corresponding area labels to the sites based on the target pixel points.

[0119] In step S602 of some embodiments, the image block may refer to the image block corresponding to the sub-region in the region distribution map. For each image block, the pixel coordinates of each pixel in the image block in the region distribution map are determined. In this way, for each pixel in the image block, a first data set can be constructed based on the pixel coordinates of the pixel and the region label of the corresponding sub-region. The first data set may refer to a set of pixel coordinates and region labels of corresponding pixels. For example, the first data set may be in the form of [X1, Y1, region label]. X1 and Y1 represent pixel coordinates, and the region label may be in any form such as numbers, letters, and characters.

[0120] In some embodiments, in step S603, after determining the position correspondence, the region label in the first dataset corresponding to the target pixel can be mapped (or associated) to the capture site based on the position correspondence. Thus, for each capture site, the region label corresponding to the capture site can be determined based on the above method.

[0121] In steps S604 to S605 of some embodiments, for each capture site, the gene expression information of the target gene at the capture site is obtained based on the spatial transcriptome data. In this way, a second data set can be constructed based on the gene expression information, the spatial location information of the capture site, and the regional label corresponding to the capture site. For example, the second data set can be in the form of [X2, Y2, regional label, gene expression information]. That is, the second data set can also include the coordinates (X2, Y2) of the target gene, and the coordinates of the target gene can be determined based on the spatial location information of the capture site. In this way, based on multiple second data sets, multidimensional data of the target gene can be obtained. The multidimensional data includes the regional label, gene expression information, and coordinates corresponding to the target gene.

[0122] In step S606 of some embodiments, for each subregion, multiple gene expression information can be obtained from the multidimensional data of the target gene based on the region label corresponding to the subregion. In this way, the obtained gene expression information can be processed by averaging, etc., to determine the regional expression level of the target gene in the subregion.

[0123] In some embodiments, step S106 may include but is not limited to the following sub-steps:

[0124] Sort the coordinates of the target genes in each subregion according to the biological gradient direction to obtain a coordinate sequence;

[0125] A gradient curve is generated based on the coordinate sequence and the regional expression of the target gene in the sub-region. The gradient curve is used to represent the changing trend of gene expression of the target gene in the direction of the biological gradient.

[0126] In the present embodiment, the coordinates of the target gene can be sorted based on the biological gradient direction. A gradient curve can then be generated based on the sorting result (i.e., the coordinate sequence) and the regional expression level of the target gene in the subregion. The gradient curve can reflect the changes in gene expression values ​​as the coordinate sequence changes, thus visually demonstrating the spatial distribution trend of the target gene's gene expression along the gradient direction.

[0127] The gene expression analysis method provided in the embodiments of this application, based on a region selection operation, enables the extraction of arbitrarily regular or irregular regions of interest, breaking through the limitation of related technologies that can only analyze biologically functional regions. Furthermore, based on gradient direction, it quantifies the spatial gradient and directional changes of gene expression, making this application applicable to the analysis of regions with continuous changes or fuzzy boundaries.

[0128] In a specific embodiment, taking a hippocampal brain sample of Alzheimer's Disease (AD) as an example, the gene expression analysis method provided in this embodiment may include the following steps:

[0129] 1. Cluster analysis. The spatial transcriptome data of AD hippocampal samples (a bin100 matrix, including 100×100 spots) are used as input data. Each spot contains spatial location information and gene expression values. Unsupervised clustering analysis of the spatial transcriptome data is performed based on the Bayesian clustering algorithm to divide the spatial transcriptome chip into multiple spatially specific regions. For example, the hippocampal sample can be divided into the following regions: CA1 region, CA2 region, CA3 region, CA4 region, DG region, and GC region. The clustering results are mapped to the two-dimensional coordinates of the chip, so that the initial brain region division map of the hippocampus can be obtained. Different regions in the initial brain region division map can be represented by different colors.

[0130] 2. Brain region delineation and contour correction. Import the clustering results into Adobe Illustrator software to perform brain region delineation based on the clustering results. In this process, irregular boundary lines can be optimized. For example, the regional division can be corrected with reference to the anatomical structure of the hippocampus to ensure that the boundaries of the CA1 region, CA2 region, CA3 region, and CA4 region conform to the actual anatomical position. In addition, the boundaries of areas with unclear cluster boundaries (such as the junction of the CA2 region and the CA3 region) can be adjusted. After adjustment, vectorization tools can be used to generate accurate brain region contours (such as Figure 2B and export it as a vector data file.

[0131] 3. Contour filling. Use different colors to fill the contour of each brain region to represent each brain region based on different colors. For example, the filling effect is as follows Figure 2C shown.

[0132] 4. Select based on brain region contours. Take the lesion area in the CA1 region of the hippocampus of AD patients as an example. Figure 2C and Figure 3A As shown, the CA1 region and the region 201 are selected to obtain a target analysis region (ie, region of interest) 301 .

[0133] In addition, in the study of AD hippocampal regions, in addition to focusing on the CA1 region, we can also focus on the changes in the lesion area between the DG region and the CA1 region. Figure 2B Therefore, the lesion area can be selected based on the interactive tool. Figure 3BAs shown, the lesion area between the CA1 region and the DG region can be selected as the target analysis area 302 .

[0134] 5. Determine the gradient direction. Figure 4A As shown in FIG, the direction from the CA1 region to the region 201 is used as the gradient direction of the target analysis region 301. Figure 5A As shown, the direction from the CA1 region to the DG region is used as the gradient direction of the target analysis region 302 .

[0135] 6. Determine the number of segments. Figure 4B As shown, expectations are based on Figure 4A The gradient direction shown divides the target analysis area into 19 gradient segments (i.e., sub-areas), so the number of segments is determined to be 19. Among them, from the CA1 area to the area 201 are segments 1 to 19, representing different gradient levels from the lesion area to the area away from the lesion. Figure 5B As shown, expectations are based on Figure 5A The gradient direction shown divides the target analysis area into 6 gradient segments, so the number of segments is determined to be 6.

[0136] 7. Generate a segmentation outline. Generate a segmentation line based on the determined number of segments (i.e., 19) in target analysis region 301 and identify the 19 gradient segments in target analysis region 301. For example, each segment can be represented by a different color to facilitate subsequent extraction of segmented outline data for analysis. Similarly, generate a segmentation line based on the determined number of segments (i.e., 6) in target analysis region 302.

[0137] 8. Contour extraction: Detect closed segmented contours from the gradient segmented image obtained in step 7 and extract the minimum closed area of ​​each segment to determine the integrity of the region.

[0138] 9. Fill the closed area. Use image processing tools (such as OpenCV's cv2.drawContours) to fill each segmented area (i.e., sub-area) and assign each segment a unique color label (e.g., light blue is labeled as gradient 1, dark blue is labeled as gradient 2, etc.). Figure 4B The effect of filling the segmented area shown is as follows Figure 7A As shown, Figure 5B The effect of filling the segmented area shown is as follows Figure 7B shown.

[0139] 10. Export fill data. Export the pixel coordinates and color labels of each pixel in each fill area and store the exported data as a standardized CSV file.

[0140] 11. Initial alignment of the filled data with spatial coordinates. Through scaling and offset operations, the pixel coordinates within the filled area are converted to the spatial coordinates corresponding to the spatial transcriptome chip to ensure that the segmented area is aligned with the actual spatial point of the chip (i.e., the capture site).

[0141] 12. Regional label transfer. The regional label of the filling area is mapped to the corresponding capture site based on the nearest neighbor method. Specifically, the distance between the aligned pixel points and the capture site is calculated, and for each capture site, the pixel point closest to the capture site is used as the target pixel point. In this way, the regional label of the filling area corresponding to the target pixel point can be mapped to the capture site. For example, the spot corresponding to the filling area representing the lesion area is assigned a regional label of gradient 1, and the spot corresponding to the filling area outside the lesion area is assigned a regional label of gradient 2 to gradient 19.

[0142] 13. Export the configuration data. Save the registration data obtained in step 12 as a CSV file.

[0143] 14. Extract regional expression. Extract the expression values ​​of target genes (such as APOE, CLU, and CST3), the coordinates of the target genes, and the regional labels from the registration data.

[0144] 15. Calculate gradient expression. Calculate the gene expression value of the target gene within the segmented region (e.g., for each segmented region, the gene expression value of the target gene within the segmented region can be averaged to obtain the gene expression value of the segmented region) and use linear interpolation to generate a continuous gradient curve.

[0145] 16. Visualization of results. Figure 8A As shown ( Figure 8A The icons DHCR24, NECAB1, PRNP, and RTN1 represent target genes. Draw the gradient curve of the target gene along the gradient direction in the target analysis area 301. Figure 8B As shown in FIG, the gradient curve of the target gene in the target analysis area 302 along the gradient direction is drawn. It can be understood that Figure 8A and Figure 8B In the figure, the Y-axis represents the average expression level of the target gene, and the X-axis represents the spatial progression along the gradient direction. Among them, the AD group represents the visualization results of the sample group, and the Con group represents the visualization results of the control group.

[0146] It is understandable that in addition to the application scenarios mentioned above, according to actual needs, the embodiments of the present application can also be applied to tumor microenvironment research, tissue development research and pathological analysis. Among them, tumor microenvironment research can refer to analyzing the gene expression gradient from the core to the edge of the tumor to reveal the spatial heterogeneity of the tumor. Tissue development research can refer to analyzing the spatial gene expression gradient of a specific tissue to analyze the key mechanisms in the development process. Pathological analysis can refer to the analysis of the gradient changes of specific genes in disease-related areas (such as inflammatory areas or necrotic areas) to provide a basis for subsequent pathological analysis.

[0147] Reference Figure 9 , the embodiment of the present application further provides a gene expression analysis device, the device comprising:

[0148] A site clustering unit 901 is used to cluster multiple capture sites according to the spatial transcriptome data of the target slice to obtain multiple biological function regions;

[0149] An image generating unit 902 is configured to generate a regional distribution map based on the spatial location information of the multiple biological functional regions and each capture site;

[0150] An analysis region determining unit 903 is configured to receive a region selection operation and determine a target analysis region, wherein the region selection operation indicates selecting all or part of the first biological function region in the region distribution map and / or selecting all or part of the second biological function region;

[0151] A region division unit 904 is used to determine the biological gradient direction and divide the target analysis region into multiple sub-regions according to the biological gradient direction;

[0152] The expression quantity acquisition unit 905 is used to obtain the regional expression quantity of the target gene in each sub-region based on the spatial transcriptome data;

[0153] The gene expression analysis unit 906 is used to perform gene expression analysis based on the biological gradient direction and the regional expression amount corresponding to each sub-region; the gene expression analysis is used to describe the change trend of the gene expression of the target gene along the biological gradient direction in the target analysis region.

[0154] It can be seen that the contents of the above-mentioned gene expression analysis method embodiment are all applicable to the embodiment of the present gene expression analysis device. The functions specifically implemented by the present gene expression analysis device embodiment are the same as those in the above-mentioned gene expression analysis method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned gene expression analysis method embodiment.

[0155] Reference Figure 10 , Figure 10The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0156] The processor 1001 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0157] The memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called by the processor 1001 to execute the gene expression analysis method of the embodiments of this application.

[0158] Input / output interface 1003, used to implement information input and output;

[0159] Communication interface 1004, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0160] Bus 1005 , which transmits information between various components of the device (e.g., processor 1001 , memory 1002 , input / output interface 1003 , and communication interface 1004 );

[0161] The processor 1001 , the memory 1002 , the input / output interface 1003 and the communication interface 1004 are connected to each other in communication within the device via the bus 1005 .

[0162] The present application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device executes the above-mentioned gene expression analysis method.

[0163] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned gene expression analysis method is implemented.

[0164] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0165] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0166] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0167] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0168] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0169] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0170] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0171] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0172] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0173] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0174] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0175] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A gene expression analysis method, characterized in that: The method comprises: Clustering multiple capture sites based on the spatial transcriptome data of the target slices to obtain multiple biological functional regions; generating a regional distribution map according to the spatial position information of the plurality of biological functional regions and each of the capture sites; receiving a region selection operation to determine a target analysis region, wherein the region selection operation indicates selecting all or part of the first biological function region and / or all or part of the second biological function region in the region distribution map; wherein the region selection operation refers to an operation of selecting a region on the region distribution map based on an interactive tool; Determining a biological gradient direction, and dividing the target analysis region into a plurality of sub-regions according to the biological gradient direction, including: segmenting the target analysis region according to the biological gradient direction and a preset number of segments to obtain a plurality of segmented contours, extracting a plurality of minimum closed regions according to the plurality of segmented contours, and performing a boundary check on each of the minimum closed regions; when a boundary check result indicates that the corresponding minimum closed region is within the target analysis region, using the minimum closed region as the sub-region; Obtaining the regional expression level of the target gene in each of the sub-regions according to the spatial transcriptome data; Gene expression analysis is performed according to the biological gradient direction and the regional expression amount corresponding to each sub-region, including: sorting the coordinates corresponding to the target gene in each sub-region according to the biological gradient direction to obtain a coordinate sequence, generating a gradient curve according to the coordinate sequence and the regional expression amount of the target gene in the sub-region, the gradient curve is used to represent the changing trend of the gene expression of the target gene in the biological gradient direction, and the gene expression analysis is used to describe the changing trend of the gene expression of the target gene along the biological gradient direction in the target analysis region.

2. The method according to claim 1, characterized in that Each of the capture sites includes gene expression information; and obtaining the regional expression level of the target gene in each of the sub-regions based on the spatial transcriptome data includes: Aligning the image coordinate system corresponding to the regional distribution map with the spatial coordinate system corresponding to the spatial transcriptome data to establish a positional correspondence between pixel points and capture sites; For each image block of the sub-region in the region distribution map, construct a first data set according to pixel coordinates of each pixel point in the image block and the region label of the sub-region; Determining, based on the positional correspondence between the pixel points and the capture sites and in combination with the first data set, a region label corresponding to the capture site; constructing a second data set based on the gene expression information, spatial location information, and the region labels corresponding to the capture sites; Based on the second data set, obtaining multidimensional data of the target gene at the corresponding capture site; the multidimensional data includes gene expression information, spatial location information and the region label; For each of the sub-regions, the regional expression level of the target gene in the sub-region is counted based on the multi-dimensional data of the target gene.

3. The method according to claim 2, characterized in that The method for establishing a position correspondence relationship between pixel points and capture sites includes: For each pixel point in the area distribution map, scaling the pixel coordinates of the pixel point according to the image resolution of the area distribution map; Determining a coordinate offset according to the image coordinate system of the regional distribution map and the spatial coordinate system corresponding to the spatial transcriptome data, and correcting the pixel coordinates after the scaling process according to the coordinate offset; The spatial position information of the capture site and the corrected pixel coordinates are used to perform distance calculation, and a position correspondence relationship between the capture site and the corresponding pixel point is established based on the distance calculation result.

4. The method according to claim 1, wherein The method for determining the direction of the biological gradient includes any of the following: Determining the direction of the biological gradient according to the expression trend of the target gene; Alternatively, a first area of ​​the target tissue in the target analysis area and a second area of ​​the non-target tissue in the target analysis area are determined, and a biological gradient direction is determined based on the first area and the second area, where the biological gradient direction represents a direction from the first area to the second area.

5. A gene expression analysis device, characterized in that: The device comprises: A site clustering unit is used to cluster multiple capture sites based on the spatial transcriptome data of the target slice to obtain multiple biological functional regions; an image generating unit, configured to generate a regional distribution map based on the spatial position information of the plurality of biological functional regions and each of the capture sites; an analysis region determination unit, configured to receive a region selection operation and determine a target analysis region, wherein the region selection operation indicates selecting all or part of the first biological function region and / or all or part of the second biological function region in the region distribution map; wherein the region selection operation refers to an operation of selecting a region on the region distribution map using an interactive tool; a region division unit, configured to determine a biological gradient direction and divide the target analysis region into a plurality of sub-regions according to the biological gradient direction, comprising: segmenting the target analysis region according to the biological gradient direction and a preset number of segments to obtain a plurality of segmented contours, extracting a plurality of minimum closed regions according to the plurality of segmented contours, and performing a boundary check on each of the minimum closed regions; and when a boundary check result indicates that the corresponding minimum closed region is within the target analysis region, using the minimum closed region as the sub-region; an expression quantity acquisition unit, configured to acquire the regional expression quantity of the target gene in each of the sub-regions based on the spatial transcriptome data; A gene expression analysis unit is used to perform gene expression analysis based on the biological gradient direction and the regional expression amount corresponding to each sub-region, including: sorting the coordinates corresponding to the target gene in each sub-region according to the biological gradient direction to obtain a coordinate sequence, generating a gradient curve based on the coordinate sequence and the regional expression amount of the target gene in the sub-region, the gradient curve is used to represent the changing trend of the gene expression of the target gene in the biological gradient direction, and the gene expression analysis is used to describe the changing trend of the gene expression of the target gene along the biological gradient direction in the target analysis region.

6. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the gene expression analysis method according to any one of claims 1 to 4 is implemented.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the gene expression analysis method according to any one of claims 1 to 4 is implemented.

8. A computer program product, comprising a computer program, wherein the computer program is read and executed by a processor of a computer device, so that the computer device executes the gene expression analysis method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Cell type unbiased positioning method and system based on space transcriptome

    CN116364180A

  • Tumor infiltration boundary analysis method and device, electronic equipment and storage medium

    CN119517166A