Organization space structure identification method and device and application

By acquiring tissue spatial transcriptome data and using feature gene sets and clustering algorithms to generate the spatial structure of lymphoid tissue, the problems of low efficiency and low accuracy in lymphoid structure image recognition are solved, enabling self-service recognition and accurate diagnosis.

CN121366413APending Publication Date: 2026-01-20BGI RESEARCH HANGZHOU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410974109.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

In existing technologies, lymphatic structure image recognition is inefficient and inaccurate, especially in the case of immature lymphatic structure image recognition, where there are large errors, making it difficult to achieve batch recognition and resulting in insufficient accuracy.

Method used

By acquiring tissue spatial transcriptome data, annotating it using a feature gene set, and combining the K-nearest neighbor algorithm and community monitoring algorithm, the spatial structure of lymphoid tissue is generated. The OPITICS algorithm is then used for spatial density clustering and pixel processing to generate a closed spatial structure.

Benefits of technology

It enables self-service identification of lymphatic structures without the need for specialized knowledge, improving identification efficiency and accuracy, reducing human error, and providing more precise pathological diagnosis and personalized treatment support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121366413A_ABST
    Figure CN121366413A_ABST
Patent Text Reader

Abstract

The invention relates to the field of biological information, in particular to a tissue space structure recognition method and device and application. The method comprises the following steps: determining a space in-situ image of a given tissue cell population according to space coordinate information of the given tissue cell population; generating a space structure of the given tissue based on the space in-situ image of the cell population of the given tissue; wherein the space coordinate information of the given tissue cell population is determined by the following steps: (a) based on tissue space transcriptome data, acquiring the tissue cell group information; and (b) annotating the tissue cell group information by using a given tissue characteristic gene set so as to determine space coordinate information of the given tissue cell group. Through the method, accurate self-service identification of the organization space structure without depending on professional expert knowledge can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biological information, in particular, the present application relates to a tissue spatial structure identification method, device and application, more particularly, the present application relates to a tissue spatial structure identification method, a tissue spatial structure identification device, a computing device and a computer readable storage medium. BACKGROUND

[0002] The lymphatic system plays a crucial role in the body's immune response, and among them, the lymphatic tissue is particularly critical in the immune response within tumor tissue. The three-level lymph node structure is not only involved in the development process of tumor tissue, but also closely related to tumor metastasis. Because there are differences in morphology, function and pathological characteristics between mature lymphatic structures and immature lymphatic structures, when performing pathological diagnosis and staging on solid tumor malignancies, the state of lymphatic structure is generally judged by artificial observation for subsequent postoperative prognosis prediction, but this traditional artificial evaluation relying on histological changes inevitably has strong subjectivity and poor repeatability.

[0003] And when faced with a large number of lymphatic structure images to be identified, it is difficult to identify a large number of lymphatic structure images based on artificial identification, and there is a problem of low identification efficiency. And based on artificial identification, there is a problem of subjectivity, so that artificial identification has a large error, resulting in low accuracy of subsequent identification.

[0004] Therefore, there is an urgent need for a lymphatic structure image identification method to solve the current problem of low identification efficiency and low accuracy of immature lymphatic structure images. SUMMARY

[0005] The present application aims to at least partially solve one of the above technical problems. To this end, one object of the present application is to provide a means for self-identifying tissue spatial structure.

[0006] In a first aspect of the present application, a tissue spatial structure identification method is provided. According to embodiments of the present application, the method comprises: determining a given tissue cell population spatial coordinate information, determining a given tissue cell population spatial in situ image; generating a spatial structure of the given tissue based on the given tissue cell population spatial in situ image; wherein the given tissue cell population spatial coordinate information is determined by the following steps: (a) obtaining the tissue cell population information based on tissue spatial transcriptome data; (b) annotating the tissue cell population information using a given tissue characteristic gene set to determine the given tissue cell population spatial coordinate information. In some examples of the present application, the present method can achieve accurate self-identification of tissue spatial structure without relying on professional expert knowledge.

[0007] In a second aspect, the present application provides a method for analyzing spatial structure of a tissue. According to an embodiment of the present application, the method comprises: identifying a spatial structure of a given tissue based on the method of the first aspect of the present application; and analyzing the spatial structure of the given tissue based on a bioinformatics method. In some examples of the present application, by identifying and analyzing the spatial structure of a given tissue, doctors can make more accurate and detailed pathological diagnoses. In addition, by analyzing the spatial structures of tissues of different patients, personalized treatment for different patients can be achieved.

[0008] In a third aspect, the present application provides a device for identifying spatial structure of a tissue. According to an embodiment of the present application, the device comprises: a tissue cell cluster information acquisition unit configured to acquire tissue cell cluster information based on tissue spatial transcriptome data; a given tissue cell group spatial coordinate information acquisition unit connected to the tissue cell cluster information acquisition unit, configured to annotate the tissue cell cluster information using a given tissue characteristic gene set to determine given tissue cell group spatial coordinate information; a given tissue cell group spatial in situ image acquisition unit connected to the given tissue cell group spatial coordinate information acquisition unit, configured to determine the given tissue cell group spatial in situ image for the given tissue cell group spatial coordinate information; and a spatial structure generation unit of the given tissue connected to the given tissue cell group spatial in situ image acquisition unit, configured to generate the spatial structure of the given tissue based on the given tissue cell group spatial in situ image. In some examples of the present application, the device can realize self-identification of the spatial structure of the tissue without human intervention, thereby improving efficiency and reducing human error.

[0009] In a fourth aspect, the present application provides a system for analyzing spatial structure of a tissue. According to an embodiment of the present application, the system comprises: the device for identifying spatial structure of a tissue according to the third aspect of the present application, configured to generate the spatial structure of the given tissue; and an analysis device connected to the device for identifying spatial structure of a tissue, configured to analyze the spatial structure of the given tissue based on a bioinformatics method. In some examples of the present application, the system can realize automatic analysis of the spatial structure of the given tissue. The system reduces the complexity of analysis, improves the efficiency of analysis, and reduces human error. In some medical application scenarios, the system makes it easier for users to understand the spatial structure of the given tissue, and provides doctors with more accurate diagnosis and treatment support.

[0010] In a fifth aspect, the present application provides a computer program product. According to an embodiment of the present application, the computer program product comprises computer instructions which, when executed on a computer, cause the method according to the first aspect or the second aspect of the present application to be performed. In some examples of the present application, the organization space structure identification and analysis are implemented by the computer program product, and the identification and analysis of the organization space structure are automatically performed, which greatly improves the efficiency and reduces the errors caused by human factors. Moreover, the results have high repeatability and verifiability.

[0011] In a fifth aspect, the present application provides a computing device. According to an embodiment of the present application, the computing device comprises a memory and a processor; the memory is configured to store a computer program; and the processor is configured to execute the computer program to implement the method according to the first aspect or the second aspect of the present application. In some examples of the present application, the processor runs the program corresponding to the executable program code stored in the memory by reading the executable program code, to implement the method of organization space structure identification or analysis.

[0012] In a sixth aspect, the present application provides a computer readable storage medium. According to an embodiment of the present application, the storage medium comprises computer instructions which, when executed on a computer, cause the computer to implement the method according to the first aspect or the second aspect of the present application. In some examples of the present application, the organization space structure identification and analysis are implemented by the computer readable storage medium, and the identification and analysis of the organization space structure are automatically performed, which greatly improves the efficiency and reduces the errors caused by human factors. Moreover, the results have high repeatability and verifiability.

[0013] It should be noted that the features and technical effects described in this paper for different aspects can be mutually borrowed, which will not be repeated here.

[0014] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter in the description of embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0015] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the appended drawings, wherein:

[0016] Figure 1 is a structural schematic diagram of an organization space structure identification device according to an embodiment of the present application;

[0017] Figure 2 is a structural schematic diagram of an organization space structure analysis system according to an embodiment of the present application;

[0018] Figure 3 is a spatial in situ illustration of immune cell groups according to an embodiment of the present application; wherein the light gray background is a nucleic acid staining background image, and the bright white dots are the expression in situ image of immune cell groups;

[0019] Figure 4 is a spatial structure fine-tuning result illustration of lymphoid tissue according to an embodiment of the present application;

[0020] Figure 5 is a spatial structure recognition result illustration of mouse embryo data according to an embodiment of the present application; wherein the upper left corner is the HE staining image of the sample, and the dark region where the eye is located is marked out with a black dashed line frame; the light gray dots in the main image are all cells with brain expression characteristics, including mature nerve cells of central nervous system and peripheral nervous system, and two main cell aggregation regions are marked out with digital labels 1 and 2 respectively, which are brain and eye. DETAILED DESCRIPTION

[0021] The embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar notations represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by reference to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as limiting the present application.

[0022] In this document, unless otherwise specified, the singular forms "a", "an", and "the" include plural referents (one or more than one); "a group" or "groups" refer to two or more than two.

[0023] In this document, unless otherwise specified, the term "comprising" or "including" is an open expression, that is, it includes the content indicated by the present application, but does not exclude other aspects.

[0024] In this document, unless otherwise specified, the terms "first", "second", "third", "fourth" and the like are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features; the features limited by "first", "second", etc. can explicitly or implicitly include one or more of the features.

[0025] In this document, the term "tissue spatial structure" refers to the spatial arrangement, distribution, and interrelationship of individual cells, cell clusters, or biomolecules in a tissue within an organism, unless otherwise specified. It involves the three-dimensional structure at the tissue level, including the location of cells, relative distance, connections between cell clusters, and the spatial distribution of biomolecules in the tissue. In some examples of the present application, the inventors develop a general method to identify the spatial structure of various types of tissues (such as mature or immature lymphoid tissues), which is of great significance for studying the interaction between tumor tissues and immune cells, understanding biological and physiological processes, and the occurrence and development of tumor diseases.

[0026] In this document, the term "spatial in situ image" refers to an image obtained by visualizing the spatial coordinate information of a cell cluster, unless otherwise specified. In some examples of the present application, the spatial in situ image provides information on the distribution, clustering, and tissue structure of different cells in the tissue.

[0027] In this document, the term "given tissue" is synonymous with the target tissue, indicating a part of the complete tissue, or the entire complete tissue, unless otherwise specified.

[0028] In this document, the term "spatial transcriptome" refers to a biological technique that combines transcriptomics and histology, providing a comprehensive understanding of the three-dimensional spatial information of gene expression by locating the spatial distribution of gene expression in tissues. In some examples of the present application, the inventors achieve self-identification of the spatial structure of tissues (such as lymphoid tissues) through spatial transcriptome data.

[0029] In this document, the term "feature gene set" refers to a set of genes that can distinguish, identify, or characterize a specific biological state, cell type, or disease state. In some examples of the present application, the inventors identify cell clusters of a given cell type based on the feature gene set.

[0030] In this document, the term "spatial density clustering" refers to a method of clustering based on the density distribution of data points in space, which divides data points into different clusters by considering the density around the data points, thereby revealing the clustering structure in the data.

[0031] In this document, the term "pixel erosion" refers to a morphological operation used to reduce or eliminate specific regions in an image. It works by sliding a structuring element (also known as a kernel) over the image, and if a part of the structuring element matches a pixel in the image, the pixel is retained at the corresponding position in the output image. Pixel erosion helps to eliminate small structures, fine objects, or separate closely connected targets.

[0032] In this document, unless otherwise indicated, the term "pixel dilation" is a morphological operation used to enlarge or enhance certain regions in an image. In contrast to pixel erosion, if a part of the structuring element matches a pixel in the image, a pixel is placed at the corresponding position in the output image. Pixel dilation helps to connect regions in an image, fill in holes, or increase the size of objects.

[0033] In this document, unless otherwise indicated, the term "K-Nearest Neighbors (KNN)" is an instance-based, non-parametric supervised learning algorithm for classification and regression. The algorithm finds the K nearest neighbors of a sample to be predicted by calculating the distance between the sample and samples in the training set, and makes a prediction based on the classes or values of these neighbors. In classification, the class of the sample to be predicted is determined by majority voting; in regression, the average or weighted average of the neighbor samples is used as the predicted value. In some examples of the present application, KNN is only used to construct a proximity network to connect each data point.

[0034] In this document, unless otherwise indicated, the term "Louvain algorithm" is a hierarchical clustering algorithm for network community detection that reveals community structure in networks by maximizing modularity. The algorithm first initializes each node as an independent community, and then iteratively performs two stages of node membership adjustment and community aggregation until the modularity no longer significantly improves. Its efficiency and ability to automatically determine the number of communities make it suitable for large-scale network analysis. In some examples of the present application, the Louvain algorithm is used to classify cells in spatial transcriptome data.

[0035] Currently, in the field of tumor research, the identification of lymphoid tissue spatial structure in tumor tissue requires high professional knowledge and relevant experience of professionals, and due to manpower limitations, there are great challenges in batch accurate identification. Although some image recognition technologies can automatically identify lymphoid tissues with clear structures, such technologies are currently only suitable for mature and complete lymphoid tissues, and the identification accuracy for lymphoid tissues with unclear features or immature development is low.

[0036] In order to reduce the skill requirements of professionals and achieve batch accurate identification of various types (mature or immature) of lymphoid tissues, the inventors designed a method capable of self-identifying the spatial structure of tissues. The core idea is that, first, the spatial transcriptome data of the tissue to be tested is obtained and classified; second, the immune cell environment constituting the target lymphoid tissue is screened using the spatial characteristics of the spatial transcriptome data; further, the potential area of the target lymphoid tissue is accurately found according to the spatial density of the immune cells, and the three-level structure of the target lymphoid tissue is self-generated through supervised clustering learning. This method can identify lymphoid tissues of any developmental state without relying on pathological knowledge and related experience; at the same time, this method can also be combined with other types of data for spatial tissue structure identification optimization and analysis. In the study of tumor tissues, this method can effectively improve the identification efficiency and accuracy, and provide more reliable basis for tumor treatment and prevention.

[0037] For ease of understanding, the method and related devices of the present application are described in detail below.

[0038] Method

[0039] In one aspect of the present application, the present application provides a method for identifying the spatial structure of a tissue. The method comprises:

[0040] (a) obtaining tissue cell cluster information based on tissue spatial transcriptome data.

[0041] First, the tissue spatial transcriptome data is filtered. In some examples of the present application, the filtering process refers to excluding data points that meet the following conditions: unique molecular identifiers (UMI) ≤ 100 and the number of gene species (features) ≤ 50. Through this step, the quality and reliability of the data are improved, and potential noise and bias are reduced.

[0042] Then, the filtered spatial transcriptome data is processed by dimension reduction. In some examples of the present application, the dimension reduction processing includes: standardizing the data points after filtering; based on the standardization result, performing principal component analysis on the standardized data points. By reducing the dimension of the spatial transcriptome data, the complexity of the calculation is reduced, and the processing efficiency of the data is improved.

[0043] In one specific example of the present application, the dimension reduction processing is obtained by dividing the filtered spatial transcriptome data by 10000, taking the natural logarithm, subtracting the mean and dividing by the variance to obtain the expression matrix after standardization. The obtained expression matrix is subjected to principal component analysis to obtain the reduced dimension representation space data of the first 50 dimensions.

[0044] Further, based on the K-Nearest Neighbor algorithm and / or the community monitoring algorithm, obtain the tissue cell cluster information.

[0045] (b) annotate the tissue cell cluster information with the given tissue characteristic gene set to determine the given tissue cell cluster spatial coordinate information.

[0046] In some examples of the present application, the given tissue is selected from immune tissue or brain tissue. In some specific examples of the present application, the given tissue is selected from immune tissue. In some preferred examples of the present application, the given tissue is selected from lymphoid tissue.

[0047] In some examples of the present application, the characteristic gene set is selected from at least one of the Cluster of Differentiation (CD) genes and the chemokine receptor genes (including CXC chemokine subfamily and CC chemokine subfamily) of immune tissue.

[0048] In one specific example of the present application, based on the characteristic gene set (the Cluster of Differentiation genes and the chemokine receptor genes of immune tissue) capable of characterizing lymphoid tissue, the cell cluster information obtained in step (a) is annotated to obtain lymphoid tissue cell cluster information. In combination with the spatial coordinate information in the spatial transcriptome data, the lymphoid tissue cell cluster spatial coordinate information is determined.

[0049] (c) determine the given tissue cell cluster spatial in-situ image based on the given tissue cell cluster spatial coordinate information.

[0050] In some examples of the present application, based on the given tissue cell cluster spatial coordinate information obtained in step (b), the pattern recognition of spatial distribution of cells is performed to generate the given tissue cell cluster spatial in-situ image.

[0051] (d) generate the spatial structure of the given tissue based on the given tissue cell cluster spatial in-situ image.

[0052] Based on the given tissue cell cluster spatial in-situ image obtained in step (c), the OPITICS algorithm is used to perform spatial density clustering on the given tissue cell cluster based on a predetermined parameter. In some examples of the present application, the predetermined parameter is selected from at least one of the minimum neighborhood sample number min_samples of the core point and the maximum distance max_eps of two points as the neighborhood relationship. In some specific examples of the present application, the characteristic value of the predetermined parameter is selected from 6-13. Optionally, 6, 7, 8, 9, 10, 11, 12 or 13. In some preferred examples of the present application, the characteristic value of the predetermined parameter is selected from 8-12, more preferably 10.

[0053] OPITICS algorithm is an extended algorithm for spatial density clustering, which generates a sorted list by calculating the core distance and reachable distance of each data point. OPITICS reveals the density changes and clustering structure in the data. Compared with the traditional DBSCAN algorithm, OPITICS does not need to specify the number of clusters in advance, and provides more ordered and detailed clustering structure information through the sorted list, making it widely used in spatial data analysis and clustering problems. Therefore, in some examples of the present application, the inventors perform spatial density clustering processing on the given tissue cell population based on the OPITICS algorithm. Through spatial density clustering processing, the recognition ability of complex data structure is improved, the dependence on algorithm parameters is reduced, and better performance is achieved in different density, shape and size of clusters.

[0054] In addition, for different types of tissues, the spatial density clustering results should meet the following conditions:

[0055] In the state of incomplete, immature or unclear features of the tissue, the spatial density clustering processing result should be consistent with the most dense region of nucleic acid staining;

[0056] In the state of complete and mature tissue, the spatial density clustering processing result should be consistent with the standard given tissue cell population structure.

[0057] By mapping the spatial clustering results that meet the above conditions back to the tissue coordinate system, the spatial structure of the given tissue is generated.

[0058] In some examples of the present application, after obtaining the spatial structure image of the given tissue, further including, closed processing of the given tissue cell population is performed to generate a closed spatial structure of the given tissue. In some examples of the present application, the closed processing includes: based on the OpenCV algorithm, pixel dilation and pixel erosion are performed on the given tissue cell population to perform morphological processing on the spatial structure image of the given tissue. Through pixel dilation, the regions in the image are connected, the pores are filled, or the size of the target is increased, so that the specific structure in the image is more prominent; through pixel erosion, the noise in the image, the small object, or the closely connected target is removed. In addition, it can also be used to remove noise in the image, smooth the image, emphasize or extract the outline of the given tissue, so as to facilitate subsequent image analysis and feature extraction.

[0059] In some examples of this application, the morphological effects of a given tissue spatial structure image are adjusted by modifying the kernel parameters used for pixel erosion and pixel dilation. Specifically, the appearance of the target structure surface can be adjusted by changing the kernel size. Smaller kernel parameters produce a finer target structure surface, while larger kernel parameters result in a smoother effect. Considering the capture rate of spatial group data, the inventors ultimately chose a kernel parameter value between 2 and 6, optionally 2, 3, 4, 5, or 6. In a preferred example of this application, a kernel parameter value of 3 is used to achieve a balance between accuracy and smoothness.

[0060] In some examples of this application, after obtaining the closed spatial structure, the method further includes: performing a correction process on the spatial structure of a given tissue based on visual information. In some examples of this application, the visual information includes at least one selected from hematoxylin and eosin tissue staining maps, fluorescent nucleic acid staining maps, and spatial co-expression feature signals. The correction process includes at least one selected from lasso processing and erasure processing. By combining this with visual information, the accuracy of the spatial structure recognition results is increased.

[0061] In other examples of this application, the spatial structure of the organization can also be finely adjusted using other visual information data, which is not specifically limited in this application.

[0062] In other examples of this application, the spatial structure identification process can also be combined with other bioinformatics analysis methods to delve deeper into gene expression patterns at specific locations within tissues through data integration. Furthermore, more in-depth spatial association analysis can be performed, contributing to a better understanding of the relationships between tumor cells and lymphoid tissues.

[0063] In some examples of this application, the organizational spatial structure identification method of this application is particularly applicable to organizational structures that cannot be distinguished by conventional annotations and to minute organizational structures whose expressive features are not specific.

[0064] In another aspect of this application, a method for analyzing tissue spatial structure is proposed. This method includes: identifying the spatial structure of a given tissue based on the aforementioned method; and analyzing the spatial structure of the given tissue based on bioinformatics methods. This method combines the three-dimensional spatial distribution information of cells within the tissue with bioinformatics methods, which helps to understand biological processes such as intercellular interactions and gene expression patterns in terms of spatial structure, providing an innovative approach for personalized medicine and the discovery of biological laws.

[0065] Device

[0066] In another aspect of this application, an organizational spatial structure identification device is proposed. For example... Figure 1As shown, the device includes: a tissue cell group information acquisition unit 101, a given tissue cell group spatial coordinate information acquisition unit 102, a given tissue cell group spatial in-situ image acquisition unit 103, and a given tissue spatial structure generation unit 104. Among them,

[0067] Unit 101: Used to obtain tissue cell group information based on tissue spatial transcriptome data;

[0068] Unit 102: Unit 102 is connected to Unit 101 and is used to annotate data from Unit 101 using a given tissue feature gene set, and combined with the spatial location information of the spatial transcriptome data to determine the spatial coordinate information of a given tissue cell population.

[0069] Unit 103: Unit 103 is connected to Unit 102 and is used to determine the in situ spatial image of a given tissue cell population based on data from Unit 102;

[0070] Unit 104, connected to Unit 103, is used to generate the spatial structure of a given organization based on data from Unit 103.

[0071] The aforementioned device can achieve automatic identification of the spatial structure of an organization without human intervention, thus improving efficiency and reducing human error.

[0072] system

[0073] In another aspect, this application proposes an organizational spatial structure analysis system. (Reference) Figure 2 The system includes: a tissue spatial structure analysis device 100 (including a tissue cell group information acquisition unit 101, a given tissue cell group spatial coordinate information acquisition unit 102, a given tissue cell group spatial in-situ image acquisition unit 103, and a given tissue spatial structure generation unit 104) and an analysis device 200.

[0074] 100 device: used to generate the spatial structure of a given organization;

[0075] Device 200, connected to device 100, is used to analyze the spatial structure of a given tissue based on bioinformatics methods.

[0076] The aforementioned system enables automated analysis of the spatial structure of a given tissue. This reduces the complexity of the analysis, improves its efficiency, and minimizes human error. In some medical applications, it makes it easier for users to understand the spatial structure of a given tissue, providing doctors with more accurate diagnostic and treatment support.

[0077] Computer program products, computing devices, computer-readable storage media

[0078] In another aspect, this application proposes a computer program product. This computer program product includes computer instructions that, when some or all of which are executed on a computer, cause the organizational spatial structure identification method or organizational spatial structure analysis method of this application to be performed. In some examples of this application, organizational spatial structure identification and analysis are achieved through a computer program product, automating the task of identifying and analyzing organizational spatial structures, greatly improving efficiency and reducing errors caused by human factors. Moreover, the results have high repeatability and verifiability.

[0079] In another aspect, this application proposes a computing device. The computing device includes: a memory and a processor; the memory for storing computer programs; and the processor for executing the computer programs to implement the organizational spatial structure identification method or organizational spatial structure analysis method of this application. In some examples of this application, the processor reads executable program code stored in the memory to run a program corresponding to the executable program code, thereby implementing the method for organizational spatial structure identification or analysis.

[0080] In another aspect, this application proposes a computer-readable storage medium. The storage medium includes computer instructions that, when executed by a computer, cause the computer to implement the organizational spatial structure identification method or organizational spatial structure analysis method of this application.

[0081] The logic and / or steps illustrated in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory. The various computer-readable storage media described in this application can represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" can include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data. It should be noted that the content included in the computer-readable storage medium can be appropriately added to or removed according to the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0082] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0083] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware, and the program can be stored in a computer-readable storage medium. When executed, the program includes one or a combination of the steps of the method embodiments.

[0084] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0085] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0086] Beneficial effects

[0087] 1) The method in this application is based on spatial omics data, which provides a seamless target spatial structure identification scheme before routine analysis for omics analysis, providing a new perspective for a deeper understanding of the development and metastasis process of tumor tissue, and providing a more effective strategy for tumor treatment and prevention;

[0088] 2) The method described in this application does not rely on expert knowledge and performs self-service identification of lymphoid tissue from the perspective of data characteristics. This avoids the time-consuming and labor-intensive nature of traditional identification methods and the problem of subjectivity.

[0089] 3) The tissue spatial structure obtained by the method of this application is a closed region, which reduces the noise generated by other image capture techniques;

[0090] 4) The method of this application can also be used in conjunction with other types of data, which facilitates the optimization and analysis of the generated closed spatial structure.

[0091] Embodiments of the present invention will now be described in more detail, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the invention, and should not be construed as limiting the invention.

[0092] Example 1: Identifying the spatial structure of lymphoid tissue based on complete samples

[0093] This embodiment uses a liver cancer sample as an example to generate a spatial structure of lymphoid tissue using the technical solution of this application.

[0094] Data source for the space group: liver cancer tissue;

[0095] A. By clustering and annotating spatial transcriptome data, the target immune cell populations—T cells and B cells—are obtained. Then, in situ spatial images of these immune cell populations are generated based on their spatial location information. Figure 3 ).

[0096] B. Perform density-based clustering on the spatial in situ images of immune cell populations, referencing... Figure 3 The range of the two mature lymphoid tissues in the upper left corner is set with a minimum neighborhood sample size (min_samples) of 8 and a maximum neighborhood distance (max_eps) of 13.

[0097] C. The clustering results are subjected to dilation and erosion operations to obtain closed lymphoid tissue. In B, the voids in the clustering results are filled, and the boundaries are smoother.

[0098] D. Filter the results of C by area (in this example, only lymphoid tissue with an area greater than or equal to 20 is retained) to obtain more reliable lymphoid tissue;

[0099] E. Perform interactive fine-tuning of the reference nucleic acid staining results for lymphoid tissue in D, such as... Figure 4 . Figure 4 The middle section shows a lymphoid tissue from result D, which differs significantly from the dark fluorescently stained area (imaging evidence); Figure 4 The upper right corner shows the adjusted in-situ spatial map of lymphoid tissue, which shows that all dark-stained areas have been covered.

[0100] F. The spatial labels of the output lymphoid tissues are saved as a list, as shown in Table 1. Table 1 records the label (TLS_label, each number represents a relatively independent lymphoid tissue) of each lymphoid tissue and the coordinates of each data point, so as to facilitate integration into the spatial transcriptome data for downstream data analysis.

[0101] Table 1 shows a partial list of spatial labels for lymphoid tissue.

[0102]

[0103]

[0104] Note: col represents the vertical axis of the spatial transcriptome; row represents the horizontal axis of the spatial transcriptome; TLS_label represents the numbering label of the lymphoid tissue.

[0105] Example 2: Identifying the spatial structure of lymphoid tissue based on incomplete samples

[0106] This embodiment analyzes a section of a mouse embryo at stage E11.5. This section is near the embryonic periphery, and the brain tissue is incomplete, so the eye and brain tissue are not well distinguished in conventional annotation. This embodiment uses mouse embryonic spatial transcriptome data combined with fluorescent staining results of biological samples to locate eye tissue mixed within the brain tissue:

[0107] A. Use the hematoxylin and eosin (HE) histological map of tissue sections as a reference background for structural identification, such as... Figure 5 ;

[0108] B. By clustering and annotating spatial transcriptome data, target cell types related to brain tissue are obtained, and the in-situ spatial information is input into this interactive system, such as... Figure 5 ;

[0109] C. Perform density-based clustering on the target cells in B, selecting a minimum neighborhood sample size (min_samples) of 9 and a maximum neighborhood distance (max_eps) of 9, to obtain... Figure 5 Region structure 1 and region structure 2 in the text;

[0110] D. Based on the background staining image, structure 1 and structure 2 can be distinguished as the brain and eye, respectively; the method of this application can further delineate the fine range of the eye based on the characteristics of the background staining image, such as... Figure 5 ;

[0111] E. Output the coordinates of structure 2 in D to integrate into the spatial transcriptome data, further improving the annotation accuracy of the spatial transcriptome data.

[0112] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0113] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.

Claims

1. A method of identifying a tissue spatial structure, characterized by, The method comprises: determining a spatial in-situ image of a given tissue cell group based on spatial coordinate information of the given tissue cell group; generating a spatial structure of the given tissue based on the spatial in-situ image of the given tissue cell group; wherein, the spatial coordinate information of the given tissue cell group is determined by the following steps: (a) obtaining the tissue cell group information based on tissue spatial transcriptome data; (b) annotating the tissue cell group information using a given tissue characteristic gene set to determine the spatial coordinate information of the given tissue cell group.

2. The method of claim 1, wherein, The given tissue is selected from immune tissue or brain tissue; Optionally, the given tissue is immune tissue, preferably lymphoid tissue.

3. The method of claim 2, wherein, After determining the spatial in-situ image of the given tissue cell group, before generating the spatial structure of the given tissue, the method comprises: performing spatial density clustering processing on the given tissue cell group based on predetermined parameters; performing closed processing on the given tissue cell group based on the spatial density clustering processing result to generate the spatial structure of the given tissue.

4. The method of claim 3, wherein, The predetermined parameters are selected from at least one of the minimum neighborhood sample number min_samples of the core point and the maximum distance max_eps of the two-point neighborhood relationship.

5. The method of claim 3, wherein, The spatial density clustering processing on the given tissue cell group is performed based on the OPITICS algorithm. Optionally, the spatial density clustering processing result should satisfy at least one of the following conditions: 1) the spatial density clustering processing result is consistent with the standard given tissue cell group structure; 2) the spatial density clustering processing result is consistent with the most dense region of nucleic acid staining.

6. The method of claim 3, wherein, The closed processing comprises: performing pixel dilation and pixel erosion on the given tissue cell group based on the OpenCV algorithm.

7. The method of claim 2, wherein, Step (a) comprises: (a-1) filtering the tissue spatial transcriptome data; (a-2) performing dimensionality reduction processing on the filtered spatial transcriptome data; (a-3) obtaining the tissue cell group information based on the K-nearest neighbor algorithm and / or community monitoring algorithm.

8. An apparatus for identifying a tissue spatial structure, characterized by The method comprises: a tissue cell group information obtaining unit for obtaining the tissue cell group information based on tissue spatial transcriptome data; a given tissue cell group spatial coordinate information obtaining unit connected to the tissue cell group information obtaining unit for annotating the tissue cell group information using a given tissue characteristic gene set to determine the spatial coordinate information of the given tissue cell group; a given tissue cell group spatial in-situ image obtaining unit connected to the given tissue cell group spatial coordinate information obtaining unit for determining the spatial in-situ image of the given tissue cell group based on the spatial coordinate information of the given tissue cell group; and a spatial structure generating unit of the given tissue connected to the given tissue cell group spatial in-situ image obtaining unit for generating a spatial structure of the given tissue based on the spatial in-situ image of the given tissue cell group. The method comprises:

9. A computing device, comprising: a memory and a processor; the memory is used to store a computer program; ​ The processor is configured to execute the computer program to implement the method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium comprises computer instructions which, when executed by a computer, cause the computer to implement the method of any one of claims 1-7.