Function unit prediction model construction method, prediction method, device and electronic equipment

By obtaining spatial transcriptome data and staining data of sample slices, and combining with machine learning module to build a functional unit prediction model, the problem of insufficient spatial resolution in the existing technology is solved, accurate prediction of functional units is achieved, and prediction accuracy is improved.

CN120431990AActive Publication Date: 2025-08-05SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510927257.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-08-05
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the spatial structure and its relationship between cells in functional units with spatial structures. Especially in biological tissues or complex systems, anatomical methods and biomedical imaging technologies lack spatial resolution and cannot fully reveal the cellular composition and spatial structure of functional units.

Method used

By obtaining spatial transcriptome data and staining data of sample slices containing functional units, combining machine learning modules, building functional unit prediction models, using spatial transcriptome data for cell classification and distance analysis, and training machine learning modules to predict the spatial structure and cell types of functional units.

Benefits of technology

It improves the accuracy of prediction of functional units with spatial structure, can clearly display the fine structure of functional units at the cellular level, and helps researchers to deeply analyze the disease mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431990A_ABST
    Figure CN120431990A_ABST
Patent Text Reader

Abstract

The invention provides a functional unit prediction model construction method and device, a prediction method and device and electronic equipment. The functional unit prediction method comprises the steps that at least one group of data sets of sample slices containing functional units are acquired, a machine learning module is trained based on the data sets of the at least one group of sample slices to construct a functional unit prediction model, and each group of data sets of the sample slices containing the functional units comprises spatial transcriptome data containing a first cell type, performing cell classification on the plurality of cells in the first sample slice based on the spatial transcriptome data of the first sample slice to obtain a first cell type of the plurality of cells in the first sample slice; and determining the second cell type of the functional unit related cells and the first intercellular distance data based on the staining data of the second sample section, namely determining the second cell type of the functional unit related cells in the second sample section and the first intercellular distance data based on the staining data of the second sample section. According to the embodiment of the invention, the prediction accuracy of the function unit with the space structure can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of functional unit prediction technology, and in particular to a functional unit prediction model construction method, a prediction method, a device, and an electronic device. Background Art

[0002] A spatially structured functional unit refers to a cell group, molecular network, or tissue structure with a clear spatial location and that performs a specific function in a biological tissue or complex system. For example, when the functional unit is the neurovascular unit (NVU), since the NVU is a functional unit composed of cells of multiple cell types working together in close collaboration, a complete functional analysis of the NVU is crucial for normal cognitive function and nervous system health. Therefore, accurately predicting functional units can help researchers more deeply analyze the gene expression characteristics of functional units, thereby assisting in understanding the occurrence and development of related diseases.

[0003] Related technologies typically rely on anatomical methods and biomedical imaging techniques to label functional units with known spatial structures and perform partial spatial analysis. However, these methods struggle to accurately predict the spatial structure and intercellular relationships within these functional units. Therefore, developing a method that can be applied to predicting functional units with spatial structures has become an urgent technical challenge. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a method for constructing a functional unit prediction model and a prediction method, device and electronic equipment, aiming to provide a prediction method that can be applied to functional units with spatial structures.

[0005] To achieve the above objectives, a first aspect of an embodiment of the present application proposes a method for constructing a functional unit prediction model, the method comprising: Step S1: obtaining at least one set of data sets of sample slices containing functional units, wherein each set of data sets of sample slices containing functional units includes: Step S1.1, spatial transcriptome data containing a first cell type, the acquisition method comprising: acquiring spatial transcriptome data of a first sample slice, and performing cell classification on a plurality of cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of the plurality of cells in the first sample slice; the spatial transcriptome data comprising expression levels and spatial distribution information of transcriptomes of the plurality of cells in the first sample slice; and Step S1.2, obtaining the second cell type and first inter-cell distance data of the functional unit-related cells, including: obtaining staining data of a second sample slice, and determining the second cell type and first inter-cell distance data of the functional unit-related cells in the second sample slice based on the staining data, wherein the first inter-cell distance data is used to describe the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or an adjacent slice of the first sample slice; Step S2: training a machine learning module based on the data set of the at least one group of sample slices containing functional units to construct a functional unit prediction model.

[0006] In some embodiments, before performing cell classification on multiple cells in the first sample slice based on the spatial transcriptome data in step S1, the functional unit prediction model construction method further includes: performing cell segmentation on the first sample slice, and the basis of the cell segmentation is the spatial transcriptome data of the first sample slice, and / or image data obtained by staining the cell nucleus and / or membrane of the first sample slice, such as an ssDNA staining image.

[0007] In some embodiments, the machine learning module includes at least one of a supervised learning module, a semi-supervised learning module, an unsupervised learning module, a regression analysis module, a reinforcement learning module, a self-learning module, a feature learning module, a sparse dictionary learning module, an anomaly detection module, a generative adversarial network, or an association rule module.

[0008] In some embodiments, the step of training a machine learning module based on a dataset of at least one set of sample slices containing functional units to construct a functional unit prediction model includes: constructing spatial adjacency data of the first sample slice according to the spatial transcriptome data containing the first cell type, wherein the spatial adjacency data is used to describe the adjacency structure between the same or different cells in the first sample slice; extracting cell space data related to the functional unit from the spatial adjacency data according to a matching result between a second cell type and the first cell type of cells related to the functional unit and the first inter-cell distance data; Determining, from the first cell type, a central cell type, an associated cell type, and distance thresholds between central cells, between associated cells, and between the central cell and the associated cells according to the spatial structure indicated by the unit spatial data; the central cell refers to a cell constituting the functional unit, and the associated cell refers to a non-central cell that has intercellular communication with the central cell; A machine learning module is trained according to the central cell type, the associated cell type and the distance threshold to construct a functional unit prediction model.

[0009] In some embodiments, performing cell classification on the plurality of cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of the plurality of cells in the first sample slice includes: Using single-cell transcriptome sequencing technology, obtaining single-cell transcriptome sequencing data of multiple cells from the sample from which the first sample slice originates, and determining common gene expression data for each cell in the first sample slice based on the expression levels of the transcriptomes of the multiple cells in the first sample slice and the single-cell transcriptome sequencing data; Cell annotation is performed on the multiple cells in the first sample slice based on the common gene expression data and preset cell marker genes to obtain a first cell type of the multiple cells in the first sample slice.

[0010] In some embodiments, determining the common gene expression data of each cell in the first sample slice based on the expression levels of the transcriptomes of the plurality of cells in the first sample slice and the single-cell transcriptome sequencing data includes: determining a first gene expression matrix of the plurality of cells in the first sample slice based on the expression levels of the transcriptomes of the plurality of cells in the first sample slice; Determine a second gene expression matrix of multiple cells in the first sample slice based on the single-cell transcriptome sequencing data; Common gene expression data of each cell in the first sample slice is obtained based on the common genes of the multiple cells in the first gene expression matrix and the second gene expression matrix.

[0011] In some embodiments, functional unit prediction is performed on a third sample slice using the functional unit prediction model and compared with staining data of a fourth sample slice to determine the accuracy level of the machine learning module, wherein the fourth sample slice is the third sample slice or an adjacent slice of the third sample slice.

[0012] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application proposes a functional unit prediction method, the method comprising: Acquiring spatial transcriptome data of the slice to be predicted, and performing cell classification on a plurality of cells in the slice to be predicted based on the spatial transcriptome data of the slice to be predicted to obtain a third cell type of the plurality of cells in the slice to be predicted, wherein the spatial transcriptome data of the slice to be predicted includes expression levels and spatial distribution information of transcriptomes of the plurality of cells in the slice to be predicted; The functional unit prediction model constructed according to the functional unit prediction model construction method of the embodiment of the present application performs functional unit prediction based on the spatial transcriptome data containing the third cell type of the slice to be predicted.

[0013] In some embodiments, the functional unit prediction model constructed according to the first aspect of the embodiments of the present application performs functional unit prediction based on the spatial transcriptome data containing the third cell type of the slice to be predicted, including: According to the functional unit prediction model constructed by the functional unit prediction model construction method of the embodiment of the present application, determining the central cell type, associated cell type, and distance thresholds between central cells, between associated cells, and between the central cell and the associated cells related to the functional unit; Determining a central cell of a functional unit from a plurality of cells in the slice to be predicted based on a matching result between the central cell type and the third cell type; Determining an associated cell of a functional unit from a plurality of cells in the slice to be predicted based on a matching result between the associated cell type and the third cell type; Acquire, according to the spatial transcriptome data, second cell distance data between the central cells, between the associated cells, and between the central cell and the associated cells in the slice to be predicted, wherein the second inter-cell distance data is used to describe the relative distances between the central cells, between the associated cells, and between the central cell and the associated cells in the slice to be predicted; Determining, based on a comparison result of the relative distance and the distance threshold, a target cell associated with the functional unit to which the central cell belongs from a plurality of cells in the slice to be predicted; the target cell includes the central cell and the associated cell; The functional units of the slice to be predicted are predicted according to the target cells and the spatial distribution information of the target cells in the slice to be predicted.

[0014] In some embodiments, after the functional unit prediction model constructed according to the functional unit prediction model construction method of the embodiments of the present application performs functional unit prediction based on the spatial transcriptome data containing the third cell type of the slice to be predicted, the functional unit prediction method further comprises: Extracting the expression level and spatial distribution information of the transcriptome of the target cell in the slice to be predicted from the spatial transcriptome data of the slice to be predicted; Density statistics and functional analysis are performed on the functional units in the slice to be predicted based on the expression level and spatial distribution information of the transcriptome of the target cell in the slice to be predicted.

[0015] To achieve the above-mentioned purpose, a third aspect of the embodiments of the present application provides a device for constructing a functional unit prediction model, the device comprising: A data set acquisition module for acquiring a data set of at least one set of sample slices containing functional units, the data set acquisition module comprising: a first cell classification module and a second cell classification module, the first cell classification module being configured to acquire spatial transcriptome data containing a first cell type, the acquisition method comprising: acquiring spatial transcriptome data of a first sample slice, and performing cell classification on a plurality of cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of the plurality of cells in the first sample slice; the spatial transcriptome data comprising expression levels and spatial distribution information of transcriptomes of the plurality of cells in the first sample slice; the second cell classification module being configured to acquire a second cell type of cells associated with the functional unit and first inter-cell distance data, the acquisition method comprising: acquiring staining data of a second sample slice, and determining the second cell type of cells associated with the functional unit in the second sample slice and first inter-cell distance data based on the staining data, the first inter-cell distance data being configured to describe the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or an adjacent slice of the first sample slice; A model building module is used to train a machine learning module based on a data set of at least one group of sample slices containing functional units to build a functional unit prediction model.

[0016] To achieve the above-mentioned purpose, a fourth aspect of the embodiments of the present application provides a functional unit prediction device, the device comprising: a third cell classification module, configured to obtain spatial transcriptome data of the slice to be predicted, and perform cell classification on a plurality of cells in the slice to be predicted based on the spatial transcriptome data of the slice to be predicted, to obtain a third cell type of the plurality of cells in the slice to be predicted, wherein the spatial transcriptome data of the slice to be predicted includes expression levels and spatial distribution information of transcriptomes of the plurality of cells in the slice to be predicted; The unit prediction module is used to predict the functional unit based on the spatial transcriptome data containing the third cell type of the slice to be predicted according to the functional unit prediction model constructed by the functional unit prediction model construction method of the embodiment of the present application.

[0017] To achieve the above-mentioned purpose, the fifth aspect of an embodiment of the present application proposes an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the functional unit prediction model construction method described in the first aspect and the functional unit prediction method described in the second aspect.

[0018] To achieve the above-mentioned purpose, the sixth aspect of an embodiment of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the functional unit prediction model construction method described in the first aspect and the functional unit prediction method described in the second aspect.

[0019] To achieve the above-mentioned purpose, the seventh aspect of an embodiment of the present application proposes a computer program product, which includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the functional unit prediction model construction method described in the first aspect and the functional unit prediction method described in the second aspect.

[0020] The functional unit prediction model construction method, prediction method, device and electronic device proposed in the embodiment of the present application can, when constructing a functional unit prediction model for predicting functional units in a slice, first obtain a data set of at least one group of sample slices containing functional units, and each group of sample slices containing functional units includes: spatial transcriptome data containing a first cell type, a second cell type of cells related to the functional unit, and first inter-cell distance data; wherein, the method for obtaining the spatial transcriptome data containing the first cell type includes: obtaining the spatial transcriptome data of the first sample slice, and performing cell classification on multiple cells in the first sample slice based on the spatial transcriptome data to obtain the first cell type of multiple cells in the first sample slice. Cell type; spatial transcriptome data includes expression levels and spatial distribution information of transcriptomes of multiple cells in a first sample slice; a method for obtaining data on the second cell type and intercellular distance of cells associated with the functional unit includes: obtaining staining data of a second sample slice, determining the second cell type and first intercellular distance data of cells associated with the functional unit in the second sample slice based on the staining data, the first intercellular distance data being used to describe the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or an adjacent slice of the first sample slice; thereafter, training a machine learning module based on a data set of at least one set of sample slices containing functional units to construct a functional unit prediction model. Since the distribution of functional units in the first sample slice and the second sample slice is similar, the embodiment of the present application can analyze the type distribution and spatial distribution between cells contained in the functional unit based on the spatial transcriptome data containing the first cell type, and the second cell type and intercellular distance data of cells associated with the functional unit, so as to construct a functional unit prediction model capable of predicting the functional unit. In this way, the embodiment of the present application can be better applied to the prediction of functional units with spatial structures and improve the accuracy of the prediction of functional units. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are used to provide a further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation to the technical solution of the present disclosure.

[0022] Figure 1 This is a flow chart of the method for constructing a functional unit prediction model provided by an embodiment of the present application; Figure 2 This is a spatial distribution map of cell types in the hippocampus provided in the examples of the present application; Figure 3 This is an immunohistochemical fluorescence image of the hippocampus provided in the examples of the present application; Figure 4 Schematic diagram of constructing a functional unit prediction model based on machine learning provided in an embodiment of the present application; Figure 5 This is a schematic diagram of the NVU model prediction convergence curve provided in an embodiment of the present application; Figure 6 It is a spatial distribution diagram of the functional units of the neurovascular unit provided in the embodiment of the present application; Figure 7 This is a schematic diagram of density changes of the functional units of the neurovascular unit provided in an embodiment of the present application; Figure 8 This is a schematic diagram of the functional unit of the neurovascular unit in response to the hypoxia pathway provided in the embodiments of the present application; Figure 9 This is a schematic diagram of the functional unit of the neurovascular unit in regulating immune response provided by the embodiments of the present application; Figure 10 This is a flow chart of a specific embodiment of the method for constructing a functional unit prediction model provided in an embodiment of the present application; Figure 11 is a schematic diagram of a functional unit prediction device provided in an embodiment of the present application; Figure 12 is a schematic diagram of a functional unit prediction device provided in an embodiment of the present application; Figure 13 This is a hardware structure diagram of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not intended to limit the present disclosure.

[0024] Before further explaining the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations: Spatial transcriptomics (ST) sequencing technology is a technique for spatially analyzing RNA. It is primarily used to determine the spatial location of RNA within a single tissue section. It uses unique molecular identifiers (UMIs) to identify copies of the same RNA molecule and can also simultaneously measure RNA expression levels. Stereo-seq (Spatial Enhanced Resolution Omics-sequencing) is an advanced spatial transcriptomics technology. Developed using DNA nanoballs (DNBs), Stereo-seq is a high-throughput, ultra-high-resolution, and large-field-of-view in situ panoramic technique. Stereo-seq enables simultaneous spatial transcriptome analysis of a single sample at four scales: tissue, cellular, subcellular, and molecular. Stereo-seq captures mRNA within a tissue using spatiotemporal microarrays and restores its spatial location using spatial barcodes (or Coordinate IDs, CIDs). This enables the spatial expression of genes within the tissue, laying a strong foundation for a deeper understanding of the relationship between cellular gene expression and morphology and the local environment.

[0025] Pathological staining images: Specific stains are used to enhance the contrast of specific structures or components within tissue samples, allowing for clearer observation and analysis of tissue cell morphology and pathological changes under a microscope. Common staining techniques include hematoxylin and eosin staining and special stains (such as mIF staining). Hematoxylin and eosin staining, referring to hematoxylin and eosin, can clearly reveal distinct structures in the cell nucleus and cytoplasm. Special staining can refer to the use of specialized staining techniques for specific purposes. Special stains include Masson's trichrome staining, PAS staining, and silver staining. These staining techniques can highlight specific tissue structures or components, such as collagen fibers, muscle tissue, and neural tissue.

[0026] Artificial Intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and create new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0027] Spatial transcriptomics is an emerging technology that combines imaging techniques with gene expression analysis to precisely localize gene expression within the spatial structure of cells or tissues. When spatial transcriptomics is linked and combined with multimodal information, it can reveal the behavior and function of cells within specific biological structures, providing a foundation for understanding complex biological processes such as disease development, tissue formation, and immune responses. Spatial transcriptomics can be used to observe the distribution of different cell types within tissues, as well as the interactions between different cell types and other functional or characteristic features.

[0028] Spatial resolution refers to the minimum distance between two adjacent objects that an imaging system can distinguish. It is primarily used to describe the image's ability to resolve detail in spatial dimensions (such as length and width for two-dimensional images, or length, width, and height for three-dimensional images). For example, in CT imaging, high spatial resolution means that tiny bone structures and soft tissue boundaries can be more clearly distinguished. If the spatial resolution is low, image details will be blurred, and adjacent structures may merge and become difficult to distinguish. Taking brain tissue as an example, high spatial resolution imaging technology can help researchers clearly see the distribution of neuronal cell bodies, the direction of nerve fibers, and the connections between them, thereby better understanding the brain's neural circuits and functions.

[0029] A spatially structured functional unit refers to a group of cells, molecular networks, or tissue structures with a well-defined spatial location and specific functions within a biological tissue or complex system. For example, the neurovascular unit (NVU) is a functional unit composed of multiple cell types working together in close collaboration. The NVU is responsible for maintaining brain homeostasis, ensuring adequate oxygen and nutrients for neurons, and removing metabolic waste. Therefore, a comprehensive functional analysis of the NVU is crucial for normal cognitive function and neurological health. These units rely not only on their constituent components (such as specific cells, proteins, and genes) but also on their spatial arrangement and interactions to achieve their function. For example, in neural tissue, neurons and glial cells form complex neural networks with specific spatial distribution and connectivity. These networks regulate and control various physiological activities of animals through the transmission of electrical and chemical signals. Different cell arrangements reflect the specific functional requirements of the tissue, and signal transduction and regulation within the tissue structure also depend on the spatial location and interrelationships of cells. For example, in neural tissue, the complex architecture of axons and dendrites allows for long-distance transmission and synchronous processing of neural signals. Functional units with spatial structure have broad applications in studying disease mechanisms and precision medicine. For example, they reveal how the spatial distribution of amyloid plaques in Alzheimer's disease affects neighboring neurons, and locate areas of drug-resistant cell enrichment in tumors, thereby guiding localized targeted therapy. Therefore, accurately predicting functional units can help researchers more deeply analyze the gene expression characteristics of functional units, thereby assisting in understanding the occurrence and development of related diseases.

[0030] However, current studies of the cellular composition and spatial organization of spatially structured functional units typically rely on anatomical methods and biomedical imaging techniques to label known spatially structured functional units and perform partial spatial analysis. However, these methods struggle to accurately predict the spatial organization and intercellular relationships within these units (i.e., they are unable to effectively predict the complex gene expression patterns and spatial distribution of these units). Anatomical methods rely on tissue sectioning, immunohistochemical staining, and microscopic observation to label and visualize the structure of these units. These methods often require extensive expertise and can be subject to subjectivity in tissue section processing and image analysis. Due to their lack of spatial resolution, these methods struggle to reveal the complete structure and cellular relationships of these units at the cellular level. Immunohistochemical staining combined with microscopic imaging is a common method for studying spatially structured functional units. By staining and labeling specific cell types (e.g., the neurovascular unit, which includes endothelial cells, astrocytes, and neurons), researchers can visualize the spatial distribution and interactions of these cells. However, this method requires demanding tissue sample processing and can be limited by section thickness and labeling accuracy when resolving complex three-dimensional spatial structures. In addition, related technologies have proposed using gene expression analysis methods (such as single-cell RNA sequencing) to provide cellular composition and gene expression information of functional units with spatial structure. However, the gene expression technologies used in related technologies generally lack spatial resolution and cannot reveal the precise location and interactions of cells in tissues.

[0031] For example, when studying brain neural tissue, the gene expression of nerve cells will change dynamically under different physiological and pathological conditions, and these changes are closely related to the interaction between cells. The methods used in related technologies can only observe the static state of cells at a certain point in time, and it is difficult to capture the dynamic changes in gene expression over time and the interaction between cells. In other words, it is impossible to fully and deeply reveal the dynamic changes in gene expression between cells within functional units with spatial structures. Therefore, how to provide a prediction method that can be applied to functional units with spatial structures has become a technical problem that needs to be solved urgently.

[0032] Based on this, the embodiments of the present application propose a functional unit prediction model construction method and a prediction method, device and electronic equipment, which can be better applied to the prediction of functional units with spatial structures and improve the prediction accuracy of functional units with spatial structures.

[0033] The following describes the method for constructing a functional unit prediction model provided in an embodiment of the present application.

[0034] Reference Figure 1In some embodiments, the functional unit prediction model construction method provided in the embodiments of the present application includes but is not limited to steps S1 to S2.

[0035] Step S1: obtaining at least one set of data sets of sample slices containing functional units, wherein each set of data sets of sample slices containing functional units includes: Step S1.1, spatial transcriptome data of a first cell type, the acquisition method comprising: acquiring spatial transcriptome data of a first sample slice, and performing cell classification on a plurality of cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of the plurality of cells in the first sample slice; the spatial transcriptome data comprising expression levels and spatial distribution information of transcriptomes of the plurality of cells in the first sample slice; and Step S1.2, obtaining second cell type and intercellular distance data of cells associated with the functional unit, including: obtaining staining data of a second sample slice, and determining the second cell type and first intercellular distance data of cells associated with the functional unit in the second sample slice based on the staining data, wherein the first intercellular distance data is used to describe the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or an adjacent slice of the first sample slice; Step S2: training a machine learning module based on a data set of at least one set of sample slices containing functional units to construct a functional unit prediction model.

[0036] Steps S1 to S2 shown in the embodiment of the present application, by introducing spatial transcriptome data, can reveal the distance and spatial distribution characteristics between cells at the single-cell level, and combined with other technical features, can provide researchers with the spatial coordinates of cells and their interactions, solving the problem of insufficient spatial resolution in related technologies, and can clearly display the fine structure of functional units with spatial structure at the cellular level, thereby accurately analyzing the composition and function of functional units with spatial structure. Specifically, the embodiment of the present application is based on the cell types of multiple cells in the first sample slice, the cell types of multiple cells in the second sample slice, and the intercellular distance data to construct a functional unit prediction model that can predict the cell type distribution and spatial distribution of functional units, which can be better applied to the prediction of functional units with spatial structure and improve the accuracy of prediction of functional units.

[0037] Among them, step S1.1 and step S1.2 can be performed in any order.

[0038] It should be noted that spatial transcriptomics technology can break through these limitations and provide more accurate data at the cellular and tissue levels, helping researchers in fields such as neuroscience, oncology, and immunology to deeply analyze the functional units of cells or tissues with spatial structures and identify their gene expression characteristics, thereby providing new perspectives for the study of disease mechanisms.

[0039] In step S1 of some embodiments, at least one dataset of sample slices containing functional units is used to train a subsequent machine learning module to construct a functional unit prediction model for predicting functions. The process of obtaining each dataset of sample slices containing functional units may specifically include: step S1.1, obtaining spatial transcriptome data containing a first cell type; and step S1.2, obtaining a second cell type and intercellular distance data of cells associated with the functional unit.

[0040] In step S1.1 of some embodiments, when obtaining spatial transcriptome data containing a first cell type, the first sample slice may refer to a tissue slice to be subjected to spatial transcriptome sequencing. The first sample slice is a biological slice obtained in accordance with relevant legal provisions, including plant and animal slices. The first sample slice refers to a specific tissue slice selected during the research process, which contains the target cell population to be studied (e.g., the first sample slice tested may contain 50,000-100,000 cells, without limitation), and serves as the basis for subsequent analysis. The spatial transcriptome data of the first sample slice refers to data obtained using spatial transcriptome technology regarding the expression levels and spatial distribution of transcriptomes of multiple cells in the first sample slice. Spatial transcriptome technology can simultaneously obtain spatial information on gene expression in tissue or cell samples. This data can reflect the specific location of cells in the tissue and their gene expression characteristics. Cell classification refers to the process of classifying multiple cells in the first sample slice into different cell types or cell subpopulations based on factors such as their gene expression characteristics and spatial location. The cell type refers to an identifier assigned to each cell in the first sample slice, which is used to indicate the specific cell type or cell subpopulation to which the cell belongs, such as vascular cells (endothelial cells and pericytes), astrocytes, neurons (inhibitory neurons and excitatory neurons), etc.

[0041] It should be noted that for the specific method of obtaining spatial transcriptome data of slices (such as sample slices, slices to be predicted, etc.) in this application, in order to supplement this disclosure, this application can adopt the patent disclosures on the method of obtaining spatial transcriptome data in the following patent applications by reference for data acquisition: European Patent Publication No. EP2697391A1, International Patent Application Publication No. WO2024086167A3, and WO2023115536A1. The structure and acquisition method of all slices mentioned in the embodiments of this application can refer to the first sample slice mentioned above and will not be repeated here.

[0042] It should be noted that the spatial transcriptome data for the first sample slice includes spatial location data for multiple cells in the first sample slice (i.e., spatial distribution information of the transcriptome in the first sample slice) and spatial gene expression data (i.e., expression levels of the transcriptome in the first sample slice). The spatial transcriptome data in this application is spatial transcriptome data with single-cell resolution, which can accurately predict functional units with spatial structure, such as spatial transcriptome data determined using Stereo-seq methods. Spatial location data refers to the specific location information of multiple cells in the tissue in the first sample slice. In spatial transcriptomics studies, researchers can divide tissue slices into multiple small regions, each of which is called a "spot." Each spot represents a small area of tissue (e.g., a cell in a tissue slice). Cells within this region are analyzed to determine their gene expression patterns. The spatial location of these cells can be described by coordinates, such as (x, y) coordinates in a two-dimensional plane or (x, y, z) coordinates in three-dimensional space. This coordinate information can clearly define the relative or absolute position of each cell in the first sample slice. Spatial gene expression data refers to the specific gene expression patterns within multiple cells in the first sample slice. In practical applications, spatial transcriptome data refers to the expression data of all genes in a cell or tissue under specific conditions and time points. Spatial gene expression data reflects the transcriptional activity of each gene within each cell in the first sample slice, indicating which genes are expressed and at what levels. This data can be presented in the form of a matrix, with rows representing genes and columns representing cells (or spots). Each element in the matrix represents the expression level of a gene in a particular cell.

[0043] In some embodiments, before performing cell classification on multiple cells in the first sample slice based on the spatial transcriptome data in step S1, the functional unit prediction model construction method further includes: performing cell segmentation on the first sample slice, and the basis of the cell segmentation is the spatial transcriptome data of the first sample slice, and / or, image data obtained by staining the cell nucleus and / or membrane of the first sample slice, such as an ssDNA staining image.

[0044] In some embodiments, the process of selecting the first sample slice specifically includes: Acquire spatial transcriptome data of multiple sample slices, where the spatial transcriptome data of each sample slice includes the expression level and spatial distribution information of the transcriptome of each cell in the sample slice; Data extraction was performed based on spatial transcriptome data to obtain the number of genes and the proportion of mitochondrial genes in each cell in each sample section; Multiple sample slices are screened based on the number of genes per cell and the ratio of mitochondrial genes to determine a first sample slice.

[0045] Among them, the present application can first screen multiple initial sample slices, that is, perform quality control on the initial sample slices, eliminate low-quality cells and slices, and use the initial sample slices that meet the quality control standards as the first sample slices to construct a functional unit prediction model for predicting functional units with spatial structures. Specifically, spatial sequencing data and corresponding spatial transcriptome data of multiple spatial sites in each initial sample slice are first obtained. The spatial transcriptome data includes spatial gene expression data of multiple cells in the initial sample slice (that is, the expression level of the transcriptome of multiple cells in the initial sample slice) and spatial position data (that is, the spatial distribution information of the transcriptome of multiple cells in the initial sample slice).

[0046] Spatial sequencing data refers to data on transcriptome expression and spatial location, obtained by performing spatial transcriptome sequencing on each initial sample slice using technologies such as Visium, Visium HD, Visium HD3', and Stereo-seq. For example, Visium uses spatial probes on a chip, each containing a spatial barcode (each position on the chip has a unique spatial barcode sequence), UMIs, and a capture sequence (poly T), to capture mRNA in tissues. The spatial probes are then extended using the mRNA as a template. The resulting extension products contain mRNA sequence information, corresponding UMIs, and spatial location information (spatial barcode sequence information). This information is decoded through sequencing (such as bridge sequencing and DNB sequencing). Each position (or spatial site) is called a spot. Generally, a spot contains a cluster of spatial probes with the same spatial barcode. UMIs are used to distinguish copies of mRNA molecules from different sources, avoiding the problem of double counting caused by polymerase chain reaction (PCR) amplification.

[0047] Furthermore, after determining the spatial transcriptome data of multiple sample slices, a spatial density image of the initial sample slice can be generated based on the spatial sequencing data and the spatial coordinates of each spatial site in the initial sample slice. Among them, the present application can summarize the total UMI count of each spot point in the spatial coordinates of the sample slice based on the spatial sequencing data (specifically, by identifying and counting the different UMI sequences on each spot point to obtain the total UMI count corresponding to the spot point. This process can be completed using specialized bioinformatics software and algorithms, without limitation), and generate a spatial density matrix for each sample slice based on the total UMI count and the spatial coordinates of each spatial site. Afterwards, the spatial density matrix is converted into an image to obtain a spatial density image, and the grayscale intensity in the image is used to reflect the number of UMIs.

[0048] It should be noted that for the spatial density matrix, the present application can associate the spatial coordinates of each spot point with the corresponding total UMI count, that is, to correspond these spatial coordinates to the corresponding total UMI count one by one to form a data set. Furthermore, according to the spatial layout of the sample slice and the distribution of the spot points, the above data set is organized into a matrix. The rows and columns of the matrix correspond to the spatial positions on the initial sample slice (such as the row corresponds to the row coordinate of the sample slice, and the column corresponds to the column coordinate of the sample slice), and each element in the matrix represents the total UMI count corresponding to the spot point at that position. In this way, a spatial density matrix is generated, and the element values in the matrix reflect the gene expression density at the corresponding position.

[0049] Furthermore, the present application can perform cell segmentation on the spatial density image based on the spatial distribution information of the transcriptomes of multiple cells in the initial sample slice in the initial sample slice, and determine the cell area of each cell in the spatial density image in the initial sample slice. Specifically, the present application can preset a cell segmentation model (such as the ESPANet model) for cell segmentation, and post-process the initial spatial gene expression data of the initial sample slice using a watershed algorithm, ultimately generating a cell-gene matrix for creating a Seurat RDS object for a cellbin (cell unit; in the field of spatial transcriptomic technology, a bin can refer to the process of merging adjacent or similar cells in spatial positions into a single unit for analysis).

[0050] The present application can also perform cell segmentation based on the optical image of the initial sample slice, for example, using H&E staining images or mIF images to perform cell segmentation.

[0051] It should be noted that the watershed algorithm is a morphologically based image segmentation algorithm that treats an image as a terrain with peaks and valleys. It segments the image by flooding seed regions and gradually expanding them. In cell segmentation, the watershed algorithm can utilize cell boundary information to further refine the segmentation results of the ESPANet model, address problems such as cell adhesion, and optimize the segmentation results, thereby more accurately determining the cell area of each cell in the spatial density image in the initial sample slice.

[0052] Furthermore, data extraction is performed on the initial spatial gene expression data based on cell regions to obtain the number of genes (nFeatures), the proportion of mitochondrial genes (percent.MT), and the total number of gene expression levels for each cell in the initial sample slice. nFeatures refers to the number of genes per cell, that is, the number of unique genes detected in each cell. The nFeatures value for each cell is obtained by counting the number of nonzero elements in each row of the cell-gene matrix. nCounts represents the sum of the expression levels of all genes in each cell, that is, the total number of UMIs detected in each cell. The nCounts value for each cell is obtained by summing the elements in each row of the cell-gene matrix. Percent.MT refers to the proportion of mitochondrial genes in each cell. This is done by first determining which genes are mitochondrial, then counting the number of mitochondrial genes in each cell, dividing the result by the nFeatures value for that cell, and finally multiplying by 100 to obtain the percent.MT value. Thus, the quality control standards of this application may include: at least 50% of the cells in a tissue section must simultaneously meet the following requirements: (1) the number of genes detected in the cells (nFeature) is greater than 100; and (2) the proportion of mitochondrial genes in the cells (percent.MT) is less than 20%. In other words, only when at least 50% of the cells in the initial sample section meet the preset quality control standards can it be considered of qualified quality, and then the initial sample section can be used as the first sample section. The quality control standards can be flexibly set as needed, but it must be ensured that at least 50% of the cells must meet the quality control standards at the same time.

[0053] In the above embodiment, the present application can obtain the expression profile of genes from these spot points, thereby analyzing the idle data, and the present application can determine the cell outline based on the enrichment of spot point genes or any method known to technicians in the field for segmenting cells, such as ssDNA staining or DAPI staining, and circle the cells (i.e., cell segmentation) to achieve the effect of single-cell resolution and obtain multiple cellbins. Specifically, the present application can generate a cell-gene matrix corresponding to a cellbin according to the SAW process (Stereo-seq Analysis Workflow), and perform quality control on multiple initial sample slices based on the matrix to select initial sample slices with better quality for subsequent processing, which can improve the prediction accuracy of functional units with spatial structures.

[0054] In some embodiments, the step of performing cell classification on the plurality of cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of the plurality of cells in the first sample slice may specifically include: Using single-cell transcriptome sequencing technology, obtaining single-cell transcriptome sequencing data of multiple cells from the source sample in the first sample slice, and determining common gene expression data for each cell in the first sample slice based on the expression levels of the transcriptomes of the multiple cells in the first sample slice and the single-cell transcriptome sequencing data; Cell annotation is performed on multiple cells in the first sample slice based on the common gene expression data and preset cell marker genes to obtain first cell types of the multiple cells in the first sample slice.

[0055] Single-cell transcriptome sequencing data refers to data generated using single-cell sequencing or single-nucleus RNA sequencing (snRNA-seq). This data is equivalent to a reference dataset, which is used to perform cell annotation. Common gene expression data refers to the intersection of the expression levels of transcriptomes of multiple cells in the first sample slice and the gene data in the single-cell transcriptome sequencing data.

[0056] Furthermore, the present application can perform cell classification on multiple cells in the first sample slice based on common gene expression data and preset cell marker genes to obtain a first cell type for each cell in the first sample slice. Preset cell marker genes are genes that are specifically or highly expressed in a particular cell type. These genes can serve as markers for identifying cell types. The first cell type is equivalent to a cell ID in the first sample slice. This is because each classified cell is assigned an ID upon being circled, representing the identity of the cell (and the expression of all genes within this range also belongs to this cell). Furthermore, after cell classification, a column annotation related to this ID can be added to the spatial expression matrix of the spatial transcriptome data (a matrix used to represent the expression levels of the transcriptomes of multiple cells in the first sample slice), such as cell1: Neuron. In this way, after classifying each cell through the algorithm, the present application can map the multiple first cell types obtained to a spatial expression matrix with single-cell resolution, that is, each cell in the first sample slice is labeled according to the determined cell ID, without having to regenerate the matrix.

[0057] It is understandable that after determining the first sample slice of qualified quality, the present application can use the Spatial-ID algorithm to classify the cells in the first sample slice of qualified quality. The algorithm combines the common genes of snRNA-seq data and Stereo-seq data (i.e., target spatial gene expression data) to predict cell types.

[0058] In some embodiments, the step of determining the common gene expression data of each cell in the first sample slice based on the expression levels of the transcriptomes of multiple cells in the first sample slice and the single-cell transcriptome sequencing data may specifically include: determining a first gene expression matrix of the plurality of cells in the first sample slice based on the expression levels of the transcriptomes of the plurality of cells in the first sample slice; determining a second gene expression matrix for a plurality of cells in the first sample slice based on the single-cell transcriptome sequencing data; Common gene expression data of each cell in the first sample slice is obtained according to the common genes of the plurality of cells in the first gene expression matrix and the second gene expression matrix.

[0059] Among them, the present application can first extract the spatial gene expression data of the first sample slice (i.e., the expression level of the transcriptome of multiple cells in the first sample slice) and the gene expression matrix corresponding to the single-cell transcriptome sequencing data, and then use the intersection of the two gene expression matrices as common gene expression data. The first gene expression matrix refers to a matrix constructed based on the expression level of the transcriptome of multiple cells in the first sample slice. The rows of the matrix usually represent genes, the columns represent cells in the first sample slice, and the elements in the matrix represent the expression level of a certain gene in a certain cell, which is used to describe the gene expression of cells in the first sample slice. The second gene expression matrix refers to a matrix constructed based on single-cell transcriptome sequencing data, which also uses genes as rows and cells as columns. The elements represent the expression level of genes in cells. The matrix does not contain spatial position information of cells. Common gene expression data refers to gene expression information in the gene expression data of each cell in the first sample slice that exists in both the first gene expression matrix and the second gene expression matrix. These genes are reflected in both data sources and can be used to comprehensively analyze the gene expression characteristics of cells.

[0060] In some specific embodiments, the present application can optimize cell type probabilities using spatial neighborhood information and gene expression profiles through a graph convolutional network (GCN) and self-supervised learning. Ultimately, the cell type of each cell is determined based on the highest probability generated by the GCN to determine the identity of each cell in the spatial chip (i.e., the first cell type).

[0061] For example, Figure 2 , which is a spatial distribution diagram of cell types in the hippocampus provided in the examples of the present application. Figure 2 The figure shows the spatial distribution of cells after cell classification for the control group chip (the chip here is the slice) and the disease group chip. Figure 2 The left figure shows the spatial distribution of cell annotation results of the control group chip. Figure 2The right figure in the figure shows the spatial distribution of the cell annotation results of the disease group chip, and different colors represent different cell types. In this way, the application can distinguish cell types according to different colors.

[0062] In step S1.2 of some embodiments, the second cell type is the cell type of cells associated with the functional unit in the second sample slice. The first inter-cell distance data is used to describe the relative distances in spatial distribution between cells of the same or different cell types in the second sample slice. This data can reflect differences in spatial distribution between cells of the same or different cell types, such as adjacent cells or cells that are far apart. Staining data refers to data obtained after a specific staining process is performed on the second sample slice. Staining can mark specific structural or molecular characteristics of cells, facilitating subsequent analysis of information such as inter-cell distances. The stained image can be a pathological staining image. The second sample slice can be the first sample slice or an adjacent slice of the first sample slice. Adjacent slices refer to continuous slices that are located close to the first sample slice in the three-dimensional structure of the tissue, without limitation. Therefore, by obtaining staining data, the cell types and cell spatial distribution required to construct the functional unit can be determined. This application uses the neurovascular unit (a functional unit) as an example for detailed description in the following embodiments.

[0063] It should be noted that in the following embodiments, the present application uses the neurovascular unit (a functional unit) as an example for detailed description. In order to enhance the robustness of the constructed functional unit prediction model in neurovascular unit analysis, the present application may also implement data enhancement strategies (such as image transformation of staining data): while maintaining cell attribution and biological structural constraints (such as neuron-vascular connection topological invariance), introduce geometric diversity through random rotation (such as rotation of ±45°), flipping, translation (such as moving a distance of ±15% of the size), etc.; and / or, combine optical adjustment (Gaussian blur to simulate defocus, Poisson noise to simulate low-light imaging, channel perturbation to adapt to staining differences) and biologically specialized deformation (local occlusion to avoid key synapses, elastic deformation to protect unit integrity).

[0064] Therefore, the present application can perform spatial analysis on the second sample slice based on immunostaining techniques to analyze the spatial distribution of cells of the same or different cell types within spatially structured functional units (such as astrocytes, endothelial cells, and neurons in the neurovascular unit). The relative distances between cells of the same or different cell types (such as the Euclidean distance between cells) are counted and extracted, and these distance feature data are organized for subsequent model construction.

[0065] It should be noted that the staining data can be image data obtained through a specific staining technique, such as immunostaining. The first sample slice is processed to render cells in the first sample slice with different colors or markers, facilitating subsequent observation and analysis of cell distribution and characteristics. Determining the first intercellular distance data can utilize image analysis software (such as ImageJ) or bioinformatics tools (such as phenoptr) to calculate the relative distances between identical or different cells in the second sample slice based on the staining data and the cells' spatial position data. For example, algorithms such as Euclidean distance can be used to obtain distance information between each cell and other cells, thereby describing differences in the spatial distribution of cells. Alternatively, the center of the cell nucleus in the stained image corresponding to the staining data can be marked, and then the image can be read using skimage methods to calculate the distance statistics. For example, since cell type labels can include vascular cells (endothelial cells and pericytes), astrocytes, neurons (inhibitory neurons and excitatory neurons), etc., by analyzing the relative spatial distribution between these cells, the distance between vascular cells and the distance between astrocytes, neurons and vascular cells can be determined, which helps to deeply understand the spatial structure of the neurovascular unit and its functional associations, and provide key distance feature data for the accurate prediction and analysis of the neurovascular unit.

[0066] It should be noted that the image corresponding to the staining data can be a staining image obtained by staining one or more of the cell nucleus, nuclear membrane, cytoplasm, and cell membrane, such as the cell nucleus and cytoplasm, and only specific cells in the immune library can be stained.

[0067] It should be noted that the dataset for each set of sample slices containing a functional unit in this application may also include spatial transcriptome data for a first cell type, as well as data for a fourth cell type and third intercellular distances for cells associated with the functional unit in a second sample slice. Another dataset is determined based on the staining data after image transformation of the second sample slice. In other words, there is no need to provide actual biochemical slices to obtain the additional dataset.

[0068] For example, Figure 3 The figure shows the immunohistochemical fluorescence image of the hippocampus provided in the embodiment of the present application. Figure 3 , which shows the spatial distribution of cells of the neurovascular unit in the immunohistochemical staining slices of the control group chip (the chip here is the slice, i.e., corresponding to the second sample slice of the present application) and the disease group chip. Figure 3 The left image shows the spatial distribution of cells in the neurovascular unit in the immunohistochemical staining section of the control group chip. Figure 3The right image shows the spatial distribution of cells within the neurovascular unit (NUVU) in immunohistochemically stained sections of the disease panel. The dashed circle indicates an artificially labeled functional unit of the NUVU. The image shows the distribution of astrocytes (Astrocytes), neurons (Neurons), and the fluorescent dye (DAPI) for the nucleus.

[0069] In step S2 of some embodiments, the present application may further train a machine learning module based on a dataset of at least one set of sample slices to construct a functional unit prediction model. Because the functional unit prediction model is constructed based on information related to the functional unit indicated by the second sample slice, it can be used to predict the spatial structure of a functional unit identical to the functional unit. After the functional unit prediction model is trained, the functional unit prediction model can be used to accurately predict the composition and function of functional units with spatial structures.

[0070] It should be noted that the machine learning module of the present application includes at least one of a supervised learning module, a semi-supervised learning module, an unsupervised learning module, a regression analysis module, a reinforcement learning module, a self-learning module, a feature learning module, a sparse dictionary learning module, anomaly detection module, a generative adversarial network or an association rule module, without limitation.

[0071] It should be noted that this application can make predictions based on transcriptomes at single-cell resolution, so it is relatively accurate. This application can analyze the specific cell types contained in the spatial structure of the predicted functional units to view the characteristics of pathological changes, enhance the ability to analyze the dynamic changes of functional units with spatial structures, and can accurately predict their changes under different physiological or pathological conditions, providing a basis for the formulation of personalized treatment plans.

[0072] In some embodiments, the step of training a machine learning module based on a dataset of at least one set of sample slices containing functional units to construct a functional unit prediction model may specifically include: constructing spatial adjacency data for the first sample slice based on the spatial transcriptome data containing the first cell type, wherein the spatial adjacency data is used to describe the adjacency structure between the same or different cells in the first sample slice; extracting cell space data related to the functional unit from the spatial adjacency data based on a matching result between the second cell type and the first cell type of cells related to the functional unit and the first inter-cell distance data; determining, from the first cell type, a central cell type, an associated cell type, and distance thresholds between central cells, between associated cells, and between central cells and associated cells according to the spatial structure indicated by the unit spatial data; central cells refer to cells constituting a functional unit, and associated cells refer to non-central cells that have intercellular communication with central cells; The machine learning module is trained based on the central cell type, associated cell type, and distance threshold to construct a functional unit prediction model.

[0073] Wherein, the spatial transcriptome data containing the first cell type at this time can be presented in the form of a matrix, and the rows of the matrix represent genes, and the columns represent cells in spatial positions, so that the spatial position of each cell in the first sample slice can be determined. The application can set a neighboring distance threshold value to determine which cells are considered to be adjacent to each other to construct spatial adjacency data, and the spatial adjacency data at this time can be in matrix form, i.e., a spatial adjacency matrix. The spatial adjacency matrix can be used to describe the adjacency structure between the same or different cells in the first sample slice. Therefore, the application can establish a spatial adjacency matrix according to the first sample slice, and extract the distance characteristics of the cells in the neurovascular unit from the immunohistochemical staining of the first sample slice or the adjacent second sample slice, and then, establish a functional unit prediction model based on the information integration of the two.

[0074] The second cell type of the functional unit-related cells can be used to determine the cell type required to construct the functional unit. The second cell type of the functional unit-related cells can then be matched with the first cell type, and cell spatial data associated with the functional unit can be extracted from the spatial adjacency data based on the relative distances between identical or different cells in the second sample slice. The cell spatial data in this case indicates the spatial structure associated with the spatial distribution of the cells associated with the functional unit matched from the first sample slice.

[0075] The central cell type is used to indicate the cell type that plays a key role in the functional unit. For example, when the target functional unit is a neurovascular unit, the corresponding central cell types may include endothelial cells and pericytes. The central cells indicated by the central cell type refer to the cells that constitute the functional unit. The associated cell type is used to indicate the cell types that are allowed to exist in the functional unit, such as neurons and astrocytes. The associated cells indicated by the associated cell type refer to non-central cells that have intercellular communication with the central cells. By matching the first cell type according to the central cell type and the associated cell type, the central cells and associated cells related to the spatial structure of the functional unit can be screened out from the cells indicated by multiple first cell types, thereby avoiding the influence of invalid cell types on the specific spatial structure prediction of the functional unit.

[0076] The distance threshold between a central cell type and its associated cell types can refer to the distance threshold between central cells, between associated cells, and between central cells and associated cells. In other words, it refers to the maximum distance between identical or different cells within the same functional unit. When the distance between a central cell and an associated cell exceeds the distance threshold, they belong to different functional units.

[0077] It should be noted that the process of constructing the functional unit prediction model based on the central cell type, associated cell type and distance threshold in this application can be specifically expressed as follows: First, a central cell set is constructed based on the central cell type, which can be expressed as " ",in, represents the central cell set selected from multiple cells in the first sample slice, j represents the cell corresponding to the central cell type in the first sample slice, and cells represents the set of multiple cells in the first sample slice. represents the cell type corresponding to cell j, Represents a set of central cell types. Further, determining the cell process that matches the second cell type associated with the functional unit in the first sample slice can be equivalent to: constructing a distance matrix based on the first inter-cell distance data using the Euclidean distance between the points (the point refers to the distance between the center points of two cells) ,like ,in, Represents the cell pairs in the distance matrix ( ) between the relative distances between the cell pairs ( ) represents cell pairs of the same cell type or different cell types in the second sample slice, and Respectively represent the spatial position data of the corresponding cells, and ( ) represents the candidate distance between each pair of cells.

[0078] Based on this, the functional unit prediction model with spatial structure constructed based on the distance matrix, central cell type, associated cell type and distance threshold can be expressed as the following formula:

[0079] in, Indicates that a cell set matching the central cell type is selected from multiple cells in the first sample slice. is the distance matrix, Indicates the matrix The data corresponding to the coordinates, d represents the distance threshold, represents a collection of associated cell types, The different associated cell types in the functional unit are indicated respectively. Represents cells in the first sample slice The corresponding first cell type belongs to an associated cell type, Indicates the number The functional unit can include multiple functional units in a sample slice. The above formula expresses that by traversing all cells in the first sample slice, combining the distance matrix, the central cell set, and the set of associated cell types, all cells contained in the functional unit are found and integrated to form at least one functional unit. For example, when the functional unit is a neurovascular unit (NVU), the constructed functional unit prediction model is an NVU model.

[0080] The distance threshold refers to the maximum distance between identical or different cells in the same functional unit. Therefore, when the intercellular distance between an associated cell and the central cell is greater than the distance threshold, the cell indicated by the associated cell type is considered to be a cell that does not meet the structural requirements of the functional unit to which the central cell belongs.

[0081] In the above embodiment, the present application can establish a distance matrix for cells within a spatially structured functional unit by combining a Euclidean distance algorithm. Combining the definition of a spatially structured functional unit with the spatial positions of cells, a machine learning approach can be used to train a functional unit prediction model. For example, the distance differences between endothelial cells, astrocytes, and neurons can reflect the functional state of the blood-brain barrier and the integrity of the neurovascular unit. The model is fed with distance features between cells, such as: celltype1-celltype2: 30 (30 distance units); celltype1-celltype2: 350 (350 distance units), to predict spatially structured functional units. The established functional unit prediction model traverses the central cell set, identifies all cell types in the distance matrix that meet the aforementioned distance feature set conditions, and then defines them as the same unit. This functional unit prediction model predicts the functional structure of a space.

[0082] For example, Figure 4 As shown, it is a schematic diagram of constructing a functional unit prediction model based on machine learning provided by an embodiment of the present application. Among them, the present application can first establish the spatial adjacency data 410 of the first sample slice, that is, the proximity matrix, based on the spatial transcriptome data of the first sample slice. Secondly, the first intercellular distance data 420 of the cells in the neurovascular unit can be extracted from the immunohistochemical staining (i.e., staining data) of the first sample slice itself or the adjacent slice (the continuous slice that is close to the position of the spatial transcriptome chip in the three-dimensional structure of the tissue), and the functional unit prediction model 430 can be established based on the integration of the information of the two.

[0083] In some embodiments, the present application can perform functional unit prediction on the third sample slice through a functional unit prediction model, and compare it with the staining data of the fourth sample slice to determine the accuracy level of the machine learning module, wherein the fourth sample slice is the third sample slice or an adjacent slice of the third sample slice.

[0084] That is, after establishing a functional unit prediction model 430 (e.g., a prediction model based on a neurovascular unit), the present application can also use an immunohistochemical (i.e., immunofluorescence)-stained experimental slice (i.e., a fourth sample slice) as a validation standard. The model prediction results (i.e., the spatial distribution of target cells corresponding to the predicted neurovascular unit in the third sample slice) are compared with the spatial distribution of the immunohistochemical staining image, and the spatial position of the predicted neurovascular unit is matched with the positions of the labeled blood vessels and neurons in the immunohistochemical staining image. Furthermore, by calculating the overlap rate (i.e., cell coincidence) between the spatial distribution of multiple cells in the predicted neurovascular unit and the spatial distribution of cells in the actual labeled functional unit, the intersection percentage of the predicted neurovascular unit area and the immunofluorescence-labeled area is determined, thereby statistically analyzing the prediction accuracy of the functional unit prediction model 430. Experimental verification results show that the overlap between the predicted neurovascular unit and the immunofluorescence staining image is approximately 75%, which also reflects the model's accuracy in spatial structure prediction. After accurately determining the cells corresponding to the functional unit, the relationship between the cells in the functional unit can be clarified by combining the cell type characteristics and their spatial location.

[0085] It should be noted that after determining the cells corresponding to the functional units, the prediction results can also be visualized to intuitively display the spatial organizational structure and functional characteristics of the functional units with spatial structures.

[0086] It should be noted that during the training of the functional unit prediction model, this application can predict and analyze neurovascular units in human hippocampal samples as an example. The original process of this application through the distance feature extraction model can be divided into four parts: data preprocessing, building a convolutional neural network (CNN) model, training the model, and extracting features. Specifically, the image data is first preprocessed to accurately label different cell types (such as Neuron, Astro, etc.) and neurovascular units, that is, to clarify the unit and category to which the cells belong, and the image pixel values are normalized to the range of 0-1. Next, build the CNN model. Specifically, you can go through the first convolutional layer (that is, using 16 3x3 convolution kernels with a stride of 1, a padding of 1, and the activation function of ReLU to extract preliminary image features), the first pooling layer (that is, using a 2x2 max pooling layer with a stride of 2 to downsample the feature map to reduce the feature dimension), the second convolutional layer (that is, using 32 3x3 convolution kernels with a stride of 1, a padding of 1, and the activation function is still ReLU to further extract more complex features), the second pooling layer (that is, using a 2x2 max pooling layer with a stride of 2 for downsampling), and the fully connected layer (that is, flattening the pooled features and connecting them to a fully connected layer with 128 neurons and the activation function of ReLU). During training, the preprocessed data can be divided into a training set and a validation set. Multiple rounds of training are performed on the training set. The loss is calculated and the model parameters are updated in each round. At the same time, the model performance is monitored on the validation set to prevent overfitting. After the model is trained, the test image is input into the model and feature vectors are obtained from the output layer. These feature vectors contain information such as the distance characteristics between different cell types and can be used for subsequent analysis.

[0087] In some embodiments, as Figure 5 The figure shows a schematic diagram of the NVU model prediction convergence curve provided in the embodiment of the present application. The figure shows the trend of the accuracy of the NVU model in predicting NVU in immunofluorescence slices as the number of training iterations increases, that is, when the number of training iterations is higher, the model's prediction results for functional units with spatial structures are more accurate and stable.

[0088] In some embodiments, Figure 2 Each cell in the chip data has its own gene expression profile and position coordinates, so this application can extract the position coordinates corresponding to each cell type and apply them to the distance-based model (such as the NVU model) to obtain Figure 6 Regarding the control group chip (such as Figure 6 left image) and disease group chips (such as Figure 6 Figure 3 (right panel) shows the spatial distribution of functional units of the neurovascular unit, with different colors representing different cell types.

[0089] In some embodiments, the present application also provides a functional unit prediction method, including but not limited to the following steps: Acquiring spatial transcriptome data of the slice to be predicted, and performing cell classification on a plurality of cells in the slice to be predicted based on the spatial transcriptome data of the slice to be predicted to obtain a third cell type of the plurality of cells in the slice to be predicted, wherein the spatial transcriptome data of the slice to be predicted includes expression levels and spatial distribution information of transcriptomes of the plurality of cells in the slice to be predicted; According to the constructed functional unit prediction model, functional unit prediction is performed based on the spatial transcriptome data containing the third cell type of the slice to be predicted.

[0090] Among them, in the actual application stage, the present application can perform cell classification on multiple cells in the slice to be predicted based on the spatial transcriptome data of the slice to be predicted, and the classification process has been described in detail in the above embodiment. The third cell type is used to characterize the cell type corresponding to each cell in the slice to be predicted. The spatial transcriptome data of the slice to be predicted includes the expression level and spatial distribution information of the transcriptomes of multiple cells in the slice to be predicted.

[0091] In some embodiments, the step of performing functional unit prediction based on the spatial transcriptome data of the slice to be predicted containing the first cell type according to the constructed functional unit prediction model may specifically include: According to the constructed functional unit prediction model, the central cell type and associated cell type related to the functional unit, as well as the distance thresholds between central cells, between associated cells, and between central cells and associated cells are determined; determining the central cell of the functional unit from a plurality of cells in the slice to be predicted based on the matching results of the central cell type and the third cell type; determining the associated cell of the functional unit from a plurality of cells in the slice to be predicted based on the matching results between the associated cell type and the third cell type; obtaining second cell distance data between central cells, between associated cells, and between central cells and associated cells in the slice to be predicted according to the spatial transcriptome data, wherein the second inter-cell distance data is used to describe the relative distances between central cells, between associated cells, and between central cells and associated cells in the slice to be predicted; According to the comparison result of the relative distance and the distance threshold, a target cell related to the functional unit to which the central cell belongs is determined from multiple cells in the slice to be predicted, and the target cell includes the central cell and the associated cell; The functional units of the slice to be predicted are predicted based on the target cells and the spatial distribution information of the target cells in the slice to be predicted.

[0092] Among them, according to the formula constructed according to the above model, the functional unit prediction model predefines the central cell type set, the associated cell type set, and the distance thresholds between central cells, between associated cells, and between central cells and associated cells corresponding to the functional unit. Therefore, it is possible to match the third cell type based on these defined sets, and screen out central cells and associated cells related to the functional unit from multiple cells of the slice to be predicted. Central cells can refer to key cells that constitute the structure of the functional unit. Associated cells can refer to non-central cells that have intercellular communication with central cells, that is, cells that the functional unit allows to exist, and the spatial distance between associated cells and central cells in the same functional unit should be less than or equal to the distance threshold.

[0093] Furthermore, target cells are cells in the slice to be predicted that can be divided into the spatial structure of functional units. After the target cells are determined, the spatial structure of the functional units of the slice to be predicted can be constructed based on the spatial distribution information of the target cells in the slice to be predicted, thus achieving the prediction of the spatial structure of the functional units.

[0094] In some embodiments, after performing functional unit prediction based on the spatial transcriptome data of the slice to be predicted containing the third cell type according to the constructed functional unit prediction model, the method further comprises: Extracting the expression level and spatial distribution information of the transcriptome of the target cell in the slice to be predicted from the spatial transcriptome data of the slice to be predicted; Based on the expression level and spatial distribution information of the transcriptome of the target cell in the slice to be predicted, density statistics and functional analysis are performed on the functional units in the slice to be predicted.

[0095] The expression level of the target cell's transcriptome in the slice to be predicted is equivalent to cellular gene expression data, which refers to data that can reflect the gene expression status within the cell and can indicate the activity level of the genes within the cell, that is, which genes are expressed and in what amounts. The spatial distribution information of the target cell's transcriptome in the slice to be predicted is equivalent to cell position data, which refers to data that can describe the specific location of each cell in the first sample slice and is typically represented by coordinates, such as coordinates in a two-dimensional plane or coordinates in three-dimensional space.

[0096] This application can determine the spatial distribution characteristics of the predicted functional units in the target slice based on the expression level and spatial distribution information of the transcriptome of the target cells in the slice to be predicted. The spatial distribution characteristics can be used to describe the spatial distribution law and pattern of the target cells in the functional unit, such as whether the cells are evenly distributed, clustered, or distributed in a specific geometric shape, so as to further realize the density statistics and functional analysis of the functional units in the predicted slice. This application can comprehensively analyze the gene expression data of the target cells, the corresponding cell types and spatial distribution characteristics, and can use a variety of analysis methods, such as differential expression analysis, enrichment analysis, cluster analysis, etc., to reveal the laws and mechanisms of gene expression in the functional units. For example, differential expression analysis can find genes that are differentially expressed in different cell types or spatial locations; enrichment analysis can determine the biological functions and signaling pathways involved in these differentially expressed genes; cluster analysis can classify cells with similar gene expression patterns and spatial distribution characteristics. In this way, by gaining a deep understanding of the gene expression regulation mechanism of the target functional unit, the interaction between cells, and its function in biological processes, it can provide important theoretical basis and practical guidance for fields such as disease diagnosis and drug development. Therefore, based on the spatial transcriptome data with single-cell resolution, this application can use the functional unit prediction model established based on the intercellular distance to predict the functional units, and analyze the composition, distribution and association of the functional units with other brain regions through the prediction results.

[0097] For example, after predicting the spatial structure of the neurovascular unit, the present application can also perform density statistics and functional analysis on the prediction results to intuitively display the spatial organizational structure and functional characteristics of the neurovascular unit. Figure 7 、 Figure 8 and Figure 9 The figure shows the density statistics and functional score changes of the neurovascular unit functional units of the control group and disease group chips provided in the embodiment of the present application. Figure 7 The horizontal axis indicates two groups of samples (control group chip and disease group chip), and the vertical axis indicates the average density (Average Density), the unit is count per square millimeter (count / mm 2 ), which is used to measure the density of related indicators on the chip. It can be seen that in terms of density changes, the average density of the disease group is significantly higher than that of the control group (the difference in the height of the bar graph is obvious), indicating that the target cells in the disease state are actively infiltrating or proliferating in the tissue. Figure 8The horizontal axis indicates the two groups of samples (control group chip and disease group chip), and the vertical axis indicates the positive regulation score (Positive Regulation Score, ranging from about -0.5 to 1.5). It can be seen that in response to the hypoxia pathway, hypoxia-related genes (such as HIF1α, VEGF, GLUT1) are highly expressed in the disease group, which may drive angiogenesis, metabolic reprogramming and tumor progression (in the case of tumor research). Figure 9 , the horizontal axis indicates the two groups of samples (control group chip and disease group chip), and the vertical axis indicates the positive regulation score (Positive Regulation Score, ranging from approximately -1 to 1). It can be seen that in terms of immune response regulation, the disease group is significantly activated (positive regulation score increased), that is, key immune pathways (such as TNFα / NF-κB, IFNγ signaling) or immune checkpoint genes (such as PD-L1, CTLA4) are upregulated, which can indicate immune microenvironment remodeling. It can be seen that this application can not only efficiently predict the structure of functional units with spatial structure, but also reveal the dynamic changes of different cell types in space, provide a deep understanding of the health status of the nervous system, and thus provide support for the early diagnosis and precise treatment of neurovascular diseases.

[0098] Reference Figure 10 In a specific embodiment, the method for constructing a functional unit prediction model provided in the embodiment of the present application may include the following steps: Step S1001: Acquire spatial transcriptome data in multiple sample slices.

[0099] The spatial transcriptome data of the initial sample slice includes spatial gene expression data and spatial position data of multiple cells in the initial sample slice. Before predicting the functional units with spatial structure, the present application can first prepare the slice and perform quality control on the slice to screen out the first sample slice that meets the quality control standards.

[0100] Step S1002, performing data extraction based on the spatial transcriptome data to obtain the number of genes and the ratio of mitochondrial genes in each cell in each sample slice; Step S1003 : screening multiple sample slices based on the number of genes and the ratio of mitochondrial genes in each cell to determine a first sample slice.

[0101] The first sample slice is the initial sample slice that meets quality control standards. These standards may include: at least 50% of the cells in a tissue slice must simultaneously meet the following requirements: the average number of genes per cell (nFeature) must be greater than 100, and the average proportion of mitochondrial genes per cell (percent.MT) must be less than 20. In other words, chips that meet these standards are considered of qualified quality and can be used for predictive analysis of neurovascular units.

[0102] Step S1004 : Acquire spatial transcriptome data of the first sample slice, and perform cell classification on a plurality of cells in the first sample slice based on the spatial transcriptome data to obtain cell type labels of the plurality of cells in the first sample slice.

[0103] Among them, after determining the first sample slice, the present application can obtain spatial transcriptome data with single-cell resolution for each first sample slice, and perform cell annotation based on the single-cell spatial transcriptome data.

[0104] Step S1005 : Acquire staining data of the second sample slice, and determine the second cell type and the first inter-cell distance data of the functional unit-related cells in the second sample slice based on the staining data.

[0105] The first inter-cell distance data is used to describe the relative distance between cells of the same cell type or different cell types in the second sample slice, and the second sample slice is the first sample slice or an adjacent slice of the first sample slice. The present application can further perform statistics on the cell distance characteristics within the functional unit, that is, the distance characteristics between different cell types can be confirmed and counted based on the immunohistochemical staining results in the sample slice itself or the corresponding adjacent slices.

[0106] Step S1006 : constructing spatial adjacency data of the first sample slice based on the spatial transcriptome data containing the first cell type.

[0107] Among them, the spatial adjacency data is used to describe the adjacency structure between the same or different cells in the first sample slice. The preset label screening conditions are equivalent to the conditions for constructing a model for structural prediction of the target functional unit. Different target functional units correspond to different preset label screening conditions.

[0108] Step S1007 : extracting unit space data related to the functional unit from the spatial adjacency data based on the matching results of the second cell type and the first cell type of the cells related to the functional unit and the first inter-cell distance data.

[0109] Step S1008 , determining the central cell type, associated cell types, and distance thresholds between central cells, between associated cells, and between central cells and associated cells from the first cell type according to the spatial structure indicated by the unit space data.

[0110] Different functional unit predictions may have different distance thresholds between central cells, between associated cells, and between central cells and associated cells. Central cells refer to cells that make up a functional unit, while associated cells refer to non-central cells that have intercellular communication with central cells.

[0111] Step S1009: training a machine learning module based on the central cell type, associated cell types, and distance thresholds to construct a functional unit prediction model.

[0112] Among them, after determining the central cell type, associated cell type and distance threshold, a spatial relationship model between cells in the functional unit can be established to quantify the relative distance between cells and their spatial distribution characteristics. By traversing each cell, the cell corresponding to the functional unit can be determined, thereby determining the spatial structure corresponding to the functional unit.

[0113] This application defines spatially structured functional units at the transcriptome level for the first time, applies this prediction method to neurological disease research, and explores the changing patterns of spatially structured functional units in neurological diseases. Based on the constructed functional unit prediction model, new experimental strategies or intervention methods are designed to promote the early diagnosis and treatment of diseases related to spatially structured functional units. In addition, this application has low professional requirements for professionals and causes less damage to samples. It can effectively obtain high-quality transcriptome data, providing great convenience for scientific researchers.

[0114] The functional unit prediction method provided in the embodiments of the present application significantly improves the spatial resolution and cell analysis accuracy of functional units with spatial structure through single-cell resolution spatial transcriptome technology, and can accurately reveal the spatial position of cells and their subtle changes, thereby improving the accuracy of analysis. Compared with related technologies, the present application reduces the reliance on slice samples, avoids complex immunohistochemical staining and microscopy operations, and by combining the Euclidean distance algorithm and machine learning methods, a more accurate functional unit prediction model can be established, which can more accurately predict the composition and function of functional units with spatial structure, thereby promoting the early diagnosis and treatment of neurological diseases (such as stroke, Alzheimer's disease, etc.). At the same time, the present invention can enhance the ability to analyze the dynamic changes of functional units with spatial structure, that is, by determining the target functional unit through intercellular distance data and cell type labels, it can accurately predict its changes under different physiological or pathological conditions, providing a basis for the formulation of personalized treatment plans. In addition, the present application uses automated machine learning analysis methods to make the prediction process more efficient and accurate, reduce manual intervention, save experimental time, reduce costs and technical barriers, and provide convenience for more scientific researchers. Therefore, the present application can be better applied to the prediction of functional units with spatial structures, improve the accuracy of prediction of functional units with spatial structures, thereby providing effective technical support for the early diagnosis and precise treatment of neurological diseases, and has broad application prospects.

[0115] Reference Figure 11 , the embodiment of the present application further provides a functional unit prediction model construction device, the device comprising: The data set acquisition module 1110 is used to acquire a data set of at least one set of sample slices containing functional units. The data set acquisition module 1110 includes: a first cell classification module 1111 and a second cell classification module 1112. The first cell classification module 1111 is used to acquire spatial transcriptome data containing a first cell type. The acquisition method includes: acquiring spatial transcriptome data of the first sample slice, and performing cell classification on multiple cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of multiple cells in the first sample slice; the spatial transcriptome data includes expression levels and spatial distribution information of the transcriptomes of the multiple cells in the first sample slice; the second cell classification module 1112 is used to acquire second cell types and first inter-cell distance data of cells associated with the functional unit. The acquisition method includes: acquiring staining data of the second sample slice, determining the second cell types and first inter-cell distance data of cells associated with the functional unit in the second sample slice based on the staining data, and the first inter-cell distance data is used to describe the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or an adjacent slice of the first sample slice; The model building module 1120 is used to train a machine learning module based on a data set of at least one set of sample slices to build a functional unit prediction model.

[0116] It can be seen that the contents of the above-mentioned functional unit prediction model construction method embodiment are all applicable to the embodiment of the present functional unit prediction model construction device. The functions specifically implemented by the present functional unit prediction model construction device embodiment are the same as those in the above-mentioned functional unit prediction model construction method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned functional unit prediction model construction method embodiment.

[0117] Reference Figure 12 , an embodiment of the present application further provides a functional unit prediction device, the device comprising: A third cell classification module 1210 is configured to obtain spatial transcriptome data of the slice to be predicted, and perform cell classification on multiple cells in the slice to be predicted based on the spatial transcriptome data of the slice to be predicted to obtain a third cell type of the multiple cells in the slice to be predicted, wherein the spatial transcriptome data of the slice to be predicted includes expression levels and spatial distribution information of the transcriptomes of the multiple cells in the slice to be predicted; The unit prediction module 1220 is configured to perform functional unit prediction based on the spatial transcriptome data of the slice to be predicted containing the third cell type according to the constructed functional unit prediction model.

[0118] It can be seen that the contents of the above-mentioned functional unit prediction method embodiment are all applicable to the embodiments of the present functional unit prediction device. The functions specifically implemented by the present functional unit prediction device embodiment are the same as those in the above-mentioned functional unit prediction method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned functional unit prediction method embodiment.

[0119] Reference Figure 13 , Figure 13 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes: The processor 1301 may be implemented as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application. The memory 1302 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1302 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1302 and is called by the processor 1301 to execute the functional unit prediction model construction method and the functional unit prediction method of the embodiments of this application. Input / output interface 1303, used to implement information input and output; Communication interface 1304, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.); Bus 1305 , which transmits information between various components of the device (e.g., processor 1301 , memory 1302 , input / output interface 1303 , and communication interface 1304 ); The processor 1301 , the memory 1302 , the input / output interface 1303 and the communication interface 1304 are connected to each other in communication within the device via a bus 1305 .

[0120] An embodiment of the present application also provides a computer-readable storage medium, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the above-mentioned functional unit prediction model construction method and functional unit prediction method are implemented.

[0121] The present application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device implements the above-mentioned functional unit prediction model construction method and functional unit prediction method.

[0122] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present disclosure described herein, for example, can be implemented in orders other than those illustrated or described herein. In addition, the terms "comprises" and "comprising," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.

[0123] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0124] It should be understood that in the description of the embodiments of the present application, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.

[0125] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0126] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0127] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0128] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0129] It should also be understood that the various implementation methods provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0130] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.

Claims

1. A method for constructing a functional unit prediction model, characterized in that: The method comprises: Step S1: obtaining at least one set of data sets of sample slices containing functional units, wherein each set of data sets of sample slices containing functional units includes: Step S1.1, spatial transcriptome data containing a first cell type, the acquisition method comprising: acquiring spatial transcriptome data of a first sample slice, and performing cell classification on a plurality of cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of the plurality of cells in the first sample slice; the spatial transcriptome data comprising expression levels and spatial distribution information of transcriptomes of the plurality of cells in the first sample slice; and Step S1.2, obtaining the second cell type and first inter-cell distance data of the functional unit-related cells, including: obtaining staining data of a second sample slice, and determining the second cell type and first inter-cell distance data of the functional unit-related cells in the second sample slice based on the staining data, wherein the first inter-cell distance data is used to describe the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or an adjacent slice of the first sample slice; Step S2: training a machine learning module based on the data set of the at least one group of sample slices containing functional units to construct a functional unit prediction model.

2. The method for constructing a functional unit prediction model according to claim 1, wherein: The machine learning module includes at least one of a supervised learning module, a semi-supervised learning module, an unsupervised learning module, a regression analysis module, a reinforcement learning module, a self-learning module, a feature learning module, a sparse dictionary learning module, an anomaly detection module, a generative adversarial network or an association rule module.

3. The method according to claim 1, characterized in that The step of training a machine learning module based on a data set of at least one set of sample slices containing functional units to construct a functional unit prediction model comprises: constructing spatial adjacency data of the first sample slice according to the spatial transcriptome data containing the first cell type, wherein the spatial adjacency data is used to describe the adjacency structure between the same or different cells in the first sample slice; extracting cell space data related to the functional unit from the spatial adjacency data according to a matching result between a second cell type and the first cell type of cells related to the functional unit and the first inter-cell distance data; Determining, from the first cell type, a central cell type, an associated cell type, and distance thresholds between central cells, between associated cells, and between the central cell and the associated cells according to the spatial structure indicated by the unit spatial data; the central cell refers to a cell constituting the functional unit, and the associated cell refers to a non-central cell that has intercellular communication with the central cell; A machine learning module is trained according to the central cell type, the associated cell type, and the distance threshold to construct a functional unit prediction model.

4. The method for constructing a functional unit prediction model according to claim 1, wherein: The performing cell classification on the plurality of cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of the plurality of cells in the first sample slice includes: Using single-cell transcriptome sequencing technology, obtaining single-cell transcriptome sequencing data of multiple cells from the sample from which the first sample slice originates, and determining common gene expression data for each cell in the first sample slice based on the expression levels of the transcriptomes of the multiple cells in the first sample slice and the single-cell transcriptome sequencing data; Performing cell annotation on the plurality of cells in the first sample slice based on the common gene expression data and preset cell marker genes to obtain a first cell type of the plurality of cells in the first sample slice.

5. The method for constructing a functional unit prediction model according to claim 4, wherein: The determining, based on the expression levels of the transcriptomes of the plurality of cells in the first sample slice and the single-cell transcriptome sequencing data, common gene expression data of each cell in the first sample slice comprises: determining a first gene expression matrix of the plurality of cells in the first sample slice based on the expression levels of the transcriptomes of the plurality of cells in the first sample slice; Determine a second gene expression matrix of multiple cells in the first sample slice based on the single-cell transcriptome sequencing data; Common gene expression data of each cell in the first sample slice is obtained based on the common genes of the multiple cells in the first gene expression matrix and the second gene expression matrix.

6. The method for constructing a functional unit prediction model according to claim 1, wherein: Functional unit prediction is performed on the third sample slice using the functional unit prediction model, and compared with staining data of the fourth sample slice to determine the accuracy level of the machine learning module, wherein the fourth sample slice is the third sample slice or an adjacent slice of the third sample slice.

7. A functional unit prediction method, characterized in that: The method comprises: Acquiring spatial transcriptome data of the slice to be predicted, and performing cell classification on a plurality of cells in the slice to be predicted based on the spatial transcriptome data of the slice to be predicted to obtain a third cell type of the plurality of cells in the slice to be predicted, wherein the spatial transcriptome data of the slice to be predicted includes expression levels and spatial distribution information of transcriptomes of the plurality of cells in the slice to be predicted; The functional unit prediction model constructed according to the functional unit prediction model construction method according to any one of claims 1 to 6 performs functional unit prediction based on the spatial transcriptome data containing the third cell type of the slice to be predicted.

8. The method according to claim 7, characterized in that The functional unit prediction model constructed according to the method for constructing a functional unit prediction model according to any one of claims 1 to 6 performs functional unit prediction based on the spatial transcriptome data containing the third cell type of the slice to be predicted, comprising: The functional unit prediction model constructed by the functional unit prediction model construction method according to any one of claims 1 to 6, determining the central cell type, associated cell type, and distance thresholds between central cells, between associated cells, and between the central cell and the associated cells related to the functional unit; Determining a central cell of a functional unit from a plurality of cells in the slice to be predicted based on a matching result between the central cell type and the third cell type; Determining an associated cell of a functional unit from a plurality of cells in the slice to be predicted based on a matching result between the associated cell type and the third cell type; Acquire, according to the spatial transcriptome data, second cell distance data between the central cells, between the associated cells, and between the central cell and the associated cells in the slice to be predicted, wherein the second inter-cell distance data is used to describe the relative distances between the central cells, between the associated cells, and between the central cell and the associated cells in the slice to be predicted; Determining, based on a comparison result of the relative distance and the distance threshold, a target cell associated with the functional unit to which the central cell belongs from a plurality of cells in the slice to be predicted; the target cell includes the central cell and the associated cell; The functional units of the slice to be predicted are predicted according to the target cells and the spatial distribution information of the target cells in the slice to be predicted.

9. The method according to claim 8, characterized in that After the functional unit prediction model constructed according to the method for constructing a functional unit prediction model according to any one of claims 1 to 6 performs functional unit prediction based on the spatial transcriptome data of the slice to be predicted containing the third cell type, the functional unit prediction method further comprises: Extracting the expression level and spatial distribution information of the transcriptome of the target cell in the slice to be predicted from the spatial transcriptome data of the slice to be predicted; According to the expression level and spatial distribution information of the transcriptome of the target cell in the slice to be predicted, density statistics and functional analysis are performed on the functional units in the slice to be predicted.

10. A functional unit prediction model construction device, characterized in that: The device comprises: A data set acquisition module for acquiring a data set of at least one set of sample slices containing functional units, the data set acquisition module comprising: a first cell classification module and a second cell classification module, the first cell classification module being configured to acquire spatial transcriptome data containing a first cell type, the acquisition method comprising: acquiring spatial transcriptome data of a first sample slice, and performing cell classification on a plurality of cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of the plurality of cells in the first sample slice; the spatial transcriptome data comprising expression levels and spatial distribution information of transcriptomes of the plurality of cells in the first sample slice; the second cell classification module being configured to acquire a second cell type of cells associated with the functional unit and first inter-cell distance data, the acquisition method comprising: acquiring staining data of a second sample slice, and determining the second cell type of cells associated with the functional unit in the second sample slice and first inter-cell distance data based on the staining data, the first inter-cell distance data being configured to describe the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or an adjacent slice of the first sample slice; A model building module is used to train a machine learning module based on a data set of at least one group of sample slices containing functional units to build a functional unit prediction model.

11. A functional unit prediction device, characterized in that: The device comprises: a third cell classification module, configured to obtain spatial transcriptome data of the slice to be predicted, and perform cell classification on a plurality of cells in the slice to be predicted based on the spatial transcriptome data of the slice to be predicted, to obtain a third cell type of the plurality of cells in the slice to be predicted, wherein the spatial transcriptome data of the slice to be predicted includes expression levels and spatial distribution information of transcriptomes of the plurality of cells in the slice to be predicted; A unit prediction module is used to predict functional units based on the spatial transcriptome data containing the third cell type of the slice to be predicted, using a functional unit prediction model constructed according to the functional unit prediction model construction method according to any one of claims 1 to 6.

12. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the function unit prediction model construction method described in any one of claims 1 to 6 and the function unit prediction method described in any one of claims 7 to 9 are implemented.

13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for constructing a functional unit prediction model according to any one of claims 1 to 6 and the method for predicting a functional unit according to any one of claims 7 to 9 are implemented.

14. A computer program product, comprising a computer program, wherein the computer program is read and executed by a processor of a computer device, so that the computer device executes the functional unit prediction model construction method described in any one of claims 1 to 6 and the functional unit prediction method described in any one of claims 7 to 9.

Citation Information

Patent Citations

  • Spatial transcriptome cell clustering and analyzing method

    CN114091603A

  • Analysis method and system for integrating single cell transcriptome and spatial transcriptome data

    CN114944193A

  • Spatial transcriptome analysis method and device based on deep learning and readable medium

    CN116994245A

  • Method and device for determining cell type based on space transcriptome deconvolution of GraphSAGE

    CN118098356A

  • Three-level lymph structure identification and classification method and application thereof in cancer treatment

    CN118314965A