Function unit prediction model construction method and prediction method, device, and electronic equipment

By constructing a functional unit prediction model and training a machine learning module using spatial transcriptome data and staining data, the problem of insufficient spatial resolution in existing technologies is solved, enabling accurate prediction and functional analysis of functional units and supporting disease research.

CN120431990BActive Publication Date: 2025-11-04SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510927257.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-11-04
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately predict the spatial structure and interrelationships between cells in functional units with spatial structures. This is especially true in biological tissues or complex systems, where anatomical methods and biomedical imaging techniques lack spatial resolution and cannot fully reveal the dynamic changes in gene expression among cells of functional units.

Method used

By acquiring spatial transcriptomic and staining data from sample slices containing functional units, a machine learning module is trained to build a functional unit prediction model. Combining single-cell transcriptomic sequencing technology and spatial transcriptomics, cell type and distance data are analyzed to accurately predict the spatial distribution and composition of functional units.

Benefits of technology

It improves the accuracy of predicting functional units with spatial structures, and can clearly display the fine structure and function of functional units at the cellular level, supporting in-depth research on disease mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431990B_ABST
    Figure CN120431990B_ABST
Patent Text Reader

Abstract

The present disclosure provides a functional unit prediction model construction method and prediction method, device and electronic equipment. The functional unit prediction method comprises: obtaining at least one set of sample slice data set containing functional units, and training a machine learning module based on at least one set of sample slice data set to construct a functional unit prediction model. Each set of sample slice data set containing functional units comprises: spatial transcriptome data containing a first cell type, i.e., performing cell classification on a plurality of cells in a first sample slice based on spatial transcriptome data of the first sample slice to obtain a first cell type of the plurality of cells in the first sample slice; and second cell type of functional unit related cells and intercellular distance data, i.e., determining the second cell type of the functional unit related cells and the first intercellular distance data in the second sample slice based on the staining data of the second sample slice. The embodiments of the present application can improve the prediction accuracy of functional units with spatial structure.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of functional unit prediction, in particular to a functional unit prediction model construction method and prediction method, device and electronic equipment. BACKGROUND

[0002] The functional unit with spatial structure refers to a cell group, a molecular network or a tissue structure with a specific spatial position and performing a specific function in a biological tissue or a complex system. For example, when the functional unit is a neurovascular unit (NVU), since the NVU is a functional unit composed of cells of multiple cell types closely cooperating with each other, a complete functional analysis of the NVU is crucial for normal cognitive function and nervous system health. Therefore, accurately predicting the functional unit can help researchers more deeply analyze the gene expression characteristics of the functional unit, thereby assisting in understanding the occurrence and development process of related diseases.

[0003] The related art usually relies on anatomical methods and biomedical imaging techniques to mark the functional unit with spatial structure known to have spatial structure and perform partial spatial analysis, but these methods are difficult to accurately predict the spatial structure and mutual relationship between cells in the functional unit with spatial structure. Therefore, how to provide a prediction method suitable for the functional unit with spatial structure has become a technical problem to be solved. SUMMARY

[0004] The main purpose of the embodiments of the present application is to provide a functional unit prediction model construction method and prediction method, device and electronic equipment which can be suitable for the prediction of the functional unit with spatial structure.

[0005] To achieve the above-mentioned purpose, a first aspect of the embodiments of the present application provides a functional unit prediction model construction method, which comprises:

[0006] Step S1, acquiring at least one set of sample slice data containing a functional unit, each set of sample slice data containing a functional unit comprising:

[0007] Step S1.1, spatial transcriptome data containing a first cell type, the acquisition method comprising: acquiring spatial transcriptome data of a first sample slice, and performing cell classification on a plurality of cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of the plurality of cells in the first sample slice; the spatial transcriptome data comprising expression level and spatial distribution information of the transcriptome of the plurality of cells in the first sample slice; and

[0008] Step S1.2, obtaining the second cell type of the functional unit related cell and the first intercellular distance data, the method comprising: obtaining staining data of a second sample slice, determining the second cell type of the functional unit related cell and the first intercellular distance data in the second sample slice based on the staining data, the first intercellular distance data being used to describe the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or an adjacent slice of the first sample slice;

[0009] Step S2, training the machine learning module based on the data set of the at least one functional unit containing sample slice, and constructing a functional unit prediction model.

[0010] In some embodiments, before the cell classification of the plurality of cells in the first sample slice based on the spatial transcriptome data in step S1, the functional unit prediction model construction method further comprises: cell segmentation of the first sample slice, and the cell segmentation is based on the spatial transcriptome data of the first sample slice, and / or image data obtained by cell nucleus and / or membrane staining, such as ssDNA staining image, etc.

[0011] In some embodiments, the machine learning module comprises at least one of a supervised learning module, a semi-supervised learning module, an unsupervised learning module, a regression analysis module, a reinforcement learning module, a self-learning module, a feature learning module, a sparse dictionary learning module, an anomaly detection module, a generative adversarial network or an association rule module.

[0012] In some embodiments, the machine learning module comprises at least one of a supervised learning module, a semi-supervised learning module, an unsupervised learning module, a regression analysis module, a reinforcement learning module, a self-learning module, a feature learning module, a sparse dictionary learning module, an anomaly detection module, a generative adversarial network or an association rule module.

[0013] According to the spatial transcriptome data containing the first cell type, the spatial adjacency data of the first sample slice is constructed, and the spatial adjacency data is used to describe the adjacency structure between the same or different cells in the first sample slice;

[0014] According to the matching result of the second cell type of the functional unit related cell and the first cell type, and the first intercellular distance data, the unit spatial data related to the functional unit is extracted from the spatial adjacency data;

[0015] According to the spatial structure indicated by the unit spatial data, the center cell type, the associated cell type, and the distance threshold between the center cells, the associated cells and between the center cells and the associated cells are determined from the first cell type; the center cell refers to the cell constituting the functional unit, and the associated cell refers to the non-center cell having intercellular communication with the center cell.

[0016] training a machine learning module according to the central cell type, the associated cell type, and the distance threshold.

[0017] In some embodiments, the cell classification of the plurality of cells in the first sample slice based on the spatial transcriptome data comprises:

[0018] obtaining single-cell transcriptome sequencing data of a plurality of cells of the sample from which the first sample slice is derived, and determining common gene expression data of each cell in the first sample slice based on expression levels of the transcriptome of the plurality of cells in the first sample slice and the single-cell transcriptome sequencing data;

[0019] performing cell annotation on the plurality of cells in the first sample slice based on the common gene expression data and preset cell marker genes, to obtain first cell types of the plurality of cells in the first sample slice.

[0020] In some embodiments, the determination of the common gene expression data of each cell in the first sample slice based on the expression levels of the transcriptome of the plurality of cells in the first sample slice and the single-cell transcriptome sequencing data comprises:

[0021] determining a first gene expression matrix of the plurality of cells in the first sample slice based on the expression levels of the transcriptome of the plurality of cells in the first sample slice;

[0022] determining a second gene expression matrix of the plurality of cells in the first sample slice based on the single-cell transcriptome sequencing data;

[0023] determining the common gene expression data of each cell in the first sample slice according to common genes of the plurality of cells in the first gene expression matrix and the second gene expression matrix.

[0024] In some embodiments, the functional unit prediction of the third sample slice by the functional unit prediction model is compared with staining data of a fourth sample slice to determine an accuracy level of the machine learning module, wherein the fourth sample slice is the third sample slice or an adjacent slice of the third sample slice.

[0025] To achieve the above object, a second aspect of the embodiments of the present application provides a functional unit prediction method, which comprises:

[0026] obtain spatial transcriptome data of a to-be-predicted slice, and perform cell classification on a plurality of cells in the to-be-predicted slice based on the spatial transcriptome data of the to-be-predicted slice, to obtain a third cell type of the plurality of cells in the to-be-predicted slice, wherein the spatial transcriptome data of the to-be-predicted slice includes expression level and spatial distribution information of a transcriptome of the plurality of cells in the to-be-predicted slice;

[0027] The functional unit prediction model constructed by the functional unit prediction model construction method according to the embodiments of the present application is used for functional unit prediction based on the spatial transcriptome data of the to-be-predicted slice containing the third cell type.

[0028] In some embodiments, the functional unit prediction model constructed according to the first aspect of the embodiments of the present application is used for functional unit prediction based on the spatial transcriptome data of the to-be-predicted slice containing the third cell type, including:

[0029] The functional unit prediction model constructed by the functional unit prediction model construction method according to the embodiments of the present application is used for determining a central cell type, an associated cell type, and a distance threshold between the central cells, between the associated cells, and between the central cells and the associated cells;

[0030] Based on a matching result of the central cell type and the third cell type, a central cell of a functional unit is determined from the plurality of cells of the to-be-predicted slice;

[0031] Based on a matching result of the associated cell type and the third cell type, an associated cell of a functional unit is determined from the plurality of cells of the to-be-predicted slice;

[0032] Second cell distance data between the central cells, between the associated cells, and between the central cells and the associated cells in the to-be-predicted slice is obtained according to the spatial transcriptome data, and the second cell distance data is used for describing relative distances between the central cells, between the associated cells, and between the central cells and the associated cells in the to-be-predicted slice;

[0033] According to a comparison result of the relative distances and the distance threshold, a target cell related to a functional unit to which the central cell belongs is determined from the plurality of cells of the to-be-predicted slice; the target cell includes the central cell and the associated cell.

[0034] According to the target cell and spatial distribution information of the target cell in the to-be-predicted slice, a functional unit of the to-be-predicted slice is predicted.

[0035] In some embodiments, after the functional unit prediction model constructed by the functional unit prediction model construction method according to the embodiments of the present application is used to predict the functional unit of the third cell type based on the spatial transcriptome data of the to-be-predicted slice containing the third cell type, the functional unit prediction method further comprises:

[0036] extracting the expression level and spatial distribution information of the transcriptome of the target cell in the to-be-predicted slice from the spatial transcriptome data of the to-be-predicted slice;

[0037] performing density statistics and functional analysis on the functional unit in the to-be-predicted slice according to the expression level and spatial distribution information of the transcriptome of the target cell in the to-be-predicted slice.

[0038] To achieve the above-mentioned purpose, a third aspect of the embodiments of the present application proposes a functional unit prediction model construction device, which comprises:

[0039] a data set acquisition module for acquiring at least one set of data set of sample slice containing functional unit, the data set acquisition module comprises: a first cell classification module and a second cell classification module, the first cell classification module is used for acquiring spatial transcriptome data containing first cell type, the acquisition method comprises: acquiring spatial transcriptome data of the first sample slice, and performing cell classification on a plurality of cells in the first sample slice based on the spatial transcriptome data, to obtain the first cell type of the plurality of cells in the first sample slice; the spatial transcriptome data comprises the expression level and spatial distribution information of the transcriptome of a plurality of cells in the first sample slice; the second cell classification module is used for acquiring the second cell type of the functional unit related cell and the first cell distance data, and the acquisition method comprises: acquiring the staining data of the second sample slice, determining the second cell type of the functional unit related cell and the first cell distance data in the second sample slice based on the staining data, the first cell distance data is used to describe the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or the adjacent slice of the first sample slice;

[0040] a model construction module for training a machine learning module based on the at least one set of data set of sample slice containing functional unit, and constructing a functional unit prediction model.

[0041] To achieve the above-mentioned purpose, a fourth aspect of the embodiments of the present application proposes a functional unit prediction device, which comprises:

[0042] a third cell classification module, configured to obtain spatial transcriptome data of a to-be-predicted slice, and perform cell classification on a plurality of cells in the to-be-predicted slice based on the spatial transcriptome data of the to-be-predicted slice, to obtain third cell types of the plurality of cells in the to-be-predicted slice, wherein the spatial transcriptome data of the to-be-predicted slice comprises expression levels and spatial distribution information of transcriptomes of the plurality of cells in the to-be-predicted slice;

[0043] a unit prediction module, configured to perform unit prediction based on the spatial transcriptome data of the to-be-predicted slice containing the third cell types according to a functional unit prediction model constructed by the functional unit prediction model construction method.

[0044] To achieve the above object, a fifth aspect of embodiments of the present application provides an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the functional unit prediction model construction method of the first aspect and the functional unit prediction method of the second aspect when executing the computer program.

[0045] To achieve the above object, a sixth aspect of embodiments of the present application provides a computer readable storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to implement the functional unit prediction model construction method of the first aspect and the functional unit prediction method of the second aspect.

[0046] To achieve the above object, a seventh aspect of embodiments of the present application provides a computer program product, the computer program product comprises a computer program, the computer program is read and executed by a processor of a computer device, so that the computer device executes the functional unit prediction model construction method of the first aspect and the functional unit prediction method of the second aspect.

[0047] The function unit prediction model construction method and the function unit prediction method, the device and the electronic equipment provided by the embodiments of the present application can be used to construct a function unit prediction model for predicting a function unit in a slice. Before the function unit prediction model is constructed, at least one set of data of a sample slice containing a function unit is acquired. Each set of data of the sample slice containing the function unit includes spatial transcriptome data containing a first cell type, a second cell type of a function unit related cell, and first cell distance data. The method for acquiring the spatial transcriptome data containing the first cell type includes acquiring spatial transcriptome data of a first sample slice, performing cell classification on a plurality of cells in the first sample slice based on the spatial transcriptome data, and obtaining the first cell type of the plurality of cells in the first sample slice. The spatial transcriptome data includes expression level and spatial distribution information of the transcriptome of the plurality of cells in the first sample slice. The method for acquiring the second cell type of the function unit related cell and the cell distance data includes acquiring staining data of a second sample slice, determining the second cell type of the function unit related cell and the first cell distance data in the second sample slice based on the staining data, and using the first cell distance data to describe the relative distance between the same or different cells in the second sample slice. The second sample slice is the first sample slice or an adjacent slice of the first sample slice. Then, a machine learning module is trained based on the at least one set of data of the sample slice containing the function unit, and the function unit prediction model is constructed. Since the distribution of the function unit in the first sample slice and the second sample slice has similarity, the spatial transcriptome data containing the first cell type and the second cell type of the function unit related cell and the cell distance data can be used to analyze the type distribution and the spatial distribution between the cells contained in the function unit, so as to construct the function unit prediction model capable of predicting the function unit. In this way, the embodiments of the present application can be better applied to the prediction of the function unit with a spatial structure, and the prediction accuracy of the function unit is improved. BRIEF DESCRIPTION OF DRAWINGS

[0048] The accompanying drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification, and are used to explain the technical solutions of the present disclosure together with the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions of the present disclosure.

[0049] Figure 1 is a flowchart of a function unit prediction model construction method provided by the embodiments of the present application;

[0050] Figure 2 is a spatial distribution diagram of a hippocampal brain region cell type provided by the embodiments of the present application;

[0051] Figure 3 is an immunofluorescence diagram of a hippocampal brain region provided by the embodiments of the present application;

[0052] Figure 4 is a schematic diagram of constructing a functional unit prediction model based on machine learning provided by an embodiment of the present application;

[0053] Figure 5 is a schematic diagram of an NVU model prediction convergence curve provided by an embodiment of the present application;

[0054] Figure 6 is a spatial distribution diagram of a neural vascular unit functional unit provided by an embodiment of the present application;

[0055] Figure 7 is a schematic diagram of the density change of a neural vascular unit functional unit provided by an embodiment of the present application;

[0056] Figure 8 is a schematic diagram of the response of a neural vascular unit functional unit to a hypoxic pathway provided by an embodiment of the present application;

[0057] Figure 9 is a schematic diagram of the immune response regulation of a neural vascular unit functional unit provided by an embodiment of the present application;

[0058] Figure 10 is a specific embodiment flowchart of a functional unit prediction model construction method provided by an embodiment of the present application;

[0059] Figure 11 is a schematic diagram of a functional unit prediction device provided by an embodiment of the present application;

[0060] Figure 12 is a schematic diagram of a functional unit prediction device provided by an embodiment of the present application;

[0061] Figure 13 is a hardware structure schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0062] In order to make the purposes, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and do not limit the present disclosure.

[0063] Before the present disclosure is further described in detail, the terms and phrases involved in the embodiments of the present disclosure are explained, and the terms and phrases involved in the embodiments of the present disclosure are applicable to the following explanations:

[0064] Spatial Transcriptomics (ST) sequencing technology: a technology for analyzing RNA from a spatial perspective, mainly used to analyze the spatial position of RNA in a single tissue section. It uses Unique Molecular Identifiers (UMI) to identify copies of the same RNA molecule, and can also detect the expression level of RNA at the same time. Stereo-seq (Spatial Enhanced Resolution Omics-sequencing) is an advanced spatial transcriptomics technology. Stereo-seq is based on DNA Nano Ball (DNB) and is a high-throughput, ultra-high-resolution, large-field in situ panoramic technology. Stereo-seq can achieve spatial transcriptome analysis of the same sample at the tissue, cell, subcellular, and molecular "four scales" simultaneously. Stereo-seq captures mRNA in tissues through a spatiotemporal chip and restores the spatial position through spatial barcodes (Spatial barcode or Coordinate ID, CID), achieving spatial expression detection of genes in tissues and providing a strong foundation for in-depth understanding of the relationship between cell gene expression and local environment.

[0065] Pathological staining images: By using specific staining agents to enhance the contrast of specific structures or components in the tissue sample, the morphology and pathological changes of the tissue cells can be observed and analyzed more clearly under a microscope. Common staining techniques include HE staining and special staining (such as mIF staining). Among them, HE staining refers to Hematoxylin and Eosin staining, which can clearly show the different structures of cell nuclei and cytoplasm. Special staining can refer to the use of special staining techniques for specific purposes. Special staining can include Masson's trichrome staining, PAS staining, silver staining, etc., which can highlight specific tissue structures or components such as collagen fibers, muscle tissue, and nerve tissue.

[0066] Artificial Intelligence (AI): is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; Artificial intelligence is a branch of computer science, artificial intelligence aims to understand the essence of intelligence, and produce a new intelligent machine that can react in a similar way to human intelligence, the research in this field includes robots, language recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computer or digital computer controlled machine to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results.

[0067] Spatial transcriptomics is an emerging technology that combines imaging techniques and gene expression analysis, enabling precise localization of gene expression in the spatial structure of cells or tissues. When spatial transcriptomics is associated and combined with multi-modal information, it can reveal the behavior and function of cells in a specific biological structure, providing a research basis for understanding complex biological processes such as disease development, tissue formation and immune response. Through spatial transcriptomics, the distribution of different cell types in tissues and the interaction between different cell types can be observed, as well as other functions or characteristics.

[0068] Spatial resolution refers to the minimum distance between two adjacent objects that an imaging system can distinguish, and is mainly used to describe the ability of an image to distinguish details in the spatial dimension (such as length and width direction for two-dimensional images; or length, width and height direction for three-dimensional images). For example, in CT imaging, high spatial resolution means that the fine structure of bones, the boundary of soft tissues, etc. can be distinguished more clearly; if the spatial resolution is low, the details in the image will be blurred, and adjacent structures may be fused and difficult to distinguish. Taking brain tissue as an example, high spatial resolution imaging technology can help researchers clearly see the distribution of neuron cell bodies, the direction of nerve fibers and their connection relationship, so as to better understand the neural circuit and function of the brain.

[0069] Functional units with spatial structure refer to cell groups, molecular networks or tissue structures with specific spatial positions and specific functions in biological tissues or complex systems. For example, when the functional unit is a neurovascular unit (NVU), since the NVU is a functional unit composed of cells of multiple cell types that work closely together, it is responsible for maintaining the homeostasis of the brain, ensuring that neurons obtain sufficient oxygen and nutrients, and removing metabolic waste, so a complete functional analysis of the NVU is crucial for normal cognitive function and nervous system health. These units not only depend on their constituent components (such as specific cells, proteins, genes), but also highly depend on their arrangement and interaction in space to realize the function. Taking neural tissue as an example, neurons and glial cells form a complex neural network according to a specific spatial distribution and connection mode, and realize the regulation and control of various physiological activities of the animal body through the transmission of electrical signals and chemical signals. Different cell arrangement modes reflect the specific functional needs of the tissue, and the signal transmission and regulation in the tissue structure also depend on the spatial position and mutual relationship of the cells. For example, in neural tissue, the complex structure of axons and dendrites allows long-distance transmission and synchronous processing of neural signals. Functional units with spatial structure have wide application significance in the study of disease mechanisms and precision treatment, such as revealing how the spatial distribution of amyloid plaques in Alzheimer's disease affects adjacent neurons, and locating drug-resistant cell enrichment areas in tumors, thereby guiding local targeted therapy. Therefore, accurate prediction of functional units can help researchers more deeply analyze the gene expression characteristics of functional units, thereby assisting in understanding the occurrence and development process of related diseases.

[0070] However, current research on the cell composition and spatial structure of functional units with spatial structure usually relies on anatomical methods and biomedical imaging techniques to label functional units known to have spatial structure and perform partial spatial analysis, but these methods are difficult to accurately predict the spatial structure and mutual relationship between cells in functional units with spatial structure (i.e., cannot effectively predict the complex gene expression pattern and spatial distribution of functional units with spatial structure). Among them, anatomical methods rely on tissue sectioning, immunohistochemical staining, microscopic observation and other means to label and observe the structure of functional units with spatial structure. These methods usually require rich professional experience and may have subjectivity in tissue section processing and image analysis. Due to the lack of spatial resolution, it is difficult to reveal the complete structure of functional units with spatial structure and the relationship between cells at the cellular level. Immunohistochemical staining combined with microscopy is a commonly used method for current research on functional units with spatial structure. By staining to label specific cell types (such as neural vascular units including vascular endothelial cells, astrocytes and neurons, etc.), researchers can observe the spatial distribution and interaction of cells. However, this method requires higher processing requirements for tissue samples, and when analyzing complex three-dimensional spatial structures, it may be limited by section thickness and labeling accuracy. In addition, related technologies also provide cell composition and gene expression information of functional units with spatial structure through gene expression analysis methods (such as single-cell RNA sequencing). However, the gene expression technology used by related technologies usually lacks spatial resolution and cannot reveal the precise location and interaction of cells in the tissue.

[0071] For example, in the study of brain neural tissue, the gene expression of neural cells changes dynamically under different physiological and pathological conditions, and these changes are closely related to the interaction between cells. The method used by related technologies can only observe the static state of cells at a certain time point, and it is difficult to capture the dynamic changes of gene expression over time and cell interaction, that is, it is difficult to comprehensively and deeply reveal the dynamic changes of gene expression between cells within functional units with spatial structure. Therefore, how to provide a prediction method suitable for functional units with spatial structure has become a technical problem to be solved.

[0072] Therefore, based on this, the embodiment of the present application provides a functional unit prediction model construction method and a prediction method, device and electronic equipment, which can be better applied to the prediction of functional units with spatial structure and improve the prediction accuracy of functional units with spatial structure.

[0073] The functional unit prediction model construction method provided by the embodiment of the present application is described below.

[0074] Reference Figure 1In some embodiments, the function unit prediction model construction method provided by the embodiments of the present application includes but is not limited to steps S1 to S2.

[0075] In step S1, at least one set of sample slice data containing a function unit is obtained, and each set of sample slice data containing a function unit includes:

[0076] In step S1.1, spatial transcriptome data containing a first cell type is obtained, and the method includes: obtaining spatial transcriptome data of a first sample slice, and performing cell classification on a plurality of cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of the plurality of cells in the first sample slice; the spatial transcriptome data includes expression level and spatial distribution information of the transcriptome of the plurality of cells in the first sample slice; and

[0077] In step S1.2, a second cell type of a function unit related cell and intercellular distance data are obtained, and the method includes: obtaining staining data of a second sample slice, determining the second cell type of the function unit related cell and first intercellular distance data based on the staining data, and the first intercellular distance data is used to describe the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or an adjacent slice of the first sample slice;

[0078] In step S2, a machine learning module is trained based on at least one set of sample slice data containing a function unit to construct a function unit prediction model.

[0079] The steps S1 to S2 shown in the embodiments of the present application can reveal the distance and spatial distribution characteristics between cells at the single cell level by introducing spatial transcriptome data, and can provide researchers with spatial coordinates of cells and their interactions in combination with other technical features, solve the problem of insufficient spatial resolution in the related art, and clearly display the fine structure of the function unit with spatial structure at the cell level, thereby accurately analyzing the composition and function of the function unit with spatial structure. Specifically, the function unit prediction model capable of predicting the cell type distribution and spatial distribution of the function unit is constructed based on the cell type of the plurality of cells in the first sample slice and the cell type and intercellular distance data of the plurality of cells in the second sample slice, which can be better applied to the prediction of the function unit with spatial structure and improve the prediction accuracy of the function unit.

[0080] Steps S1.1 and S1.2 can be performed in any order.

[0081] It should be noted that spatial transcriptome technology can break through these limitations, providing more accurate data at the cellular and tissue levels, helping researchers in neuroscience, oncology, immunology, and other fields to analyze the functional units with spatial structure of cells or tissues, and identify their gene expression characteristics, thereby providing a new perspective for the study of disease mechanisms.

[0082] In step S1 of some embodiments, the data set of at least one group of sample slices containing functional units is used to train the subsequent machine learning module to construct a functional unit prediction model for predicting functions. Wherein, the process of obtaining the data set of each group of sample slices containing functional units can include: step S1.1, obtaining spatial transcriptome data containing a first cell type; and step S1.2, obtaining a second cell type of functional unit related cells and intercellular distance data.

[0083] In step S1.1 of some embodiments, when obtaining spatial transcriptome data containing a first cell type, the first sample slice therein can refer to a tissue slice to be subjected to spatial transcriptome sequencing. The first sample slice is a biological slice obtained in accordance with relevant legal regulations, including plant slices and animal slices. The first sample slice refers to a specific tissue slice selected during the research process, which contains the target cell population that needs to be studied (such as the first sample slice tested which can contain 50,000-100,000 cells, not limited). It is the basis sample for subsequent analysis. The spatial transcriptome data of the first sample slice refers to the data obtained by spatial transcriptome technology about the expression level and spatial distribution information of the transcriptome of multiple cells in the first sample slice. Spatial transcriptome technology can simultaneously obtain spatial information of gene expression in tissue or cell samples. These data can reflect the specific location of cells in the tissue and their gene expression characteristics. Cell classification refers to the process of dividing multiple cells in the first sample slice into different cell types or cell subpopulations according to their gene expression characteristics, spatial location, etc. Cell type refers to an identifier assigned to each cell in the first sample slice, indicating the specific cell type or cell subpopulation to which the cell belongs, such as vascular cells (endothelial cells and pericytes), astrocytes, neurons (inhibitory neurons and excitatory neurons), etc.

[0084] It should be noted that, in order to supplement the present disclosure, the specific acquisition method of the spatial transcriptome data of the slice (such as the sample slice, the to-be-predicted slice, etc.) in the present application can adopt the patent disclosure of the spatial transcriptome data acquisition method in the following patent applications by reference: European Patent Publication No. EP2697391A1, International Patent Application Publication No. WO2024086167A3, and WO2023115536A1. The structure and acquisition method of all the slices mentioned in the embodiments of the present application can refer to the first sample slice described above, and will not be repeated here.

[0085] It should be noted that the spatial transcriptome data of the first sample slice includes spatial position data (i.e., spatial distribution information of the transcriptome in the first sample slice) and spatial gene expression data (i.e., expression level of the transcriptome in the first sample slice) of a plurality of cells in the first sample slice. Among them, the spatial transcriptome data of the present application is spatial transcriptome data with single-cell resolution, which can realize accurate prediction of functional units with spatial structure, such as spatial transcriptome data determined by Stereo-seq method. The spatial position data refers to the specific position information of the plurality of cells in the first sample slice in the tissue. In the study of spatial transcriptomics, researchers can divide the tissue slice into multiple small regions, each region is called a "spot", and each spot represents a small piece of tissue (for example, a cell of the tissue slice). The cells in this region are analyzed to determine their gene expression patterns. The spatial position of these cells can be described by coordinates, such as (x, y) coordinates in a two-dimensional plane or (x, y, z) coordinates in a three-dimensional space. These coordinate information can clearly indicate the relative or absolute position of each cell in the first sample slice. The spatial gene expression data refers to the specific situation of gene expression in the plurality of cells in the first sample slice. In practical applications, spatial transcriptome data refers to the expression data of all genes in cells or tissues at a specific condition and time point. Spatial gene expression data reflects the activity level of each gene in each cell in the first sample slice at the transcription level, which genes are expressed and how much they are expressed. These data can be presented in the form of a matrix, where the rows represent genes, the columns represent cells (or spots), and the elements in the matrix represent the expression level of a certain gene in a certain cell.

[0086] In some embodiments, before the cell classification of the plurality of cells in the first sample slice based on the spatial transcriptome data in step S1, the functional unit prediction model construction method further comprises: performing cell segmentation on the first sample slice, and the basis for cell segmentation is the spatial transcriptome data of the first sample slice, and / or image data obtained by performing nucleus and / or membrane staining on the first sample slice, such as ssDNA staining images, etc.

[0087] In some embodiments, the process of screening the first sample slice specifically comprises:

[0088] spatial transcriptome data of a plurality of sample slices are obtained, and the spatial transcriptome data of each sample slice includes expression level and spatial distribution information of the transcriptome of each cell in the sample slice;

[0089] data extraction is performed based on the spatial transcriptome data to obtain the gene number and mitochondrial gene proportion of each cell in each sample slice;

[0090] The plurality of sample slices are screened based on the gene number and mitochondrial gene proportion of each cell, and a first sample slice is determined.

[0091] In the present application, a plurality of initial sample slices can be screened first, that is, quality control is performed on the initial sample slices, low-quality cells and slices are removed, and the initial sample slices meeting the quality control standard are used as the first sample slices for constructing the functional unit prediction model for predicting the functional unit with spatial structure. Specifically, spatial sequencing data and corresponding spatial transcriptome data of a plurality of spatial sites in each initial sample slice are obtained, and the spatial transcriptome data includes spatial gene expression data of a plurality of cells in the initial sample slice (that is, expression level of the transcriptome of the plurality of cells in the initial sample slice) and spatial position data (that is, spatial distribution information of the transcriptome of the plurality of cells in the initial sample slice).

[0092] The spatial sequencing data refers to the data about the expression and spatial position of the transcriptome obtained by performing spatial transcriptome sequencing technology (such as Visium, Visium HD, Visium HD3', Stereo-seq, etc.) on each initial sample slice. For example, Visium uses spatial probes including spatial barcodes (each position on the chip has a unique spatial barcode sequence), UMI, and capture sequences (poly T) on the chip to capture mRNA in the tissue. The spatial probes are extended with mRNA as a template, and the obtained extension product includes sequence information of mRNA, corresponding UMI information, and spatial position information (spatial barcode sequence information) of mRNA. The obtained extension product is sequenced (such as bridge sequencing, DNB sequencing) to decode these information. Each position (i.e., spatial site) is also called a spot. Generally, a spot has a spatial probe cluster with the same spatial barcode fixed thereon. UMI is used to distinguish different mRNA molecule sources to avoid the repeated counting problem caused by polymerase chain reaction (PCR) amplification.

[0093] Further, after determining the spatial transcriptome data of the plurality of sample slices, a spatial density image of the initial sample slice can be generated based on the spatial sequencing data and the spatial coordinates of each spatial locus in the initial sample slice. Wherein, the present application can aggregate the total UMI count of each spot point in the spatial coordinates of the sample slice based on the spatial sequencing data (specifically, the total UMI count of each spot point can be obtained by identifying and counting different UMI sequences on each spot point. This process can be completed using special bioinformatics software and algorithms, which are not limited), and generate a spatial density matrix of each sample slice based on the total UMI count and the spatial coordinates of each spatial locus. Then, the spatial density matrix is converted into an image to obtain the spatial density image, and the gray intensity in the image is used to reflect the number of UMIs.

[0094] It should be noted that for the spatial density matrix, the present application can associate the spatial coordinates of each spot point with the corresponding total UMI count, i.e., one-to-one correspondence between these spatial coordinates and the corresponding total UMI count, forming a data set. Further, according to the spatial layout of the sample slice and the distribution of the spot points, the above data set is organized into a matrix. The rows and columns of the matrix correspond to the spatial positions on the initial sample slice (such as the rows correspond to the row coordinates of the sample slice, and the columns correspond to the column coordinates of the sample slice), and each element in the matrix represents the total UMI count corresponding to the spot point at the position. In this way, the spatial density matrix is generated, and the element value in the matrix reflects the gene expression density of the corresponding position.

[0095] Further, the present application can perform cell segmentation on the spatial density image based on the spatial distribution information of the transcriptome of the plurality of cells in the initial sample slice in the initial sample slice, and determine the cell region of each cell in the initial sample slice in the spatial density image. Specifically, the present application can preset a cell segmentation model (such as ESPANet model) for cell segmentation, and perform post-processing on the initial spatial gene expression data of the initial sample slice through a watershed algorithm, and finally generate a cell-gene matrix for creating a Seurat RDS object of cellbin (a cell unit, in the field of spatial transcript technology, bin can refer to the process of merging adjacent or similar cells in spatial positions into a unit for analysis).

[0096] The present application can also perform cell segmentation based on the optical image of the initial sample slice, for example, using H&E staining image or mIF image for cell segmentation.

[0097] It should be noted that the watershed algorithm is a morphological-based image segmentation algorithm, which can regard the image as a terrain with peaks and valleys, and segment the image by flooding seed regions and gradually expanding them. In cell segmentation, the watershed algorithm can utilize the boundary information of the cells to further refine the segmentation results of the ESPANet model, solve the problem of cell adhesion, optimize the segmentation results, and thus more accurately determine the cell region of each cell in the spatial density image in the initial sample slice.

[0098] Further, based on the cell region, data extraction is performed on the initial spatial gene expression data to obtain the gene number (nFeatures), the proportion of mitochondrial genes (percent.MT), and the total number of gene expression amounts of each cell in the initial sample slice. Wherein, nFeatures refers to the gene number of each cell, i.e. the number of different genes detected in each cell, and the nFeatures value of each cell can be obtained by counting the number of non-zero elements in each row of the cell-gene matrix. nCounts represents the total number of expression amounts of all genes in each cell, i.e. the total number of UMIs detected in each cell. The nCounts value of each cell can be obtained by summing the elements in each row of the cell-gene matrix. Percent.mt refers to the proportion of mitochondrial genes in each cell, i.e. first determine which genes are mitochondrial genes, then count the number of mitochondrial genes in each cell, divide by the nFeatures value of the cell, and finally multiply by 100 to obtain the percent.mt value. In this way, the quality control standard of the present application can include: at least 50% of the cells in a tissue slice need to meet the following requirements at the same time: (1) the number of genes detected in the cell (nFeature) is greater than 100; and (2) the proportion of mitochondrial genes (percent.MT) in the cell is less than 20%. That is, when at least 50% of the cells in the initial sample slice meet the preset quality control standard, it can be considered as qualified, and then the initial sample slice can be regarded as the first sample slice, and the quality control standard can be flexibly set according to needs, but at least 50% of the cells need to meet the quality control standard at the same time.

[0099] In the above embodiments, the application can obtain the expression profile of the genes from the spot points, analyze the data of idling, and determine the cell contour and circle the cells (i.e., cell segmentation) according to the spot gene enrichment or ssDNA staining or DAPI staining or any method known to those skilled in the art to segment cells, so as to achieve the effect of single cell resolution and obtain a plurality of cellbins. Specifically, the application can generate a cell-gene matrix corresponding to the cellbin according to the SAW process (Stereo-seq Analysis Workflow), and perform quality control on a plurality of initial sample slices according to the matrix to select the initial sample slices with better quality for subsequent processing, which can improve the prediction accuracy of functional units with spatial structure.

[0100] In some embodiments, the step of performing cell classification on the plurality of cells in the first sample slice based on the spatial transcriptome data to obtain the first cell type of the plurality of cells in the first sample slice can specifically include:

[0101] Using single-cell transcriptome sequencing technology, single-cell transcriptome sequencing data of the plurality of cells of the source sample in the first sample slice is obtained, and based on the expression level of the transcriptome of the plurality of cells in the first sample slice and the single-cell transcriptome sequencing data, common gene expression data of each cell in the first sample slice is determined.

[0102] Based on the common gene expression data and the preset cell marker gene, the plurality of cells in the first sample slice are cell annotated to obtain the first cell type of the plurality of cells in the first sample slice.

[0103] Wherein, the single-cell transcriptome sequencing data refers to the data determined by single-cell sequencing or single-nucleus sequencing (single nucleus RNA sequencing, snRNA-seq) method, which is equivalent to a reference data set, and the cell annotation is performed by referring to the reference data set. The common gene expression data refers to the intersection data of the expression level of the transcriptome of the plurality of cells in the first sample slice and the gene data in the single-cell transcriptome sequencing data.

[0104] Further, the application can perform cell classification on the plurality of cells in the first sample slice based on the common gene expression data and preset cell marker genes, to obtain a first cell type of each cell in the first sample slice. The preset cell marker genes refer to genes that are specifically expressed or highly expressed in a specific cell type. These genes can serve as markers for identifying cell types. The first cell type corresponds to a cell ID in the first sample slice, because each cell after classification is given an ID when it is circled, representing the identity of the cell (at the same time, the expression of all genes in this range also belongs to this cell). Further, the cell classification can add a column of annotations about the ID to the spatial expression matrix of the spatial transcriptome data (a matrix used to represent the expression levels of the transcriptomes of the plurality of cells in the first sample slice). For example, cell1: Neuron (neuron). In this way, after classifying each cell by algorithm, the application can map the obtained plurality of first cell types to the single-cell resolution spatial expression matrix, i.e., give each cell in the first sample slice a label according to the determined cell ID, without regenerating the matrix.

[0105] It can be understood that after determining the first sample slice that meets the quality, the application can use the Spatial-ID algorithm to classify the cells in the first sample slice that meets the quality, which combines the common genes of snRNA-seq data and Stereo-seq data (i.e., target spatial gene expression data) for cell type prediction.

[0106] In some embodiments, based on the expression levels of the transcriptomes of the plurality of cells in the first sample slice and the single-cell transcriptome sequencing data, the step of determining the common gene expression data of each cell in the first sample slice can specifically include:

[0107] Based on the expression levels of the transcriptomes of the plurality of cells in the first sample slice, a first gene expression matrix of the plurality of cells in the first sample slice is determined;

[0108] Based on the single-cell transcriptome sequencing data, a second gene expression matrix of the plurality of cells in the first sample slice is determined;

[0109] According to the common genes of the plurality of cells in the first gene expression matrix and the second gene expression matrix, the common gene expression data of each cell in the first sample slice is obtained.

[0110] In the application, the spatial gene expression data of the first sample slice (i.e., the expression levels of the transcriptomes of the plurality of cells in the first sample slice) and the gene expression matrix corresponding to the single-cell transcriptome sequencing data can be extracted first, and then the intersection of the two gene expression matrices is taken as the common gene expression data. The first gene expression matrix refers to a matrix constructed based on the expression levels of the transcriptomes of the plurality of cells in the first sample slice. The rows of the matrix usually represent genes, and the columns represent cells in the first sample slice. The elements in the matrix represent the expression levels of a certain gene in a certain cell, and are used to describe the gene expression of the cells in the first sample slice. The second gene expression matrix refers to a matrix constructed according to the single-cell transcriptome sequencing data. Similarly, the rows are genes, and the columns are cells. The elements represent the expression levels of the genes in the cells. This matrix does not contain the spatial position information of the cells. The common gene expression data refers to the gene expression information of each cell in the first sample slice that exists in both the first gene expression matrix and the second gene expression matrix. These genes are reflected in both data sources and can be used for comprehensive analysis of the gene expression characteristics of the cells.

[0111] In some embodiments, the application can use spatial neighborhood information and gene expression profiles to optimize cell type probabilities through a graph convolutional network (GCN) and self-supervised learning. Finally, the cell type of each cell is determined according to the highest probability generated by the GCN, and the identity of each cell in the spatial chip (i.e., the first cell type) is determined.

[0112] As shown in FIG. 1, the application provides a spatial distribution map of hippocampal brain cell types. Figure 2 As shown in FIG. 2, the spatial distribution maps of the control chip (here, the chip refers to the slice) and the disease chip after cell classification are shown. Figure 2 As shown in the left graph of FIG. 3, the spatial distribution of the cell annotation results of the control chip is shown. Figure 2 As shown in the right graph of FIG. 3, the spatial distribution of the cell annotation results of the disease chip is shown, and different colors represent different cell types. In this way, the application can distinguish cell types according to different colors. Figure 2

[0113] ​In step S1.2 of some embodiments, the second cell type is a cell type of cells in the second sample slice that are related to the functional unit. The first intercellular distance data are used to describe the relative distance between cells of the same cell type or different cell types in the spatial distribution in the second sample slice, which can reflect the difference in spatial distribution between cells of the same cell type or different cell types, such as adjacent cells, cells far apart, etc. The staining data refer to data obtained after specific staining of the second sample slice. The staining can mark specific structures or molecular characteristics of cells, which is helpful for subsequent analysis of intercellular distance and other information. The staining image can be a pathological staining image. The second sample slice can be the first sample slice or a neighboring slice of the first sample slice. The neighboring slice refers to a continuous slice that is close to the first sample slice in the three-dimensional structure of the tissue, without limitation. Therefore, by obtaining the staining data, the cell types and cell spatial distribution required for constructing the functional unit can be known, and the present application will be described in detail below with the neural vascular unit (a functional unit) as an example.

[0114] It should be noted that the present application will be described in detail below with the neural vascular unit (a functional unit) as an example. In order to enhance the robustness of the constructed functional unit prediction model in the analysis of the neural vascular unit, the present application can also implement a data enhancement strategy (such as image transformation of the staining data): under the premise of maintaining the cell attribution and biological structure constraints (such as the invariance of the neuron-vascular connection topology), geometric diversity is introduced by random rotation (such as rotation ± 45°), flipping, translation (such as moving a distance of ± 15% of the size), etc.; and / or, combined with optical adjustment (Gaussian blur simulates out-of-focus, Poisson noise simulates low-light imaging, channel disturbance adapts to staining differences) and biological specialization deformation (local occlusion avoids key synapses, elastic deformation protects the integrity of the unit).

[0115] Therefore, the present application can perform spatial analysis on the second sample slice based on immunostaining technology, analyze the spatial distribution between cells of the same or different cell types in the functional unit with spatial structure (such as astrocytes, vascular endothelial cells, neurons, etc. in the neural vascular unit). The relative distance between cells of the same or different cell types (such as the Euclidean distance between cells) is counted and extracted, and these distance feature data are sorted for subsequent model construction.

[0116] It should be noted that the staining data can be image data obtained by a specific staining technique, such as an immunostaining technique, and the first sample slice is processed so that the cells in the first sample slice exhibit different colors or markers, thereby facilitating subsequent observation and analysis of the distribution and characteristics of the cells. The process of determining the first intercellular distance data can utilize image analysis software (such as ImageJ) or bioinformatics tools (such as phenoptr) to calculate the relative distance between the same or different cells in the second sample slice based on the staining data and the spatial position data of the cells. For example, the Euclidean distance algorithm or the like can be used to obtain the distance information between each cell and other cells, thereby describing the spatial distribution difference between the cells; or the center of the nucleus in the staining image corresponding to the staining data can be marked first, and then the distance can be calculated by reading the image using the skimage method. For example, since the cell type label can include vascular cells (endothelial cells and pericytes), astrocytes, neurons (inhibitory neurons and excitatory neurons), etc., by analyzing the relative spatial distribution between these cells, the distance between the vascular cells, as well as the distance between the astrocytes, neurons and the vascular cells can be determined, thereby helping to deeply understand the spatial structure of the neurovascular unit and its functional association, and providing key distance feature data for accurate prediction and analysis of the neurovascular unit.

[0117] It should be noted that the image corresponding to the staining data can be a staining image obtained by staining one or more of the cell nucleus, nuclear membrane, cytoplasm, and cell membrane, such as the cell nucleus and cytoplasm, and the immunological library can only stain a few specific cells.

[0118] It should be noted that each data set of the sample slice containing the functional unit of the present application can also include spatial transcriptome data containing the first cell type, and fourth cell type and third intercellular distance data of the functional unit related cells in the second sample slice, and another data set is determined based on the transformed staining data of the image of the second sample slice. That is, it is not necessary to provide a real biochemical slice to obtain the other data set.

[0119] As shown in the left image of FIG. 8A, Figure 3 As shown in the left image of FIG. 8A, Figure 3 The left image of FIG. 8A shows the cell spatial distribution of the neurovascular unit in the immunohistochemical staining slice of the control chip (here, the chip is a slice, i.e., the second sample slice of the present application), and the right image of FIG. 8A shows the cell spatial distribution of the neurovascular unit in the immunohistochemical staining slice of the disease chip. Figure 3 The left image of FIG. 8A shows the cell spatial distribution of the neurovascular unit in the immunohistochemical staining slice of the control chip, Figure 3The right part of the figure in FIG. 1 shows the spatial distribution of the cells of the neurovascular unit in the immunohistochemical staining section of the disease group chip, and the functional unit of a manually marked neurovascular unit is circled by a dashed line. In the figure, the distribution of astrocytes (Astro), neurons, and the fluorescent dye of the cell nucleus (DAPI) is shown.

[0120] In step S2 of some embodiments, further, the application can train a machine learning module based on the data set of the at least one sample section to construct a functional unit prediction model. Since the functional unit prediction model is constructed based on the relevant information of the functional unit indicated by the second sample section, it can be used to predict the spatial structure of the same functional unit. After the training of the functional unit prediction model is completed, the functional unit prediction model can be used to accurately predict the composition and function of the functional unit with spatial structure.

[0121] It should be noted that the machine learning module of the application includes at least one of a supervised learning module, a semi-supervised learning module, an unsupervised learning module, a regression analysis module, a reinforcement learning module, a self-learning module, a feature learning module, a sparse dictionary learning module, an anomaly detection module, a generative adversarial network, or a correlation rule module, without limitation.

[0122] It should be noted that the application can be based on single-cell resolution transcriptome for prediction, so it is more accurate. The application can analyze the specific cell types contained in the predicted spatial structure of the functional unit, view the change characteristics of the pathology, enhance the analysis ability of the dynamic changes of the functional unit with spatial structure, and accurately predict its changes under different physiological or pathological conditions, providing a basis for the development of individualized treatment plans.

[0123] In some embodiments, the step of training a machine learning module based on the data set of the at least one sample section containing a functional unit to construct a functional unit prediction model can specifically include:

[0124] Constructing spatial adjacency data of the first sample section according to the spatial transcriptome data containing the first cell type, the spatial adjacency data being used to describe the adjacency structure between the same or different cells in the first sample section;

[0125] Extracting unit spatial data related to the functional unit from the spatial adjacency data according to the matching results of the second cell type and the first cell type of the functional unit related cells, and the first cell distance data;

[0126] According to the spatial structure indicated by the unit space data, center cell types, associated cell types, and distance thresholds between center cells, between associated cells, and between center cells and associated cells are determined from the first cell types; the center cells refer to cells that construct the functional unit, and the associated cells refer to non-center cells that have intercellular communication with the center cells;

[0127] The machine learning module is trained according to the center cell types, the associated cell types, and the distance thresholds, and a functional unit prediction model is constructed.

[0128] At this time, the spatial transcriptome data containing the first cell types can be presented in the form of a matrix, and the rows of the matrix represent genes, and the columns represent cells at spatial positions. In this way, the spatial positions of each cell in the first sample slice can be determined. The present application can set a proximity distance threshold for determining which cells are considered adjacent to each other to construct spatial adjacency data. At this time, the spatial adjacency data can be in the form of a matrix, i.e., a spatial adjacency matrix. The spatial adjacency matrix can be used to describe the adjacency structure between the same or different cells in the first sample slice. Therefore, the present application can establish a spatial adjacency matrix according to the first sample slice, and extract distance features of cells within a neural vascular unit from immunohistochemical staining of the first sample slice or the adjacent second sample slice. Then, a functional unit prediction model is established according to the information integration of the two.

[0129] At this time, the second cell types of the functional unit related cells can know the cell types required to construct the functional unit. At this time, the second cell types of the functional unit related cells and the first cell types can be type matched, and the relative distances between the same or different cells in the second sample slice can be used to extract unit space data related to the functional unit from the spatial adjacency data. At this time, the unit space data is used to indicate the spatial structure associated with the spatial distribution of the matched cells related to the functional unit from the first sample slice.

[0130] The center cell types are used to indicate the cell types that have a key role in the functional unit. For example, when the target functional unit is a neural vascular unit, the corresponding center cell types can include endothelial cells and pericytes. The center cell types indicated by the center cells refer to the cells that construct the functional unit. The associated cell types are used to indicate the cell types that are allowed to exist in the functional unit, such as neurons and astrocytes. The associated cell types indicated by the associated cells refer to non-center cells that have intercellular communication with the center cells. According to the matching of the center cell types and the associated cell types with the first cell types, the center cells and the associated cells related to the spatial structure of the functional unit can be screened from the cells indicated by the plurality of first cell types, avoiding the influence of invalid cell types on the specific spatial structure prediction of the functional unit.

[0131] The distance threshold between central cell types and associated cell types can refer to the distance threshold between central cells, between associated cells, and between central cells and associated cells; that is, the maximum distance between identical or different cells within the same functional unit. When the distance between a central cell and an associated cell is greater than the distance threshold, they belong to different functional units.

[0132] It should be noted that the process of constructing the functional unit prediction model based on the central cell type, associated cell types, and distance threshold in this application can be specifically represented as follows: First, construct a set of central cells based on the central cell type, which can be represented as " ",in, This represents the set of central cells selected from multiple cells in the first sample slice, where j represents the cells corresponding to the central cell type in the first sample slice, and cells represents the set of multiple cells in the first sample slice. This indicates the cell type corresponding to cell j. This represents the set of central cell types. Further, the process of identifying cells in the first sample slice that match the second cell type associated with the functional unit can be equivalent to: constructing a distance matrix based on the first inter-cell distance data using the Euclidean distance between points (points refer to the distance between the center points of two cells). ,like ,in, Represents cell pairs in the distance matrix ( The relative distance between cells () () indicates cell pairs of the same or different cell types in the second sample slice. and These represent the spatial location data of the corresponding cells, and ( () represents the candidate distance between each pair of cells.

[0133] Based on this, the spatially structured functional unit prediction model constructed using the distance matrix, central cell type, associated cell type, and distance threshold can be expressed as the following formula:

[0134]

[0135] in, This indicates the selection of a set of cells from multiple cells in the first sample slice that match the central cell type. The distance matrix is... In the matrix The data corresponding to the coordinates, where d represents the distance threshold. This represents a set composed of related cell types. Respectively indicate different associated cell types in the functional unit. Indicate cells in the first sample slice The corresponding first cell type belongs to an associated cell type, Indicate the functional unit numbered , and a sample slice can include multiple functional units. The above formula expresses that all cells contained in the functional unit are found out by traversing all cells in the first sample slice, combining the distance matrix, the set of central cells, and the set of associated cell types, and at least one functional unit is formed by integration. For example, when the functional unit is a neurovascular unit NVU, the constructed functional unit prediction model is an NVU model.

[0136] Among them, the distance threshold refers to the maximum distance between the same or different cells in the same functional unit. Therefore, when the intercellular distance between the associated cell and the central cell is greater than the distance threshold, it is considered that the cell indicated by the associated cell type is a cell that does not meet the structural requirements of the functional unit to which the central cell belongs.

[0137] In the above embodiment, the present application can establish a distance matrix of cells in a functional unit with a spatial structure by combining the Euclidean distance algorithm, and combine the definition of the functional unit with a spatial structure and the spatial position between cells to train a functional unit prediction model using a machine learning method. For example, the distance difference between vascular endothelial cells and astrocytes, neurons can reflect the functional state of the blood-brain barrier and the integrity of the neurovascular unit. The distance characteristics between cells are input into the model, such as: celltype1-celltype2: 30 (30 distance units); celltype1-celltype2: 350 (350 distance units) to predict the functional unit with a spatial structure. The established functional unit prediction model is to traverse the central cell set, find out all cell types in the distance matrix that meet the above distance feature set conditions, and then define them as the same unit. The functional unit prediction model predicts a spatial functional structure.

[0138] For example, as shown in Figure 4 , a schematic diagram of constructing a functional unit prediction model based on machine learning provided by the embodiment of the present application. Among them, the present application can first establish the spatial adjacency data 410 of the first sample slice, that is, the adjacency matrix, according to the spatial transcriptome data of the first sample slice. Secondly, the first intercellular distance data 420 of the cells in the neurovascular unit (that is, the distance matrix D) can be extracted from the immunohistochemical staining (that is, the staining data) of the first sample slice itself or the adjacent slice (the continuous slice close to the position of the spatial transcriptome chip in the three-dimensional structure of the tissue), and the functional unit prediction model 430 is established according to the information integration of the two.

[0139] In some embodiments, the application can perform functional unit prediction on the third sample slice through a functional unit prediction model, compare the staining data of the fourth sample slice, to determine the accuracy level of the machine learning module, wherein the fourth sample slice is the third sample slice or an adjacent slice of the third sample slice.

[0140] That is, after establishing the functional unit prediction model 430 (such as the prediction model based on neural vascular units), the application can also use the immunohistochemical (i.e. immunofluorescence) stained experimental slice (i.e. the fourth sample slice) as a verification standard, compare the model prediction result (i.e. the spatial distribution of the target cells determined in the predicted neural vascular unit in the third sample slice) with the spatial distribution of the immunohistochemical staining image, and match the spatial position of the predicted neural vascular unit with the position of the labeled blood vessels and neurons in the immunohistochemical staining image. Further, by calculating the overlap rate (i.e. cell coincidence degree) between the spatial distribution of the cells in the predicted neural vascular unit and the spatial distribution of the cells in the actual labeled functional unit, the proportion of the intersection of the predicted neural vascular unit region and the immunofluorescence labeled region is determined, so that the prediction accuracy of the functional unit prediction model 430 can be statistically analyzed. The results of experimental verification show that the coincidence degree between the predicted neural vascular unit and the immunofluorescence staining image is about 75%, which can also reflect that the model has a certain accuracy in spatial structure prediction. After accurately determining the cells corresponding to the functional unit, the mutual relationship between the cells in the functional unit can be determined in combination with the type characteristics of the cells and their positions in space.

[0141] It should be noted that after determining the cells corresponding to the functional unit, the application can also visualize the prediction results to intuitively display the organizational structure and functional characteristics of the functional unit with spatial structure in space.

[0142] It should be noted that, during the training process of the functional unit prediction model, this application can use the neurovascular unit in the human hippocampus sample as an example to predict and analyze the neurovascular unit. The original process of the distance feature extraction model in this application can be divided into four parts: data preprocessing, constructing a Convolutional Neural Network (CNN) model, training the model, and feature extraction. Specifically, the image data is first preprocessed to accurately label different cell types (such as Neuron, Astro, etc.) and neurovascular units, that is, to clarify the unit and category to which the cell belongs, and to normalize the image pixel values ​​to the range of 0-1. Next, a CNN model is constructed, specifically through the following sequence: a first convolutional layer (using 16 3x3 kernels, stride 1, padding 1, and ReLU activation function to extract initial image features), a first pooling layer (using 2x2 max pooling with a stride 2 to downsample the feature map and reduce feature dimensionality), a second convolutional layer (using 32 3x3 kernels, stride 1, padding 1, and ReLU activation function to further extract more complex features), a second pooling layer (again using 2x2 max pooling with a stride 2 for downsampling), and a fully connected layer (flattening the pooled features and connecting them to a fully connected layer; the number of neurons can be set to 128, and the activation function is ReLU). During training, the preprocessed data can be divided into training and validation sets. Multiple rounds of training are performed on the training set, with loss calculated and model parameters updated in each round. Simultaneously, model performance is monitored on the validation set to prevent overfitting. After the model is trained, the test image is input into the model, and feature vectors are obtained from the output layer. These feature vectors contain information such as distance features between different cell types, which can be used for subsequent analysis.

[0143] In some embodiments, such as Figure 5 The figure shows a schematic diagram of the NVU model prediction convergence curve provided in the embodiment of this application. The figure shows the trend of the NVU model's accuracy in predicting NVU on immunofluorescence slides as a function of the number of training iterations. That is, the higher the number of training iterations, the more accurate and stable the model's prediction results for functional units with spatial structures.

[0144] In some embodiments, Figure 2 Each cell in the microarray data has its own gene expression profile and location coordinates. Therefore, this application can extract the location coordinates corresponding to each cell type and apply them to a distance-based model (such as the NVU model) to obtain... Figure 6 Regarding the control group chip (e.g.) Figure 6 (The image on the left) and disease group chips (such as...) Figure 6 The diagram on the right shows the spatial distribution of the neurovascular unit, with different colors representing different cell types.

[0145] In some embodiments, the application also provides a functional unit prediction method, including but not limited to the following steps:

[0146] Obtaining spatial transcriptome data of a to-be-predicted slice, and performing cell classification on a plurality of cells in the to-be-predicted slice based on the spatial transcriptome data of the to-be-predicted slice to obtain a third cell type of the plurality of cells in the to-be-predicted slice, the spatial transcriptome data of the to-be-predicted slice including expression level and spatial distribution information of the transcriptome of the plurality of cells in the to-be-predicted slice;

[0147] According to the constructed functional unit prediction model, performing functional unit prediction based on the spatial transcriptome data of the to-be-predicted slice containing the third cell type.

[0148] In actual application stage, the application can perform cell classification on a plurality of cells in the to-be-predicted slice based on the spatial transcriptome data of the to-be-predicted slice, and the classification process has been described in detail in the above embodiments. The third cell type is used to represent the cell type corresponding to each cell in the to-be-predicted slice. The spatial transcriptome data of the to-be-predicted slice includes expression level and spatial distribution information of the transcriptome of the plurality of cells in the to-be-predicted slice.

[0149] In some embodiments, according to the constructed functional unit prediction model, the step of performing functional unit prediction based on the spatial transcriptome data of the to-be-predicted slice containing the first cell type can specifically include:

[0150] According to the constructed functional unit prediction model, determining a central cell type, an associated cell type, and a distance threshold between the central cells, between the associated cells, and between the central cells and the associated cells;

[0151] Based on the matching result of the central cell type and the third cell type, determining the central cells of the functional unit from the plurality of cells in the to-be-predicted slice;

[0152] Based on the matching result of the associated cell type and the third cell type, determining the associated cells of the functional unit from the plurality of cells in the to-be-predicted slice;

[0153] According to the spatial transcriptome data, obtaining second cell distance data between the central cells, between the associated cells, and between the central cells and the associated cells in the to-be-predicted slice, the second cell distance data being used to describe the relative distance between the central cells, between the associated cells, and between the central cells and the associated cells in the to-be-predicted slice;

[0154] According to the comparison result of the relative distance and the distance threshold, determining target cells related to the functional unit to which the central cells belong from the plurality of cells in the to-be-predicted slice, the target cells including the central cells and the associated cells.

[0155] According to the target cells and the spatial distribution information of the target cells in the to-be-predicted slice, a functional unit of the to-be-predicted slice is predicted.

[0156] According to the formula constructed according to the above model, it is known that the functional unit prediction model predefines a set of central cell types corresponding to the functional unit, a set of associated cell types, and distance thresholds between the central cells, between the associated cells, and between the central cells and the associated cells, and thus the central cells and the associated cells related to the functional unit can be screened out from the plurality of cells of the to-be-predicted slice based on the matching of the defined sets and the third cell type. The central cell can refer to a key cell that constructs a functional unit structure. The associated cell can refer to a non-central cell that has intercellular communication with the central cell, that is, a cell allowed to exist in the functional unit, and the spatial distance between the associated cells and the central cells in the same functional unit should be less than or equal to the distance threshold.

[0157] Further, the target cell refers to a cell in the plurality of cells of the to-be-predicted slice that can be divided into the spatial structure of the functional unit. After the target unit is determined, the spatial structure of the functional unit of the to-be-predicted slice can be constructed according to the spatial distribution information of the target cell in the to-be-predicted slice, that is, the prediction of the spatial structure of the functional unit is realized.

[0158] In some embodiments, after the functional unit prediction is performed based on the spatial transcriptome data of the to-be-predicted slice containing the third cell type according to the constructed functional unit prediction model, the method further comprises:

[0159] extracting expression levels and spatial distribution information of the transcriptome of the target cell in the to-be-predicted slice from the spatial transcriptome data of the to-be-predicted slice;

[0160] According to the expression levels and spatial distribution information of the transcriptome of the target cell in the to-be-predicted slice, density statistics and functional analysis are performed on the functional unit in the to-be-predicted slice.

[0161] The expression levels of the transcriptome of the target cell in the to-be-predicted slice correspond to the cell gene expression data, which refers to data that can reflect the gene expression in the cell and can reflect the activity of the gene in the cell, that is, which genes are expressed and how much they are expressed. The spatial distribution information of the transcriptome of the target cell in the to-be-predicted slice corresponds to the cell position data, which refers to data that can describe the specific position of each cell in the first sample slice, which can usually be represented by coordinates, such as coordinates in a two-dimensional plane or coordinates in a three-dimensional space.

[0162] The present application can determine the spatial distribution characteristics of the predicted functional units in the target section according to the expression level and spatial distribution information of the transcriptome of the target cells in the to-be-predicted section, which can be used to describe the distribution rule and mode of the target cells in the functional units in space, such as whether the cells are uniformly distributed, aggregated distributed, or distributed in a specific geometric shape, and the like, so as to further realize the density statistics and functional analysis of the functional units in the to-be-predicted section. The present application can comprehensively analyze the gene expression data, corresponding cell types and spatial distribution characteristics of the target cells, and can use various analysis methods such as differential expression analysis, enrichment analysis, clustering analysis, and the like, to reveal the rules and mechanisms of gene expression in the functional units. For example, differential expression analysis can be used to find genes that are differentially expressed in different cell types or spatial positions; enrichment analysis can be used to determine the biological functions and signal pathways involved in these differentially expressed genes; and clustering analysis can be used to classify cells with similar gene expression patterns and spatial distribution characteristics. In this way, by deeply understanding the gene expression regulation mechanisms, cell-cell interactions and functions in biological processes of the target functional units, important theoretical basis and practical guidance can be provided for the fields of disease diagnosis, drug development and the like. Therefore, on the basis of spatial transcriptome data at single-cell resolution, the present application can use the functional unit prediction model established based on the distance between cells to predict the functional units, and analyze the composition, distribution and association with other brain regions of the functional units through the prediction results.

[0163] For example, after predicting the spatial structure corresponding to the neurovascular unit, the present application can also perform density statistics and functional analysis on the prediction results, so as to intuitively show the spatial organization structure and functional characteristics of the neurovascular unit. As shown in FIGS. 1-3, FIGS. 4-6, and FIGS. 7-9, they are schematic diagrams of the density statistics and functional scoring changes of the neurovascular unit functional units of the control group and disease group chips provided by the embodiments of the present application. For FIGS. 1-3, Figure 7 , Figure 8 and Figure 9 As shown in FIGS. 1-3, FIGS. 4-6, and FIGS. 7-9, they are schematic diagrams of the density statistics and functional scoring changes of the neurovascular unit functional units of the control group and disease group chips provided by the embodiments of the present application. For FIGS. 1-3, Figure 7 , the horizontal axis indicates two groups of samples (control group chips and disease group chips), and the vertical axis indicates the average density (Average Density), which is in units of count per square millimeter (count / mm 2 ), which is used to measure the density of the relevant indicators on the chip. It can be seen that in terms of density changes, the average density of the disease group is significantly higher than that of the control group (the difference in the height of the column chart is obvious), indicating that the target cells are actively infiltrating or proliferating in the tissue under the disease state. For FIGS. 4-6, Figure 8The horizontal axis represents the two sample groups (control group chip and disease group chip), and the vertical axis represents the positive regulation score (range approximately -0.5 to 1.5). It can be seen that in the response to hypoxia pathways, hypoxia-related genes (such as HIF1α, VEGF, and GLUT1) are highly expressed in the disease group, potentially driving angiogenesis, metabolic reprogramming, and tumor progression (in the case of tumor research). For Figure 9 The horizontal axis represents the two sample groups (control group chip and disease group chip), and the vertical axis represents the positive regulation score (range approximately -1 to 1). It is evident that the disease group shows significant activation in immune response regulation (increased positive regulation score), indicating upregulation of key immune pathways (such as TNFα / NF-κB and IFNγ signaling) or immune checkpoint genes (such as PD-L1 and CTLA4), suggesting immune microenvironment remodeling. Therefore, this application not only efficiently predicts the structure of functional units with spatial structures but also reveals the dynamic changes of different cell types in space, providing a profound understanding of neurological health and thus supporting the early diagnosis and precision treatment of neurovascular diseases.

[0164] Reference Figure 10 In a specific embodiment, the method for constructing a functional unit prediction model provided in this application may specifically include the following steps:

[0165] Step S1001: Obtain spatial transcriptome data from multiple sample slices.

[0166] The initial spatial transcriptome data of the sample slices includes spatial gene expression data and spatial location data of multiple cells within the initial sample slices. Before predicting functional units with spatial structures, this application can prepare and perform quality control on the slices to screen out the first sample slices that meet the quality control standards.

[0167] Step S1002: Data extraction is performed based on spatial transcriptome data to obtain the number of genes and the proportion of mitochondrial genes in each cell of each sample slice;

[0168] Step S1003: Based on the number of genes and the proportion of mitochondrial genes in each cell, multiple sample slices are screened to determine the first sample slice.

[0169] The first sample slice is an initial sample slice meeting a quality control standard, and the quality control standard can include that at least 50% of cells in a tissue slice meet the following requirements at the same time: an average value of a gene number (nFeature) of each cell should be greater than 100, and an average value of a mitochondrial gene proportion (percent.MT) in each cell should be less than 20. That is, a chip meeting the standard is considered to be qualified in quality, and can be used for prediction analysis of a neurovascular unit.

[0170] In step S1004, spatial transcriptome data of the first sample slice is acquired, and cells in the first sample slice are classified based on the spatial transcriptome data to obtain cell type labels of the cells in the first sample slice.

[0171] After the first sample slice is determined, spatial transcriptome data with single-cell resolution of each first sample slice can be acquired, and cell annotation can be performed based on the single-cell spatial transcriptome data.

[0172] In step S1005, staining data of the second sample slice is acquired, and a second cell type of a functional unit related cell in the second sample slice and first intercellular distance data are determined based on the staining data.

[0173] The first intercellular distance data is used to describe relative distances between cells of the same cell type or different cell types in the second sample slice, and the second sample slice is the first sample slice or an adjacent slice of the first sample slice. The cell distance features in the functional unit can be further counted, that is, the distance features between different cell types can be confirmed and counted according to the immunohistochemical staining results in the sample slice itself or the corresponding adjacent slice.

[0174] In step S1006, spatial adjacency data of the first sample slice is constructed according to the spatial transcriptome data containing the first cell type.

[0175] The spatial adjacency data is used to describe adjacency structure preset label screening conditions of the same or different cells in the first sample slice. The preset label screening conditions correspond to conditions for constructing a model for structure prediction of a target functional unit, and the preset label screening conditions of different target functional units are different.

[0176] In step S1007, unit spatial data related to the functional unit is extracted from the spatial adjacency data according to a matching result of the second cell type of the functional unit related cell and the first cell type, and the first intercellular distance data.

[0177] Step S1008, according to the spatial structure indicated by the cell space data, determining the center cell type, the associated cell type, and the distance threshold between the center cells, the associated cells, and the center cells and the associated cells from the first cell type.

[0178] Wherein, there may be differences in the distance threshold between the center cells, the associated cells, and the center cells and the associated cells in different functional unit predictions. The center cell refers to the cell that constitutes the functional unit, and the associated cell refers to the non-center cell that has cell-to-cell communication with the center cell.

[0179] Step S1009, training the machine learning module according to the center cell type, the associated cell type, and the distance threshold, and constructing a functional unit prediction model.

[0180] Wherein, after determining the center cell type, the associated cell type, and the distance threshold, a spatial relationship model between the cells in the functional unit can be established, the relative distance between the cells and the spatial distribution characteristics thereof are quantified, and the corresponding cells of the functional unit are determined by traversing each cell, so as to determine the spatial structure corresponding to the functional unit.

[0181] The present application first defines the functional unit with spatial structure at the level of the transcriptome, applies the prediction method to the research of neurological diseases, and explores the change rule of the functional unit with spatial structure in neurological diseases. Based on the constructed functional unit prediction model, a new experimental strategy or intervention method is designed to promote the early diagnosis and treatment of the related diseases of the functional unit with spatial structure. Moreover, the present application has a lower requirement for the professional degree of the professional personnel, and causes less damage to the sample, so that high-quality transcriptome data can be effectively obtained, thereby providing great convenience for the researchers.

[0182] The function unit prediction method provided by the embodiments of the present application significantly improves the spatial resolution and cell analysis accuracy of the function unit with spatial structure by using the single-cell resolution spatial transcriptome technology, can accurately reveal the spatial position and slight changes of the cells, and thus improves the analysis accuracy. Compared with the related art, the present application reduces the dependence on the section sample, can avoid complex immunohistochemical staining and microscope operation, and by combining the Euclidean distance algorithm and the machine learning method, a more accurate function unit prediction model can be established, which can more accurately predict the composition and function of the function unit with spatial structure, thereby promoting the early diagnosis and treatment of neurological diseases (such as stroke, Alzheimer's disease, etc.). At the same time, the present application can enhance the analysis capability of the dynamic changes of the function unit with spatial structure, i.e., by using the intercellular distance data and the cell type label to determine the target function unit, the changes of the target function unit under different physiological or pathological conditions can be accurately predicted, which provides a basis for the development of individualized treatment plan. In addition, the present application uses an automatic machine learning analysis method, so that the prediction process is more efficient and accurate, reduces manual intervention, saves experimental time, reduces cost and technical threshold, and can provide convenience for more researchers. Therefore, the present application can be better applied to the prediction of the function unit with spatial structure, improve the prediction accuracy of the function unit with spatial structure, and thus provide effective technical support for the early diagnosis and precise treatment of neurological diseases, and has a wide application prospect.

[0183] Reference Figure 11 The embodiments of the present application also provide a function unit prediction model construction device, which comprises:

[0184] The data set acquisition module 1110 is configured to acquire at least one data set of a sample section containing a function unit, and the data set acquisition module 1110 comprises a first cell classification module 1111 and a second cell classification module 1112. The first cell classification module 1111 is configured to acquire spatial transcriptome data containing a first cell type, and the acquisition method comprises acquiring spatial transcriptome data of a first sample section, and performing cell classification on a plurality of cells in the first sample section based on the spatial transcriptome data to obtain the first cell type of the plurality of cells in the first sample section. The spatial transcriptome data comprises the expression level and spatial distribution information of the transcriptome of the plurality of cells in the first sample section. The second cell classification module 1112 is configured to acquire a second cell type of a function unit related cell and first intercellular distance data, and the acquisition method comprises acquiring staining data of a second sample section, and determining the second cell type of the function unit related cell and the first intercellular distance data in the second sample section based on the staining data. The first intercellular distance data is used to describe the relative distance between the same or different cells in the second sample section. The second sample section is the first sample section or an adjacent section of the first sample section.

[0185] The model construction module 1120 is configured to train the machine learning module based on the data set of the at least one group of sample slices to construct the functional unit prediction model.

[0186] It can be seen that the content in the above functional unit prediction model construction method embodiments is applicable to the embodiments of the functional unit prediction model construction apparatus, the functional unit prediction model construction apparatus embodiments specifically implement the same functions as the above functional unit prediction model construction method embodiments, and achieve the same beneficial effects as the above functional unit prediction model construction method embodiments.

[0187] With reference to Figure 12 The embodiments of the present application also provide a functional unit prediction apparatus, which comprises:

[0188] The third cell classification module 1210 is configured to obtain spatial transcriptome data of a to-be-predicted slice, and perform cell classification on a plurality of cells in the to-be-predicted slice based on the spatial transcriptome data of the to-be-predicted slice to obtain third cell types of the plurality of cells in the to-be-predicted slice, wherein the spatial transcriptome data of the to-be-predicted slice comprises expression levels and spatial distribution information of transcriptomes of the plurality of cells in the to-be-predicted slice.

[0189] The functional unit prediction module 1220 is configured to perform functional unit prediction based on the spatial transcriptome data containing the third cell types of the to-be-predicted slice according to the constructed functional unit prediction model.

[0190] It can be seen that the content in the above functional unit prediction method embodiments is applicable to the embodiments of the functional unit prediction apparatus, the functional unit prediction apparatus embodiments specifically implement the same functions as the above functional unit prediction method embodiments, and achieve the same beneficial effects as the above functional unit prediction method embodiments.

[0191] With reference to Figure 13 , Figure 13 The hardware structure of an electronic device of another embodiment is shown, and the electronic device comprises:

[0192] The processor 1301 can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided by the embodiments of the present application.

[0193] The memory 1302 can be implemented in the form of Read Only Memory (ROM), static storage device, dynamic storage device or Random Access Memory (RAM), etc. The memory 1302 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 1302 and are called and executed by the processor 1301 to implement the function unit prediction model construction method and the function unit prediction method of the embodiments of the present application;

[0194] The input / output interface 1303 is configured to realize information input and output.

[0195] The communication interface 1304 is configured to realize the communication interaction between the device and other devices, and the communication can be realized by wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0196] The bus 1305 transmits information between various components (such as the processor 1301, the memory 1302, the input / output interface 1303 and the communication interface 1304) of the device.

[0197] The processor 1301, the memory 1302, the input / output interface 1303 and the communication interface 1304 are connected to each other through the bus 1305 for internal communication connection in the device.

[0198] The embodiments of the present application also provide a computer readable storage medium, including a memory and a processor, the memory stores a computer program, and the processor executes the computer program to realize the function unit prediction model construction method and the function unit prediction method described above.

[0199] The embodiments of the present application also provide a computer program product, which includes a computer program. The processor of the computer device reads the computer program and executes it, so that the computer device executes the function unit prediction model construction method and the function unit prediction method described above.

[0200] The terms "first", "second", "third", "fourth" and the like in the description of the disclosure and the above drawings, if any, are used to distinguish similar objects, and do not necessarily have to be described in a particular order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the disclosure described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "contain" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device containing a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0201] It should be understood that in the present disclosure, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the relationship between the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that there are three cases: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0202] It should be understood that in the description of the embodiments of the present application, the meaning of multiple (or multiple) is two or more, greater than, less than, more than, etc. is not included in the number, and above, below, etc. is included in the number.

[0203] In several embodiments provided by the present disclosure, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed objects can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0204] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0205] In addition, each functional unit in various embodiments of the present disclosure can be integrated into one processing unit, or each unit can exist physically, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0206] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present disclosure essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present disclosure. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0207] It should also be appreciated that the various embodiments provided by the present application can be combined in any way to achieve different technical effects.

[0208] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present disclosure, and these equivalent modifications or replacements are included in the scope defined by the claims of the present disclosure.

Claims

1. A method for constructing a functional unit prediction model, characterized in that, The method includes: Step S1: Obtain at least one set of sample slices containing functional units, wherein each set of sample slices containing functional units includes: Step S1.1, spatial transcriptome data containing the first cell type, is obtained by: acquiring spatial transcriptome data of a first sample slice, and classifying multiple cells in the first sample slice based on the spatial transcriptome data to obtain the first cell type of the multiple cells in the first sample slice; the spatial transcriptome data includes the expression level and spatial distribution information of the transcriptome of multiple cells in the first sample slice; and Step S1.2, obtaining the second cell type and first inter-cell distance data of the functional unit related cells, the method includes: obtaining staining data of the second sample slice, determining the second cell type and first inter-cell distance data of the functional unit related cells in the second sample slice based on the staining data, wherein the first inter-cell distance data is used to describe the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or an adjacent slice of the first sample slice; Step S2: Train the machine learning module based on the dataset of at least one set of sample slices containing functional units to construct a functional unit prediction model.

2. The method for constructing a functional unit prediction model according to claim 1, characterized in that, The machine learning module includes at least one of the following: supervised learning module, semi-supervised learning module, unsupervised learning module, regression analysis module, reinforcement learning module, self-learning module, feature learning module, sparse dictionary learning module, anomaly detection module, generative adversarial network, or association rule module.

3. The method according to claim 1, characterized in that, The step of training a machine learning module based on the dataset containing at least one set of sample slices with functional units to construct a functional unit prediction model includes: Spatial adjacency data of the first sample slice is constructed based on spatial transcriptome data containing the first cell type. The spatial adjacency data is used to describe the adjacency structure between the same or different cells in the first sample slice. Based on the matching results of the second cell type and the first cell type of the related cells of the functional unit, and the distance data between the first cells, the unit spatial data related to the functional unit is extracted from the spatial adjacency data; Based on the spatial structure indicated by the unit spatial data, the central cell type, associated cell type, and distance thresholds between central cells, between associated cells, and between the central cell and the associated cell are determined from the first cell type; the central cell refers to the cell constituting the functional unit, and the associated cell refers to a non-central cell that has inter-cell communication with the central cell. The machine learning module is trained based on the central cell type, the associated cell type, and the distance threshold to construct a functional unit prediction model.

4. The method for constructing a functional unit prediction model according to claim 1, characterized in that, The cell classification of multiple cells in the first sample slice based on the spatial transcriptome data to obtain the first cell type of the multiple cells in the first sample slice includes: Using single-cell transcriptome sequencing technology, single-cell transcriptome sequencing data of multiple cells from the sample from which the first sample slice is derived are obtained. Based on the expression level of the transcriptome of the multiple cells in the first sample slice and the single-cell transcriptome sequencing data, common gene expression data of each cell in the first sample slice are determined. Based on the common gene expression data and preset cell marker genes, cell annotation is performed on multiple cells in the first sample slice to obtain the first cell type of the multiple cells in the first sample slice.

5. The method for constructing a functional unit prediction model according to claim 4, characterized in that, The method for determining common gene expression data for each cell in the first sample slice based on the transcriptome expression levels of the multiple cells in the first sample slice and the single-cell transcriptome sequencing data includes: Based on the expression levels of the transcriptomes of the multiple cells in the first sample slice, the first gene expression matrix of the multiple cells in the first sample slice is determined. The second gene expression matrix of multiple cells in the first sample slice was determined based on the single-cell transcriptome sequencing data. Based on the common genes of the multiple cells in the first gene expression matrix and the second gene expression matrix, the common gene expression data of each cell in the first sample slice is obtained.

6. The method for constructing a functional unit prediction model according to claim 1, characterized in that, The functional unit prediction model is used to predict the functional units of the third sample slice, and the prediction is compared with the coloring data of the fourth sample slice to determine the accuracy level of the machine learning module. The fourth sample slice is the third sample slice or a neighboring slice of the third sample slice.

7. A method for predicting functional units, characterized in that, The method includes: Spatial transcriptome data of the slice to be predicted is obtained, and multiple cells in the slice to be predicted are classified based on the spatial transcriptome data of the slice to be predicted to obtain the third cell type of multiple cells in the slice to be predicted. The spatial transcriptome data of the slice to be predicted includes the expression level and spatial distribution information of the transcriptome of multiple cells in the slice to be predicted. The functional unit prediction model constructed according to the functional unit prediction model construction method according to any one of claims 1 to 6 predicts functional units based on the spatial transcriptome data containing a third cell type of the slice to be predicted.

8. The method according to claim 7, characterized in that, The functional unit prediction model constructed according to the functional unit prediction model construction method of any one of claims 1 to 6, performs functional unit prediction based on spatial transcriptome data containing a third cell type of the slice to be predicted, including: The functional unit prediction model constructed according to any one of claims 1 to 6 determines the central cell type, associated cell type, and distance thresholds between central cells, between associated cells, and between the central cell and the associated cell in relation to the functional unit. Based on the matching results of the central cell type and the third cell type, the central cell of the functional unit is determined from multiple cells in the slice to be predicted; Based on the matching results of the associated cell type and the third cell type, the associated cells of the functional unit are determined from multiple cells in the slice to be predicted; Based on the spatial transcriptome data, second cell distance data are obtained between the central cells, between the associated cells, and between the central cells and the associated cells in the slice to be predicted. The second cell distance data is used to describe the relative distances between the central cells, between the associated cells, and between the central cells and the associated cells in the slice to be predicted. Based on the comparison between the relative distance and the distance threshold, target cells related to the functional unit to which the central cell belongs are identified from multiple cells in the slice to be predicted; the target cells include the central cell and the associated cells. Based on the target cells and their spatial distribution information in the slice to be predicted, the functional units of the slice to be predicted are predicted.

9. The method according to claim 8, characterized in that, After the functional unit prediction model constructed according to the functional unit prediction model construction method according to any one of claims 1 to 6 performs functional unit prediction based on the spatial transcriptome data of the slice to be predicted containing a third cell type, the functional unit prediction method further includes: Extract the expression level and spatial distribution information of the target cell's transcriptome in the predicted slice from the spatial transcriptome data of the predicted slice; Based on the expression level and spatial distribution information of the target cell transcriptome in the slice to be predicted, density statistics and functional analysis are performed on the functional units in the slice to be predicted.

10. A device for constructing a functional unit prediction model, characterized in that, The device includes: A dataset acquisition module is used to acquire at least one set of sample slices containing functional units. The dataset acquisition module includes a first cell classification module and a second cell classification module. The first cell classification module is used to acquire spatial transcriptome data containing a first cell type. The acquisition method includes: acquiring spatial transcriptome data of the first sample slice, and classifying multiple cells in the first sample slice based on the spatial transcriptome data to obtain a first cell type of the multiple cells in the first sample slice; the spatial transcriptome data includes the expression level and spatial distribution information of the transcriptome of multiple cells in the first sample slice. The second cell classification module is used to acquire second cell types of cells related to functional units and first inter-cell distance data. The acquisition method includes: acquiring staining data of the second sample slice, and determining the second cell types of cells related to functional units and first inter-cell distance data in the second sample slice based on the staining data; the first inter-cell distance data describes the relative distance between the same or different cells in the second sample slice; wherein the second sample slice is the first sample slice or an adjacent slice of the first sample slice. The model building module is used to train the machine learning module based on the dataset containing at least one set of sample slices with functional units, and to build a functional unit prediction model.

11. A functional unit prediction device, characterized in that, The device includes: The third cell classification module is used to acquire spatial transcriptome data of the slice to be predicted, and classify multiple cells in the slice to be predicted based on the spatial transcriptome data of the slice to be predicted to obtain the third cell type of multiple cells in the slice to be predicted. The spatial transcriptome data of the slice to be predicted includes the expression level and spatial distribution information of the transcriptome of multiple cells in the slice to be predicted. The functional unit prediction module is used to construct a functional unit prediction model according to the functional unit prediction model construction method according to any one of claims 1 to 6, and to perform functional unit prediction based on the spatial transcriptome data containing a third cell type of the slice to be predicted.

12. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the functional unit prediction model construction method according to any one of claims 1 to 6, and the functional unit prediction method according to any one of claims 7 to 9.

13. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the functional unit prediction model construction method according to any one of claims 1 to 6, and the functional unit prediction method according to any one of claims 7 to 9.

14. A computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the functional unit prediction model construction method according to any one of claims 1 to 6 and the functional unit prediction method according to any one of claims 7 to 9.

Citation Information

Patent Citations

  • Method and product for localised or spatial detection of nucleic acid in a tissue sample

    EP2697391A1

  • Method for generating labeled nucleic acid molecular population and kit thereof

    WO2023115536A1

  • Methods, compositions, and kits for determining the location of an analyte in a biological sample

    WO2024086167A3

  • Spatial transcriptome analysis method and device based on deep learning and readable medium

    CN116994245A

  • Three-level lymph structure identification and classification method and application thereof in cancer treatment

    CN118314965A