Methods, systems, media, and electronic devices for identifying spatially specific regulatory molecules of adenomyosis.
By using spatial transcriptome sequencing and cell communication analysis, spatially specific regulatory molecules of adenomyosis were identified, solving the problem that existing technologies could not resolve spatially specific regulatory molecules of adenomyosis, providing new diagnostic and therapeutic targets, and enabling precise research on the microenvironment of adenomyosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2026-03-13
AI Technical Summary
Existing research methods are insufficient to fully utilize spatial location information to analyze the molecular characteristics of different developmental stages of adenomyosis. They lack specific diagnostic indicators and therapeutic targets. Existing technologies lose spatial information and cannot study the spatially specific regulatory molecules of adenomyosis from the perspective of different microenvironments in disease development.
Using spatial transcriptome sequencing technology combined with Bayesian models, zero-distribution algorithms, and NicheNet algorithms, and integrating uterine tissue and single-cell data, we screened out cellular communication signals with significantly differentially expressed ligands, receptors, and target genes in the adenomyosis microenvironment, and identified spatially specific regulatory molecules for adenomyosis.
By using spatial transcriptome sequencing technology and cell communication analysis, we can accurately identify spatially specific regulatory molecules related to the microenvironment of adenomyosis and different developmental processes. This solves the problem that existing technologies cannot fully utilize spatial location information and provides new diagnostic and therapeutic targets.
Smart Images

Figure CN118335207B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioengineering technology and relates to adenomyosis, particularly to a method, system, medium, and electronic device for identifying spatially specific regulatory molecules of adenomyosis. Background Technology
[0002] Adenomyosis refers to the presence of ectopic, non-neoplastic endometrial glands and stroma within the myometrium. Typical symptoms include dysmenorrhea, menorrhagia, chronic pelvic pain, and infertility. Its incidence is high, exceeding 34% in nulliparous women of childbearing age. Adenomyosis negatively impacts female fertility, assisted reproductive outcomes, and pregnancy outcomes: the clinical pregnancy rate for patients with adenomyosis is only 18.2%; it increases the risk of miscarriage after assisted reproductive technology and reduces live birth rates. The clinical pregnancy rate in patients with adenomyosis is significantly lower than in women without adenomyosis, with significantly reduced implantation rates, clinical pregnancy rates per cycle, clinical pregnancy rates per embryo transfer, sustained pregnancy rates, and live birth rates; adenomyosis also increases the risk of miscarriage, premature birth, small for gestational age infants, and preeclampsia. However, there is currently a lack of specific laboratory diagnostic indicators, specific causal treatment drugs, and unified classification and grading standards related to prognosis and treatment regimens for adenomyosis.
[0003] Adenomyosis is a typical disease with spatially specific distribution. Adenomyosis lesions originate from endometrial invagination in the basal layer of the endometrium. Imbalance at the junction of the endometrium and myometrium is the initiating factor in the pathogenesis of adenomyosis, which may be caused by a combination of increased endometrial invasiveness and myometrial damage. Simultaneously, adenomyosis is a fibrotic disease; fibrosis around the lesions is a characteristic feature of adenomyosis and a significant reason for the current poor efficacy of drug treatment. The endometrial invagination microenvironment and the lesion microenvironment at different depths represent different processes in the development of adenomyosis, each with its own molecular characteristics. Elucidating these molecular characteristics is crucial for the development of specific non-invasive diagnostic methods and intervention targets for adenomyosis. The spatially specific regulatory molecules of adenomyosis play a role in different microenvironments and are expressed in a manner correlated with the spatial distribution of adenomyosis lesions. However, currently, there are no reported methods, devices, or equipment for systematically studying the spatially specific regulatory molecules of adenomyosis. There is an urgent need for new research methods, devices, and equipment to comprehensively and meticulously study the spatially specific regulatory molecules of different microenvironments in adenomyosis, so as to provide direction for the development of new diagnostic and therapeutic targets.
[0004] High-throughput sequencing technology has provided us with a more comprehensive way to observe the pathophysiological mechanisms of diseases. Sequencing mRNA can fully reveal the changes in the crucial process of transcription, the central dogma. Microarrays, transcriptome sequencing, single-cell transcriptome sequencing, and spatial transcriptome sequencing are the four main mRNA sequencing technologies, each with its own focus. In previous studies, microarrays, transcriptome sequencing, and single-cell transcriptome sequencing have revealed the overall transcriptional changes and specific cell type roles in adenomyosis. However, the different processes of adenomyosis pathogenesis and progression have not yet been studied from the perspective of spatially specific regulatory molecules.
[0005] Microarray chips, by hybridizing pre-designed probes with sample cDNA, can obtain relative expression profiles of a range of known genes, but their ability to detect low-abundance transcripts is insufficient. Transcriptomics, on the other hand, directly sequences cDNA without relying on pre-designed probes, offering high sensitivity and the ability to detect novel gene changes. Single-cell transcriptomics sequencing the transcriptional information of individual cells allows for precise observation of physiological and pathological processes down to the cellular level. However, the biggest drawback of these technologies is the loss of spatial information, preventing the study of dynamic changes and interactions of important pathological processes within different microenvironments of disease development, as well as the spatially specific regulatory molecules that play a crucial role in these processes. Adenomyosis is a typical spatially specific disease, making omics technologies incorporating spatial information of significant importance for its research. Summary of the Invention
[0006] The purpose of this invention is to provide a method, system, medium, and electronic device for identifying spatially specific regulatory molecules in adenomyosis, in order to solve the problem that existing research methods cannot fully utilize spatial location information to analyze the molecular characteristics of different developmental stages of adenomyosis.
[0007] This invention provides a method for identifying space-specific regulatory molecules of adenomyosis, comprising the following steps:
[0008] Step S1: Obtain uterine tissue lesion samples and uterine tissue control samples. Use spatial transcriptome sequencing technology to perform mRNA sequencing detection on uterine tissue lesion samples and uterine tissue control samples to obtain gene expression data containing spatial location information in uterine tissue lesion samples and uterine tissue control samples.
[0009] Step S2: Using a Bayesian model, gene expression data containing spatial location information from uterine tissue lesion samples and uterine tissue control samples are integrated and analyzed with uterine single-cell data from the same menstrual cycle to obtain multiple cell subpopulations containing spatial location information, and each obtained cell subpopulation is defined as the target subpopulation.
[0010] Step S3: Use the zero distribution algorithm to obtain the interaction probability distribution and the significance of cell interaction probability among each target subpopulation, and define the spatial location of the target subpopulation with a cell interaction probability significance value of less than 0.05 and with typical adenomyosis morphological characteristics as the adenomyosis microenvironment.
[0011] Step S4: Using the NicheNet algorithm, cell communication analysis is performed on ligands, receptors, and target genes in the adenomyosis microenvironment. Cell communication signals with significantly differential expression of ligands, receptors, and target genes in the adenomyosis microenvironment are screened out, and the ligands of the cell communication signals that promote the pathogenesis of adenomyosis are identified as space-specific regulatory molecules of adenomyosis.
[0012] Furthermore, in step S2, a Bayesian model is used to integrate and analyze gene expression data containing spatial location information from uterine tissue lesion samples and uterine tissue control samples with single-cell data from the same menstrual cycle to obtain multiple cell subpopulations containing spatial location information, including:
[0013] First, single-cell data of uterine tissue in the same menstrual cycle as the uterine tissue lesion sample were obtained, and the obtained single-cell data of uterine tissue was standardized by using the SCTransform algorithm to obtain standardized single-cell data of uterine tissue.
[0014] The standardized uterine single-cell data were then integrated using the CCA analysis algorithm to obtain integrated uterine single-cell data.
[0015] Then, the PCA algorithm is used to perform linear dimensionality reduction on the integrated uterine single-cell data to obtain PCA dimensionality-reduced data of uterine single cells;
[0016] Then, the UMAP nonlinear dimensionality reduction algorithm is used to perform nonlinear dimensionality reduction on the PCA dimensionality reduction data of uterine single cells to obtain the UMAP dimensionality reduction data of uterine single cells;
[0017] Then, the K-nearest neighbor algorithm is used to perform cluster analysis on the UMAP dimensionality reduction data of uterine single cells to obtain uterine single cell cluster data;
[0018] Then, based on the expression of classic genes of cell type, the uterine single-cell cluster data is defined to obtain the type of each uterine single cell in the uterine single-cell data;
[0019] Then, negative binomial regression was used to analyze the uterine single-cell data to obtain the transcriptional characteristics of each cell type in the uterine single-cell data;
[0020] Based on the transcriptional features of each cell type in the single-cell uterine data, a Bayesian model was used to deconvolve the gene expression data containing spatial location information to obtain multiple cell subpopulations containing spatial location information.
[0021] Furthermore, the cell types in the uterine single-cell data include at least one of the following: epithelial cells, matrix fibroblasts, supporting cells, endothelial cells, or immune cells.
[0022] Furthermore, the types of cell subpopulations that contain spatial location information include: epithelial cells, matrix fibroblasts, supporting cells, endothelial cells, or immune cells.
[0023] Furthermore, after the step of integrating and analyzing gene expression data containing spatial location information from uterine tissue lesion samples and uterine tissue control samples using a Bayesian model with uterine single-cell data from the same menstrual cycle in step S2, and before the step of obtaining multiple cell subpopulations containing spatial location information in step S2 and defining each obtained cell subpopulation as a target subpopulation, the method for determining the spatially specific regulatory molecules of adenomyosis further includes: using Pearson correlation analysis to obtain the correlation of differentially expressed genes in different cell types from spatial transcriptome data and single-cell data to determine cell subpopulations containing spatial location information, and then matching the cell subpopulations containing spatial location information with preset cell gene markers to define each obtained cell subpopulation containing spatial location information as a target subpopulation.
[0024] Furthermore, in step S3, the significance of the interaction probabilities and cell interaction probabilities between each target subpopulation is obtained using the zero-distribution algorithm. The spatial confidence place of the target subpopulation with a significance value less than 0.05 and exhibiting typical adenomyosis morphological characteristics is defined as the adenomyosis microenvironment, including:
[0025] The expression levels of receptors and ligands in each cell of each target subpopulation were calculated. Based on the expression levels of receptors and ligands in each cell, combined with the receptor-ligand pairing information in the receptor-ligand database and the spatial location of the cells, receptor-ligand analysis based on the zero-distribution algorithm was used to calculate the interaction probability between cells at the spatial location. The interaction probability was then mapped onto the tissue images contained in the spatial transcriptome data to obtain the cell interaction probability distribution and the significance of the cell interaction probability for different spatial locations in uterine tissue lesion samples and uterine tissue control samples. Regions with a cell interaction probability significance correction value of less than 0.05 and with typical morphological features of adenomyosis were identified as the adenomyosis microenvironment.
[0026] Furthermore, step S4 includes:
[0027] We constructed a cell communication model that includes ligands, receptors, and target genes using ligand-receptor, ligand-target gene, and receptor-target gene information from public databases.
[0028] The Wilcoxon test was used to calculate the differential gene expression in different cell types within the adenomyosis microenvironment. Genes with a corrected p-value less than 0.05 and a fold change greater than 1.2 were defined as differentially expressed genes. These differentially expressed genes were then input into the cell communication model, and the NicheNet algorithm was used to calculate the activity of ligand signals received by different cell types within the adenomyosis microenvironment. The top 50 most active ligand signals received by different cell types within the adenomyosis microenvironment were defined as significant communication signals.
[0029] The significant communication signals were then searched in databases and literature to obtain their functional characteristics. The functional characteristics of the significant communication signals were compared with the pathogenic biological process of adenomyosis, and the ligands of the microenvironment-specific communication signals that promote adenomyosis were identified as the spatially specific regulatory molecules of adenomyosis.
[0030] Furthermore, the significant communication signal includes at least one of the following: IHH signaling pathway, WNT signaling pathway, TGFβ signaling pathway, HIF1-α signaling pathway, estrogen receptor signaling pathway, oxytocin signaling pathway, ILK signaling pathway, and integrin signaling pathway.
[0031] The present invention also provides a system for identifying spatially specific regulatory molecules of adenomyosis, comprising:
[0032] The transcriptome sequencing module is used to obtain uterine tissue lesion samples and uterine tissue control samples. Spatial transcriptome sequencing technology is used to perform mRNA sequencing detection on uterine tissue lesion samples and uterine tissue control samples to obtain gene expression data containing spatial location information in uterine tissue lesion samples and uterine tissue control samples.
[0033] The cell subpopulation detection module has a built-in Bayesian model, which is used to integrate and analyze gene expression data containing spatial location information in uterine tissue lesion samples and uterine tissue control samples with uterine single cell data in the same menstrual cycle to obtain multiple cell subpopulations containing spatial location information, and define each obtained cell subpopulation as the target subpopulation.
[0034] The adenomyosis microenvironment detection module has a built-in zero-distribution algorithm model, which is used to obtain the interaction probability distribution and the significance of cell interaction probability between various target subgroups using the zero-distribution algorithm. The spatial location of the target subgroup with a cell interaction probability significance correction value of less than 0.05 and typical adenomyosis morphological characteristics is defined as the adenomyosis microenvironment.
[0035] The adenomyosis space-specific regulatory molecule detection module incorporates the NicheNet algorithm model. It utilizes the NicheNet algorithm to perform cell communication analysis on ligands, receptors, and target genes in the adenomyosis microenvironment, screening out cell communication signals with significantly differential expression of ligands, receptors, and target genes in the adenomyosis microenvironment. The ligands of the communication signals that promote the pathogenesis of adenomyosis in the screened cell communication signals are identified as adenomyosis space-specific regulatory molecules.
[0036] Furthermore, the cell subpopulation detection module also includes:
[0037] The SCTransform algorithm model is used to standardize uterine single-cell data to obtain standardized uterine single-cell data.
[0038] The CCA analysis algorithm model is used to integrate standardized uterine single-cell data to obtain integrated uterine single-cell data.
[0039] The PCA algorithm model is used to perform linear dimensionality reduction on the integrated uterine single-cell data to obtain PCA dimensionality-reduced data of uterine single cells.
[0040] The UMAP nonlinear dimensionality reduction algorithm model is used to perform nonlinear dimensionality reduction on PCA dimensionality reduction data of uterine single cells to obtain UMAP dimensionality reduction data of uterine single cells.
[0041] The K-nearest neighbor algorithm model is used to perform cluster analysis on UMAP dimensionality reduction data of uterine single cells to obtain uterine single cell cluster data;
[0042] The Uterine Single-Cell Type Definition Submodule is used to define the uterine single-cell cluster data and obtain the type of each uterine single cell in the uterine single-cell data;
[0043] A negative binomial regression model was used to analyze uterine single-cell data to obtain the transcriptional characteristics of each cell type in the uterine single-cell data.
[0044] The present invention also provides an electronic device, the electronic device comprising: a processor and a memory;
[0045] The memory is used to store computer programs;
[0046] The processor is configured to execute a computer program stored in the memory to cause the electronic device to perform the method described above for identifying spatially specific regulatory molecules of adenomyosis.
[0047] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by an electronic device, implements the method for determining spatially specific regulatory molecules of adenomyosis as described in any of the preceding claims.
[0048] As described above, the method, system, medium, and electronic device for determining the spatially specific regulatory molecules of adenomyosis according to the present invention have the following beneficial effects:
[0049] (1) Compared with the prior art, the present invention addresses the problem that the molecular characteristics of the microenvironment of adenomyosis, a typical disease with spatial specificity, are unclear and that existing research methods are difficult to conduct a comprehensive study of the spatial specific regulatory molecules of adenomyosis from the perspective of the microenvironment. It provides a method, system, medium and electronic device based on spatial transcriptome sequencing technology to determine the spatial specific regulatory molecules of adenomyosis, and solves the problem that existing research methods are difficult to fully utilize spatial location information to analyze the molecular characteristics of different development processes of adenomyosis.
[0050] (2) This invention helps to solve the problem that the spatially specific regulatory molecules of adenomyosis cannot be accurately found due to the use of microarrays, transcriptomes, and single-cell transcriptomes that lose spatial information of tissue samples.
[0051] (3) This invention uses spatial transcriptome sequencing to sequence the expression of uterine tissue samples of adenomyosis. This can make full use of the spatial information of the tissue to reflect the differences in gene expression at different spatial locations of the disease, thereby more accurately reflecting the specific regulatory information of the microenvironment. At the same time, based on the deconvolution algorithm of combined single-cell data and cell communication analysis, spatially specific regulatory molecules related to the microenvironment of adenomyosis and different development processes of adenomyosis can be accurately identified. Attached Figure Description
[0052] Figure 1 The flowchart shown is a method for determining spatially specific regulatory molecules of adenomyosis according to an embodiment of the present invention.
[0053] Figure 2 The diagram shown is a structural schematic of the system for determining the spatially specific regulatory molecules of adenomyosis according to an embodiment of the present invention. Detailed Implementation
[0054] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0055] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. The illustrations only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0056] See Figure 1 and Figure 2 The following embodiments of the present invention provide a method, system, medium, and electronic device for identifying spatially specific regulatory molecules in adenomyosis. Compared with the prior art, the present invention addresses the problem that the molecular characteristics of the microenvironment of adenomyosis, a typical spatially specific disease, are unclear, and that existing research methods are difficult to comprehensively study the spatially specific regulatory molecules of adenomyosis from the perspective of the microenvironment. It provides a method, system, medium, and electronic device based on spatial transcriptome sequencing technology to identify spatially specific regulatory molecules in adenomyosis, solving the problem that existing research methods cannot fully utilize spatial location information to analyze the molecular characteristics of different developmental stages of adenomyosis. The present invention helps to solve the problem that the use of microarrays, transcriptomes, and single-cell transcriptomes that lose spatial information of tissue samples cannot accurately identify spatially specific regulatory molecules in adenomyosis. The present invention uses spatial transcriptome sequencing to sequence the expression of uterine tissue samples with adenomyosis, which can fully utilize the spatial information of the tissue to reflect the differences in gene expression at different spatial locations of the disease, thereby more accurately reflecting the specific regulatory information of the microenvironment. Simultaneously, based on the deconvolution algorithm combined with single-cell data and cell communication analysis, spatially specific regulatory molecules related to the microenvironment of adenomyosis and different developmental stages of adenomyosis can be accurately identified.
[0057] The technical solutions of the present invention will now be described in detail with reference to the accompanying drawings.
[0058] like Figure 1 As shown, in one embodiment, the present invention provides a method for identifying spatially specific regulatory molecules of adenomyosis, the specific steps of which are as follows:
[0059] Step S1: Obtain uterine tissue lesion samples and uterine tissue control samples. Use spatial transcriptome sequencing technology to perform mRNA sequencing detection on uterine tissue lesion samples and uterine tissue control samples to obtain gene expression data containing spatial location information in uterine tissue lesion samples and uterine tissue control samples.
[0060] Step S2: Using a Bayesian model, gene expression data containing spatial location information from uterine tissue lesion samples and uterine tissue control samples are integrated and analyzed with uterine single-cell data from the same menstrual cycle to obtain multiple cell subpopulations containing spatial location information, and each obtained cell subpopulation is defined as the target subpopulation.
[0061] Step S3: Use the zero distribution algorithm to obtain the interaction probability distribution and the significance of cell interaction probability among each target subpopulation, and define the spatial location of the target subpopulation with a cell interaction probability significance value of less than 0.05 and with typical adenomyosis morphological characteristics as the adenomyosis microenvironment.
[0062] Step S4: Using the NicheNet algorithm, cell communication analysis is performed on ligands, receptors, and target genes in the adenomyosis microenvironment. Cell communication signals with significantly differential expression of ligands, receptors, and target genes in the adenomyosis microenvironment are screened out, and the ligands of the cell communication signals that promote the pathogenesis of adenomyosis are identified as space-specific regulatory molecules of adenomyosis.
[0063] In step S1, the gene expression data containing spatial location information obtained from uterine tissue lesion samples and uterine tissue control samples include:
[0064] After wiping away surface mucus with lint-free paper, uterine tissue lesion samples and uterine tissue control samples were embedded in OCT, sectioned, stained with hematoxylin and eosin (HE), and scanned to obtain sample images of the uterine tissue lesion samples and uterine tissue control samples. Based on the sample images of the uterine tissue lesion samples and uterine tissue control samples, uterine tissue with typical morphological characteristics of adenomyosis was selected from the uterine tissue lesion samples, and uterine tissue far from the uterine fibroid site was selected from the uterine tissue control samples for spatial transcriptome sequencing to obtain gene expression data containing spatial location information of the uterine tissue lesion samples and uterine tissue control samples.
[0065] The spatial transcriptome sequencing technology is an existing technology, documented in many publications, such as patent document CN116564419B. Spatial transcriptome sequencing utilizes techniques such as barcode labeling to achieve high-throughput transcriptome sequencing, obtaining transcriptome information including the spatial location of tissues. By using frozen tissue sections and high-throughput in-situ detection of RNA expression patterns and spatial distribution in different regions of tissue sections, gene expression data can be integrated with immunochemical staining images of tissue sections of interest. This allows for the localization of gene expression information from different cells within the tissue to their original spatial location, enabling direct observation of differences in gene expression in functional regions of different parts of the tissue. The specific process of spatial transcriptome sequencing is as follows: Fresh tissue is taken, frozen, sectioned, and subjected to HE staining and imaging to optimize and determine whether the section covers the target region; the tissue section is placed on a glass slide containing RNA-binding capture probes, and fixed and permeabilized to release mRNA from the cells and bind to the corresponding capture probes; cDNA is synthesized and sequencing library is prepared using the captured RNA as a template; the prepared sequencing library is subjected to next-generation high-throughput short-read sequencing; and the HE results are used to determine which genes are expressed, their expression levels, and the spatial location information of these genes.
[0066] Uterine tissue lesion samples were obtained from patients who underwent total hysterectomy for adenomyosis, while uterine tissue control samples were obtained from patients who underwent total hysterectomy for uterine fibroids. It should be noted that, in order to preserve the structural characteristics of the endometrium and myometrium, the samples used for spatial transcriptome sequencing were whole uterine samples from patients who underwent total hysterectomy. Whole uterine samples from healthy individuals are difficult to obtain. Therefore, in clinical practice, whole uterine samples from patients who underwent total hysterectomy for uterine fibroids are often used as control samples for adenomyosis.
[0067] Uterine tissue control samples should be taken at the same time of the menstrual cycle as uterine tissue lesion samples, and the selection of uterine tissue control samples should follow the sample inclusion criteria and sample exclusion criteria.
[0068] The sample inclusion criteria were: a history of natural pregnancy and childbirth, age 50 or younger, regular menstrual cycles, and a body mass index of 18.5-23.9 kg / m². 2 Furthermore, they have no history of smoking and no history of infectious diseases.
[0069] The exclusion criteria for the samples were: a history of hormone use within 3 months, or a history of intrauterine device use within 2 months, or a history of tumors, or a history of endometriosis, or a history of hydrosalpinx, or a history of endometritis.
[0070] By defining the inclusion and exclusion criteria for samples, the uterine tissue control samples are made comparable and are not affected by other pathological or iatrogenic factors.
[0071] Uterine tissue samples with typical morphological characteristics of adenomyosis were selected from uterine tissue lesion samples, including: endometrial and myometrial samples from the endometrial invasion sites of adenomyosis, superficial adenomyosis lesion samples, deep adenomyosis lesion samples, and endometrial and myometrial samples from non-fibroid sites of control uterine fibroids.
[0072] The site of endometrial invasion refers to the area where the junction between the endometrium and the myometrium is discontinuous, and the endometrium is surrounded by the myometrium in three directions.
[0073] The endometrial and myometrial samples from the site of endometrial invasion in adenomyosis are tissues containing the entire endometrial layer and showing endometrial invasion into the myometrium; the superficial lesion samples of adenomyosis do not contain the endometrium but contain lesions near the endometrial-myometrial junction; the deep lesion samples of adenomyosis do not contain the endometrium but contain lesions away from the endometrial-myometrial junction; the endometrial and myometrial samples from the non-fibroid sites of the control uterine fibroids are tissues containing the entire endometrial layer and adjacent myometrium away from the fibroids.
[0074] mRNA sequencing reflects gene expression during the transcription process of the central dogma. mRNA is the product of gene transcription and is an intermediate step in the transformation of genetic material from DNA to proteins that perform biological functions. Detecting the abundance of mRNA in tissues can be used to analyze changes in gene expression, the correlation between different genes, the mutual influence of gene expression, and genes that occupy a core position in biological function. It is of great significance in the study of pathophysiological mechanisms.
[0075] In spatial transcriptome sequencing, each site with a specific spatial probe, i.e., a spatial minimum marker unit, may contain multiple cells. It is necessary to integrate single-cell data with similar transcriptional characteristics to determine the cell type of each spatial minimum marker unit. In this embodiment, the spatial transcriptome sequencing platform used is the 10x Genomics Visium spatial transcriptome sequencing platform. The diameter of a spatial minimum marker unit is 55 micrometers, and it typically contains 1-30 cells.
[0076] By using spatial transcriptome sequencing, we can obtain gene expression data containing spatial location information by sequencing four types of tissue samples: endometrial and myometrial samples from the adenomyosis invasion site, superficial adenomyosis lesions, deep adenomyosis lesions, and endometrial and myometrial samples from non-fibroid sites in the control uterine fibroids.
[0077] The Bayesian model used in step S2 is existing technology. The Bayesian model is a common method for implementing deconvolution, documented in many publications, such as patent document CN106599427B. The Bayesian model is a method that infers probabilities using pre-existing probabilities and sample data. Using the transcriptional features of different cell types in spatial transcriptome data and single-cell data as input data, the cell type abundance of each spatial minimum labeled unit is calculated. After obtaining the cell composition of each spatial minimum labeled unit, the cell subpopulation containing spatial information can be determined based on the cell proportions. The cell type with the highest cell proportion in each spatial minimum labeled unit is determined as the cell type of that spatial minimum labeled unit.
[0078] In step S2, a Bayesian model is used to integrate gene expression data containing spatial location information from uterine tissue lesion samples and uterine tissue control samples with single-cell data from the same menstrual cycle to obtain multiple cell subpopulations containing spatial location information, including:
[0079] First, single-cell data of uterine tissue in the same menstrual cycle as the uterine tissue lesion sample were obtained, and the obtained single-cell data of uterine tissue was standardized by using the SCTransform algorithm to obtain standardized single-cell data of uterine tissue.
[0080] The standardized uterine single-cell data were then integrated using the CCA analysis algorithm to obtain integrated uterine single-cell data.
[0081] Then, the PCA algorithm is used to perform linear dimensionality reduction on the integrated uterine single-cell data to obtain PCA dimensionality-reduced data of uterine single cells;
[0082] Then, the UMAP nonlinear dimensionality reduction algorithm is used to perform nonlinear dimensionality reduction on the PCA dimensionality reduction data of uterine single cells to obtain the UMAP dimensionality reduction data of uterine single cells;
[0083] Then, the K-nearest neighbor algorithm is used to perform cluster analysis on the UMAP dimensionality reduction data of uterine single cells to obtain uterine single cell cluster data;
[0084] Then, based on the expression of classic genes of cell type, the uterine single-cell cluster data is defined to obtain the type of each uterine single cell in the uterine single-cell data;
[0085] Then, negative binomial regression was used to analyze the uterine single-cell data to obtain the transcriptional characteristics of each cell type in the uterine single-cell data;
[0086] Based on the transcriptional features of each cell type in the single-cell uterine data, a Bayesian model was used to deconvolve the gene expression data containing spatial location information to obtain multiple cell subpopulations containing spatial location information.
[0087] The single-cell data of the uterus in the same menstrual cycle can be obtained by single-cell transcriptome sequencing of uterine tissue lesion samples and uterine tissue control samples, or from published single-cell databases of adenomyosis uterine tissue and normal uterine tissue in the same menstrual cycle. Single-cell transcriptome sequencing is an existing sequencing technology and is described in many patents, such as patent document with publication number CN117079726B.
[0088] In this embodiment, data standardization, integration, dimensionality reduction, cluster analysis, and definition of uterine single-cell data are performed using the R package Seurat. Seurat is a commonly used R package based on the R language for quality control and analysis of single-cell transcriptome and spatial transcriptome data, which includes functions such as data standardization, data integration, dimensionality reduction, clustering, and differential expression analysis.
[0089] Data standardization is a method to eliminate systemic differences in each cell caused by variations in transcript capture and amplification efficiency and sequencing depth, making cell gene expression more comparable and unaffected by technological differences.
[0090] The SCTransform algorithm is a commonly used data standardization algorithm, documented in numerous publications, including patent document CN116469473B. SCTransform is highly effective at correcting sequencing depth and can also be used to correct for the influence of factors such as mitochondria. Using the SCTransform algorithm within the Seurat package, comparable standardized single-cell data can be obtained.
[0091] Data integration is a method of combining multiple single-cell data into a single processing object. The purpose of data integration is to remove batch effects caused by differences in sequencing platforms, sequencing times, etc.; batch effects are technical differences that are unrelated to biological processes.
[0092] The CCA analysis algorithm is an existing technology, a correction method used to combine multiple datasets, and it is documented in many documents, such as the patent document with publication number CN113687083B. Data integration based on the CCA analysis algorithm enables joint analysis of single-cell data from multiple experiments.
[0093] By integrating and analyzing standardized single-cell data, integrated single-cell data that has been freed from batch effects can be obtained.
[0094] In single-cell data, the expression of each gene constitutes a dimension of variation. Dimensionality reduction is a method of selecting genes that better represent the overall differences and using as few dimensions as possible to display the true structure of the data. Linear dimensionality reduction and nonlinear dimensionality reduction are two commonly used methods. Linear dimensionality reduction maps high-dimensional data to a low-dimensional space through linear transformations, using dozens of principal components to reflect the linear changes in the data. Principal components refer to differentially expressed genes that play a major role in the features of single-cell data. Nonlinear dimensionality reduction maps high-dimensional data to a low-dimensional space through nonlinear transformations, which can better preserve the structure of the data.
[0095] PCA (Programmatical Conversion) is one of the most commonly used dimensionality reduction methods. Its goal is to map high-dimensional data to a low-dimensional space through a linear projection, aiming to maximize the variance of the data in the projected dimension, thereby using fewer dimensions while retaining more dimensions of the original data. PCA is an existing technology and is documented in many documents, such as the patent document with publication number CN117637080B.
[0096] The UMAP method is a nonlinear dimensionality reduction algorithm suitable for processing data with inconsistent shapes. It offers advantages such as preserving local structure, high efficiency, and interpretability, and can be applied to scenarios including data visualization, data preprocessing, clustering, and classification. The UMAP method learns the structural features of a high-dimensional space to find a low-dimensional representation of those features. The UMAP method is existing technology and is documented in many publications, such as the patent document with publication number CN116864012B.
[0097] By performing linear and nonlinear dimensionality reduction on the integrated single-cell data, we can obtain a low-dimensional spatial distribution that reflects the main changes and differences in the single-cell data, as well as differentially expressed genes that play a major role in the single-cell data.
[0098] For example, if a single-cell dataset detects 20,000 genes, the cells can form 20,000 dimensions. By using PCA linear dimensionality reduction, the top few dozen genes that play a major role in the differences between cells can be obtained, reducing the 20,000 dimensions to less than 50. By using UMAP nonlinear dimensionality reduction, the similarity of gene expression between cells can be represented in two-dimensional space, reducing the 20,000 dimensions to 2 dimensions.
[0099] The K-Nearest Neighbors (KNN) algorithm is an instance-based learning algorithm used for classification and regression tasks. Its core idea is that if a sample's k nearest neighbors in the feature space mostly belong to a certain category, then the sample also belongs to that category. KNN is an existing technology and a common single-cell data clustering method, documented in many publications, such as the patent document with publication number CN111292807B. KNN is used to analyze the principal component information obtained from dimensionality reduction, grouping cells with similar gene expression patterns. Through cluster analysis, cluster information of cells with similar gene expression characteristics can be obtained.
[0100] Defining uterine single-cell cluster data based on the expression of classic genes of cell types is an existing technique and a conventional definition method in the field. For example: if a cluster shows high expression of the genes EPCAM, KRT8, and WFDC2, then the cluster is identified as epithelial cells; if a cluster shows high expression of the genes IGF1, PCOLCE, and SFRP1, then the cluster is identified as stromal fibroblasts; if a cluster shows high expression of the genes ACTA2 and MYH11, then the cluster is identified as supporting cells; if a cluster shows high expression of the genes PECAM1 and VWF, then the cluster is identified as endothelial cells; and if a cluster shows high expression of the gene PTPRC, then the cluster is identified as immune cells.
[0101] Negative binomial regression is an existing technique, documented in numerous publications, such as patent publication CN114974435B. This embodiment uses the Python package `cell2location` to perform negative binomial regression analysis to obtain the transcriptional features of each cell type in single-cell data of uterine cell types. `cell2location` is a Python package used to map the transcriptional features of single-cell cell types to spatial transcriptome datasets; negative binomial regression is a generalized linear model and is the default statistical method for calculating transcriptional features in the `cell2location` package. Through negative binomial regression analysis using `cell2location`, the transcriptional features of different cell types in single-cell data can be obtained.
[0102] Transcriptional signatures refer to the unique gene expression patterns of each cell type compared to other cell types, manifested in consistent gene expression levels among cells of the same type but distinct from those of other cell types. Uterine single-cell data must include at least one of the following cell types: epithelial cells, stromal fibroblasts, supporting cells, endothelial cells, or immune cells.
[0103] Cellular subpopulations containing spatial location information include: epithelial cells, matrix fibroblasts, supporting cells, endothelial cells, or immune cells.
[0104] After the step S2, which integrates and analyzes gene expression data containing spatial location information from uterine tissue lesion samples and uterine tissue control samples using a Bayesian model with uterine single-cell data from the same menstrual cycle, and before the step S2, which obtains multiple cell subpopulations containing spatial location information and defines each obtained cell subpopulation as a target subpopulation, the method for determining the spatially specific regulatory molecules of adenomyosis further includes: using Pearson correlation analysis to obtain the correlation of differentially expressed genes in different cell types from spatial transcriptome data and single-cell data to determine cell subpopulations containing spatial location information; and then matching the cell subpopulations containing spatial location information with preset cell gene markers to define each obtained cell subpopulation containing spatial location information as a target subpopulation.
[0105] Correlation analysis refers to the analysis of two or more variable elements to measure the degree of closeness between them. In other words, it involves performing correlation analysis on gene expression data of cell subpopulations and single cells containing spatial location information. The stronger the correlation, the more accurate the definition of the cell subpopulation containing spatial location information. Pearson correlation analysis is a commonly used correlation analysis method, documented in many publications, such as patent document CN117454212B. Pearson correlation analysis is used to measure the relationship between normally distributed data, with correlation coefficients ranging from -1 to +1.
[0106] Correlation analysis was performed on the spatial smallest marker unit and single-cell transcriptome data for spatial transcriptome data of cell types such as epithelial cells, matrix fibroblasts, supporting cells, endothelial cells, or immune cells. The closer the absolute value of the correlation coefficient is to 1, the more correlated the spatial transcriptome data and single-cell transcriptome data are.
[0107] In this embodiment, the preset cell gene marker is a marker gene recognized in the art for cell type characteristic expression.
[0108] For example, by presenting the expression of all marker genes across all cell types in the spatial transcriptome, the data can be compared to determine whether different cell types in the spatial transcriptome data exhibit characteristic expression of marker genes. If the spatial transcriptome data shows high expression of the following three genes: EPCAM, KRT8, and WFDC2 in epithelial cells; high expression of the following three genes: IGF1, PCOLCE, and SFRP1 in stromal fibroblasts; high expression of the following two genes: ACTA2 and MYH11 in supporting cells; high expression of the following two genes: PECAM1 and VWF in endothelial cells; and high expression of the following gene: PTPRC in immune cells, then the cell subpopulation containing spatial information is considered representative.
[0109] In step S3, the significance of the interaction probabilities and cell interaction probabilities between each target subpopulation is obtained using the zero-distribution algorithm. The spatial confidence place of the target subpopulation with a significance correction value less than 0.05 and exhibiting typical adenomyosis morphological characteristics is defined as the adenomyosis microenvironment, including:
[0110] The expression levels of receptors and ligands in each cell of each target subpopulation were calculated. Based on the expression levels of receptors and ligands in each cell, combined with the receptor-ligand pairing information in the receptor-ligand database and the spatial location of the cells, receptor-ligand analysis based on the zero-distribution algorithm was used to calculate the interaction probability between cells at the spatial location. The interaction probability was then mapped onto the tissue images contained in the spatial transcriptome data to obtain the cell interaction probability distribution and the significance of the cell interaction probability for different spatial locations in uterine tissue lesion samples and uterine tissue control samples. Regions with a cell interaction probability significance correction value of less than 0.05 and with typical morphological features of adenomyosis were identified as the adenomyosis microenvironment.
[0111] This embodiment uses the Python software stLearn for receptor-ligand analysis. It should be noted that stLearn is a Python package used to calculate receptor-ligand interactions, distribution, and functional analysis between the smallest spatially labeled units in spatial transcriptome data. Cell-cell interaction probabilities are calculated based on the co-expression of paired receptor-ligands at different spatial locations within a tissue, combined with bioinformatics databases, to determine the probability of cell-cell interactions at different tissue spatial locations.
[0112] The null distribution algorithm is a common method for receptor-ligand analysis, documented in numerous publications, including patent document CN112466403B. This algorithm infers the spatial distribution probability of ligand-receptor interactions by statistically analyzing the expression and pairing of receptors and ligands in the smallest spatially labeled units of the spatial transcriptome, combined with bioinformatics databases and spatial transcriptome HE images. Ligand-receptor analysis can help understand cell interactions at different locations within tissues.
[0113] After obtaining multiple cell subpopulations containing spatial information, receptor-ligand analysis can be used to determine the probability of cell interaction in tissue space, thereby identifying the regions with spatial distribution related to disease development and high cell interaction probability as the adenomyosis microenvironment.
[0114] Using the method in this embodiment, high-frequency cell interactions were found at the junction of the endometrium and myometrium in endometrial and myometrial samples from the adenomyosis invasion site of adenomyosis and the endometrium and myometrium in control uterine fibroid samples from the non-fibroid site. This suggests that the junction of the endometrium and myometrium in endometrial and myometrial samples from the adenomyosis invasion site of adenomyosis and the endometrium and myometrium in control uterine fibroid samples from the non-fibroid site represents the invasion microenvironment of adenomyosis. High-frequency cell interactions were also found at the junction of the lesions and myometrium in superficial and deep adenomyosis lesions, suggesting that the junction of the lesions and myometrium in superficial and deep adenomyosis lesions represents the lesion microenvironment of adenomyosis.
[0115] Step S4 includes:
[0116] We constructed a cell communication model that includes ligands, receptors, and target genes using ligand-receptor, ligand-target gene, and receptor-target gene information from public databases.
[0117] The Wilcoxon test was used to calculate the differential gene expression in different cell types within the adenomyosis microenvironment. Genes with a corrected p-value less than 0.05 and a fold change greater than 1.2 were defined as differentially expressed genes. These differentially expressed genes were then input into the cell communication model, and the NicheNet algorithm was used to calculate the activity of ligand signals received by different cell types within the adenomyosis microenvironment. The top 50 most active ligand signals received by different cell types within the adenomyosis microenvironment were defined as significant communication signals (i.e., the significantly differentially expressed cell communication signals described in step S4).
[0118] The significant communication signals were then searched in databases and literature to obtain their functional characteristics. The functional characteristics of the significant communication signals were compared with the pathogenic biological process of adenomyosis, and the ligands of the microenvironment-specific communication signals that promote adenomyosis were identified as the spatially specific regulatory molecules of adenomyosis.
[0119] This example uses the Nichenet package in R for cell communication analysis. It should be noted that the Nichenet package is an R package used to calculate cell interactions in different spatial microenvironments. Its key feature is the integration of databases to comprehensively calculate the expression of ligands, receptors, and target genes. The public databases used to construct the cell communication model are those linked to by the Nichenet package.
[0120] The Wilcoxon test is an existing technique and a non-parametric test method. It is a commonly used statistical test method for differential expression analysis of single-cell data and is documented in many documents, such as the patent document with publication number CN114974435B.
[0121] Cell communication analysis involves statistically analyzing the differential and characteristic expression of receptors, ligands, and target genes in the microenvironment across different cell types, and combining this data with bioinformatics databases to infer cell communication signals characteristic of the microenvironment. Cell communication analysis can help understand the communication signals of significant changes in the microenvironment and analyze the regulatory molecules within it.
[0122] The NicheNet algorithm is a commonly used method for cell communication analysis, documented in numerous publications, including patent CN112466403B. Ligand signal activity is the primary outcome of the NicheNet algorithm. Based on a cell communication model and differentially expressed genes in different cell types, it calculates the activity of ligand signals received by different receptor cells by integrating ligand expression in ligand cells, receptor expression in receptor cells, and target gene expression in receptor cells. The cell communication model is constructed from information on ligand-receptor pairs, receptor-target gene pathways, and ligand-target gene regulatory networks in public databases.
[0123] Communication signals consist of ligands, receptors, and target genes. The ligand signal is the initiation step, and the cell that emits the ligand signal is the ligand cell. The receptor is the medium through which the cell receives the ligand signal, and the target gene represents the transcriptional changes that occur after the cell receives the ligand signal. The cell that receives the ligand signal through the receptor and expresses the target gene is the receptor cell. The saliency of a communication signal is reflected in the overall differences among the ligand, receptor, and target gene. The functional characteristics of significant communication signals are the functional descriptions of the ligand in the GeneCard database and the functional enrichment characteristics of the target gene in the Gene Ontology database.
[0124] Differential gene expression in different cell types within the adenomyosis microenvironment reflects the expression of their ligands, receptors, and target genes. The same cell can be either a ligand cell or a receptor cell depending on the specific receptor and ligand interaction.
[0125] Ligands are the initiation link of cell communication signals and play a one-to-one or one-to-many role in the regulation of cell function. After identifying the microenvironment of adenomyosis, cell communication analysis can reveal the characteristic ligands, receptors, and target gene expression of different cell types in the microenvironment of adenomyosis. Thus, the ligands that promote the communication signals of the pathogenesis of adenomyosis can be identified as space-specific regulatory molecules of adenomyosis.
[0126] The method in this embodiment found that in the invasion microenvironment of adenomyosis, CCL14, IHH, TRH, CCL19, CXCL12 related to immune regulation and IHH, LAMC2, NUCB2, TNFSF10, MDK, COL1A1, CXCL12, CCL5, and CCL19 related to epithelial-mesenchymal transition are spatially specific regulatory molecules of adenomyosis. Furthermore, in the lesion microenvironment of adenomyosis, fibrosis-related A2M, ADAM9, FGF9, GAS6, INHBB, LAMA1, CCL21, CTSG, PDGFB, SLIT2, LAMA1, PDGFC, SLIT2, TGFB3, TIMP1, WNT4, WNT5A, HSPG2, THBS4, TNFSF12, and TGFB1 are spatially specific regulatory molecules of adenomyosis.
[0127] The pathogenic biological processes of adenomyosis are the classic biological processes reported and associated with the pathogenesis of adenomyosis, including: cell proliferation, angiogenesis, epithelial-mesenchymal transition, invasion, migration, immune regulation, fibroblast transformation into myofibroblasts, and fibrosis. By comparing the functional characteristics of significant communication signals with the pathogenic biological processes of adenomyosis, microenvironment-specific adenomyosis-related communication signals can be obtained.
[0128] Signaling pathways that promote the pathogenesis of adenomyosis refer to those reported in the literature that promote the pathogenesis of adenomyosis. These pathways include at least one of the following: IHH signaling pathway, WNT signaling pathway, TGFβ signaling pathway, HIF1-α signaling pathway, estrogen receptor signaling pathway, oxytocin signaling pathway, ILK signaling pathway, and integrin signaling pathway.
[0129] After identifying the ligands corresponding to the signals in cell communication signals that promote the pathogenic biological processes of adenomyosis as the spatially specific regulatory molecules of adenomyosis, the spatially specific regulatory molecules of adenomyosis can be intervened on through organoid platforms, primary cell models of adenomyosis, and mouse models of adenomyosis to identify potential therapeutic targets.
[0130] Organoids are 3D in vitro cell culture systems that can replicate the complex spatial morphology of tissues, contain multiple cell types, and simulate the interactions between cells and between cells and the matrix. They can better simulate the physiological and pathological processes of human tissues.
[0131] Primary cell models of adenomyosis refer to the process of obtaining endometrial epithelial cells, stromal cells, and uterine smooth muscle cells that retain certain disease characteristics of adenomyosis by performing tissue cutting, tissue digestion, cell filtration and separation, cell plating, and cell culture on the endometrium of patients with adenomyosis who have undergone total hysterectomy. Through intervention, phenotypic identification, and downstream gene expression detection, these models provide direction for exploring the mechanisms of adenomyosis.
[0132] Adenomyosis mouse models refer to mouse models of adenomyosis induced by mechanical or hormonal means. These models provide an in vivo animal model for exploring the mechanisms of adenomyosis and studying the effects of interventions.
[0133] By using organoid platforms, primary cell models of adenomyosis, and mouse models of adenomyosis, interventions can be made on spatially specific regulatory molecules of adenomyosis to discover their impact on key biological processes and incidence of adenomyosis, explore drug potential, and contribute to the development of new adenomyosis targets.
[0134] The scope of protection of the method for determining the spatially specific regulatory molecules of adenomyosis described in the embodiments of the present invention is not limited to the order of steps listed in this embodiment. Any solution achieved by adding, subtracting, or replacing steps in the prior art based on the principles of the present invention is included within the scope of protection of the present invention.
[0135] This invention also provides a system for determining spatially specific regulatory molecules of adenomyosis. This system can implement the method for determining spatially specific regulatory molecules of adenomyosis described in this invention. However, the apparatus for implementing the method for determining spatially specific regulatory molecules of adenomyosis described in this invention includes, but is not limited to, the structure of the system for determining spatially specific regulatory molecules of adenomyosis listed in this embodiment. All structural modifications and substitutions of the prior art made in accordance with the principles of this invention are included within the protection scope of this invention.
[0136] like Figure 2 As shown, embodiments of the present invention provide a system for identifying spatially specific regulatory molecules of adenomyosis, characterized in that it comprises:
[0137] Transcriptome sequencing module 21 is used to obtain uterine tissue lesion samples and uterine tissue control samples. Spatial transcriptome sequencing technology is used to perform mRNA sequencing detection on uterine tissue lesion samples and uterine tissue control samples to obtain gene expression data containing spatial location information in uterine tissue lesion samples and uterine tissue control samples.
[0138] The cell subpopulation detection module 22 has a built-in Bayesian model, which is used to integrate and analyze gene expression data containing spatial location information in uterine tissue lesion samples and uterine tissue control samples with uterine single cell data in the same menstrual cycle using the Bayesian model to obtain multiple cell subpopulations containing spatial location information, and define each obtained cell subpopulation as the target subpopulation.
[0139] The adenomyosis microenvironment detection module 23 has a built-in zero distribution algorithm model, which is used to obtain the interaction probability distribution and the significance of cell interaction probability between various target subgroups using the zero distribution algorithm. The spatial location of the target subgroup with a cell interaction probability significance correction value of less than 0.05 and with typical adenomyosis morphological characteristics is defined as the adenomyosis microenvironment.
[0140] The adenomyosis space-specific regulatory molecule detection module 24 incorporates the NicheNet algorithm model. It is used to perform cell communication analysis on ligands, receptors, and target genes in the adenomyosis microenvironment using the NicheNet algorithm. It screens out cell communication signals with significantly differential expression of ligands, receptors, and target genes in the adenomyosis microenvironment, and identifies the ligands of the communication signals that promote the pathogenesis of adenomyosis in the screened cell communication signals as adenomyosis space-specific regulatory molecules.
[0141] The cell subpopulation detection module also has the following built-in:
[0142] The SCTransform algorithm model is used to standardize uterine single-cell data to obtain standardized uterine single-cell data.
[0143] The CCA analysis algorithm model is used to integrate standardized uterine single-cell data to obtain integrated uterine single-cell data.
[0144] The PCA algorithm model is used to perform linear dimensionality reduction on the integrated uterine single-cell data to obtain PCA dimensionality-reduced data of uterine single cells.
[0145] The UMAP nonlinear dimensionality reduction algorithm model is used to perform nonlinear dimensionality reduction on PCA dimensionality reduction data of uterine single cells to obtain UMAP dimensionality reduction data of uterine single cells.
[0146] The K-nearest neighbor algorithm model is used to perform cluster analysis on UMAP dimensionality reduction data of uterine single cells to obtain uterine single cell cluster data;
[0147] The Uterine Single-Cell Type Definition Submodule is used to define the uterine single-cell cluster data and obtain the type of each uterine single cell in the uterine single-cell data;
[0148] A negative binomial regression model was used to analyze uterine single-cell data to obtain the transcriptional characteristics of each cell type in the uterine single-cell data.
[0149] This invention also provides an electronic device, comprising: a processor and a memory; the memory for storing a computer program; and the processor for executing the computer program stored in the memory to enable the electronic device to perform the above-described method for determining spatially specific regulatory molecules of adenomyosis.
[0150] This invention also provides a computer-readable storage medium storing a computer program that, when executed by an electronic device, implements the above-described method for determining spatially specific regulatory molecules of adenomyosis.
[0151] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing a processor. The program can be stored in a computer-readable storage medium, which is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof. The storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. This available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state drive (SSD)).
[0152] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, or methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules / units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or units may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of apparatuses or modules or units may be electrical, mechanical, or other forms.
[0153] The modules / units described as separate components may or may not be physically separate. The components shown as modules / units may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules / units can be selected to achieve the objectives of the embodiments of the present invention, depending on actual needs. For example, the functional modules / units in the various embodiments of the present invention may be integrated into one processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into one module / unit.
[0154] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0155] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.
[0156] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A method of determining a spatially specific regulatory molecule of adenomyosis, characterized by, The method comprises the following steps: Step S1, obtaining a uterine tissue lesion sample and a uterine tissue control sample, and performing mRNA sequencing detection on the uterine tissue lesion sample and the uterine tissue control sample by using a spatial transcriptome sequencing technology to obtain gene expression data containing spatial position information in the uterine tissue lesion sample and the uterine tissue control sample; Step S2, integrating and analyzing the gene expression data containing spatial position information in the uterine tissue lesion sample and the uterine tissue control sample by using a Bayesian model and uterine single-cell data in the same menstrual cycle to obtain a plurality of cell subpopulations containing spatial position information, and defining each cell subpopulation obtained as a target subpopulation; Step S3, obtaining interaction probability distribution and cell interaction probability significance between each target subpopulation by using a zero distribution algorithm, and defining a spatial position of a target subpopulation having a corrected value of cell interaction probability significance less than 0.05 and a typical adenomyosis morphological feature as an adenomyosis microenvironment; Step S4, performing cell communication analysis on ligands, receptors and target genes in the adenomyosis microenvironment by using a NicheNet algorithm, screening cell communication signals in which ligands, receptors and target genes are significantly differentially expressed in the adenomyosis microenvironment, and determining a ligand of a communication signal promoting a pathogenic biological process of adenomyosis as an adenomyosis spatial specificity regulating molecule.
2. The method of determining a spatially specific regulatory molecule of adenomyosis according to claim 1, wherein, In the step S2, the gene expression data containing spatial position information in the uterine tissue lesion sample and the uterine tissue control sample are integrated and analyzed by using the Bayesian model and the uterine single-cell data in the same menstrual cycle to obtain a plurality of cell subpopulations containing spatial position information, which comprises: First, uterine single-cell data in the same menstrual cycle as the uterine tissue lesion sample is obtained, and the obtained uterine single-cell data is standardized by using an SCTransform algorithm to obtain standardized uterine single-cell data; Then, the standardized uterine single-cell data is integrated by using a CCA analysis algorithm to obtain integrated uterine single-cell data; Then, the integrated uterine single-cell data is subjected to linear dimension reduction processing by using a PCA algorithm to obtain PCA dimension reduction data of the uterine single cells; Then, the PCA dimension reduction data of the uterine single cells is subjected to non-linear dimension reduction by using a UMAP non-linear dimension reduction algorithm to obtain UMAP dimension reduction data of the uterine single cells; Then, the UMAP dimension reduction data of the uterine single cells is subjected to clustering analysis by using a K-nearest neighbor algorithm to obtain uterine single-cell clustering data; Then, the uterine single-cell clustering data is defined according to the expression of cell type classical genes to obtain the type of each uterine single cell in the uterine single-cell data; Then, the uterine single-cell data is analyzed by using a negative binomial regression to obtain transcriptional characteristics of each cell type in the uterine single-cell data; Then, the gene expression data containing spatial position information is deconvoluted by using the Bayesian model based on the transcriptional characteristics of each cell type in the uterine single-cell data to obtain a plurality of cell subpopulations containing spatial position information.
3. The method of determining a spatially specific regulatory molecule of adenomyosis according to claim 2, wherein, The cell types of the cell subpopulations comprising spatial location information include any one of epithelial cells, stromal fibroblasts, supporting cells, endothelial cells, or immune cells.
4. The method of determining a spatially specific regulatory molecule of adenomyosis according to claim 1, wherein, The cell types of the cell subpopulations comprising spatial location information include any one of epithelial cells, stromal fibroblasts, supporting cells, endothelial cells, or immune cells.
5. The method of determining a spatially specific regulatory molecule of adenomyosis according to claim 1, wherein, Before the step of obtaining multiple cell subpopulations comprising spatial location information in the step S2 of integrating and analyzing the gene expression data comprising spatial location information in the uterine tissue lesion sample and the uterine tissue control sample with the uterine single-cell data of the same menstrual cycle, and before the step of defining each of the obtained cell subpopulations as a target subpopulation, the method for determining the spatial-specific regulatory molecules of adenomyosis further comprises: obtaining the correlation of differential expression genes between the spatial transcriptome data and the single-cell data in different cell types by Pearson correlation analysis, to determine the cell subpopulations comprising spatial location information, and matching the cell subpopulations comprising spatial location information with preset cell gene markers, to define each of the obtained cell subpopulations comprising spatial location information as a target subpopulation.
6. The method of determining a spatially specific regulatory molecule of adenomyosis according to claim 1, wherein, Before the step of obtaining multiple cell subpopulations comprising spatial location information in the step S2 of integrating and analyzing the gene expression data comprising spatial location information in the uterine tissue lesion sample and the uterine tissue control sample with the uterine single-cell data of the same menstrual cycle, and before the step of defining each of the obtained cell subpopulations as a target subpopulation, the method for determining the spatial-specific regulatory molecules of adenomyosis further comprises: obtaining the correlation of differential expression genes between the spatial transcriptome data and the single-cell data in different cell types by Pearson correlation analysis, to determine the cell subpopulations comprising spatial location information, and matching the cell subpopulations comprising spatial location information with preset cell gene markers, to define each of the obtained cell subpopulations comprising spatial location information as a target subpopulation. Before the step of obtaining multiple cell subpopulations comprising spatial location information in the step S2 of integrating and analyzing the gene expression data comprising spatial location information in the uterine tissue lesion sample and the uterine tissue control sample with the uterine single-cell data of the same menstrual cycle, and before the step of defining each of the obtained cell subpopulations as a target subpopulation, the method for determining the spatial-specific regulatory molecules of adenomyosis further comprises: obtaining the correlation of differential expression genes between the spatial transcriptome data and the single-cell data in different cell types by Pearson correlation analysis, to determine the cell subpopulations comprising spatial location information, and matching the cell subpopulations comprising spatial location information with preset cell gene markers, to define each of the obtained cell subpopulations comprising spatial location information as a target subpopulation.
7. The method of determining spatially specific regulatory molecules of adenomyosis according to claim 1, wherein, The step S4 comprises: constructing a cell communication model comprising ligands, receptors, and target genes using ligand-receptor, ligand-target gene, and receptor-target gene information from public databases; using Wilcoxon test to calculate the gene differential expression of different cell types in the adenomyosis microenvironment, defining genes with a corrected P value less than 0.05 and a differential fold value greater than 1.2 as differential expression genes; and inputting the differential expression genes into the cell communication model to calculate the activity of ligand signals received by different cell types in the adenomyosis microenvironment using NicheNet algorithm; and defining the top 50 communication signals of the ligand signal activity received by different cell types in the adenomyosis microenvironment as significant communication signals; The significant communication signal is subjected to database and literature retrieval to obtain a functional characteristic of the significant communication signal; the functional characteristic of the significant communication signal is compared with the pathogenic biological process of the adenomyosis, and a ligand of a microenvironment-specific communication signal promoting the adenomyosis is determined as the spatial-specific regulatory molecule of the adenomyosis.
8. The method of determining a spatially specific regulatory molecule of adenomyosis according to claim 7, wherein, The significant communication signal at least includes any one of the following: an IHH signal pathway, a WNT signal pathway, a TGFβ signal pathway, a HIF1-α signal pathway, an estrogen receptor signal pathway, an oxytocin signal pathway, an ILK signal pathway, and an integrin signal pathway.
9. A system for determining spatially specific regulatory molecules of adenomyosis, characterized by, The method comprises the following steps: The transcriptome sequencing module is configured to obtain a uterine tissue lesion sample and a uterine tissue control sample, and perform mRNA sequencing detection on the uterine tissue lesion sample and the uterine tissue control sample by using a spatial transcriptome sequencing technology to obtain gene expression data containing spatial position information in the uterine tissue lesion sample and the uterine tissue control sample; The cell subpopulation detection module is internally provided with a Bayesian model, and is configured to integrate and analyze the gene expression data containing spatial position information in the uterine tissue lesion sample and the uterine tissue control sample with uterine single-cell data in the same menstrual cycle by using the Bayesian model to obtain a plurality of cell subpopulations containing spatial position information, and define each obtained cell subpopulation as a target subpopulation; The adenomyosis microenvironment detection module is internally provided with a zero distribution algorithm model, and is configured to obtain interaction probability distribution and significance of cell interaction probability between each target subpopulation by using the zero distribution algorithm, and define a spatial position of a target subpopulation having a corrected value of cell interaction probability significance less than 0.05 and a typical adenomyosis morphological feature as an adenomyosis microenvironment; The adenomyosis spatial-specific regulatory molecule detection module is internally provided with a NicheNet algorithm model, and is configured to perform cell communication analysis on ligands, receptors and target genes in the adenomyosis microenvironment by using the NicheNet algorithm, screen out cell communication signals in which the ligands, receptors and target genes are significantly differentially expressed in the adenomyosis microenvironment, and determine a ligand of a communication signal promoting the pathogenic biological process of the adenomyosis as an adenomyosis spatial-specific regulatory molecule.
10. The system for determining spatially specific regulatory molecules of adenomyosis according to claim 9, wherein, The cell subpopulation detection module is further internally provided with the following: The SCTransform algorithm model is configured to perform data standardization processing on the uterine single-cell data to obtain standardized uterine single-cell data; The CCA analysis algorithm model is configured to integrate the standardized uterine single-cell data to obtain integrated uterine single-cell data; The PCA algorithm model is configured to perform linear dimension reduction processing on the integrated uterine single-cell data to obtain PCA dimension reduction data of the uterine single-cell; The UMAP nonlinear dimension reduction algorithm model is configured to perform nonlinear dimension reduction on the PCA dimension reduction data of the uterine single-cell to obtain UMAP dimension reduction data of the uterine single-cell; The K-nearest neighbor algorithm model is configured to perform clustering analysis on the UMAP dimension reduction data of the uterine single-cell to obtain uterine single-cell clustering data; The uterus single cell type definition submodule is configured to define the uterus single cell clustering data to obtain the type of each uterus single cell in the uterus single cell data. The negative binomial regression model is configured to analyze the uterus single cell data to obtain the transcriptional features of each cell type in the uterus single cell data.
11. An electronic device, comprising: The electronic device comprises a processor and a memory; The memory is configured to store a computer program; The processor is configured to execute the computer program stored in the memory, so that the electronic device executes the method for determining the spatial-specific regulatory molecule of adenomyosis of any one of claims 1 to 8.
12. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the electronic device to implement the method for determining the spatial-specific regulatory molecule of adenomyosis of any one of claims 1 to 8.
Citation Information
Patent Citations
A Wave Information Prediction Method Based on Bayesian Theory and Hovercraft Attitude Information
CN106599427B
A method for analyzing two-cell structures from single-cell transcriptome data
CN111292807B
A cell communication analysis method and system
CN112466403B
A Deep Learning-Based Early Prediction Method and System for Diabetic Nephropathy
CN113687083B
A cell similarity measurement method that unifies cell type and state characteristics
CN114974435B