New method for screening myocardial therapeutic targets for ischemic heart failure by using single-cell sequencing

By using single-cell sequencing technology, Pdgfb and Tnfsf12 genes were screened as therapeutic targets for myocardial fibrosis in ischemic heart failure, which solves the problem that existing treatment strategies cannot improve cardiac function and provides a new method for controlling malignant fibrosis.

WO2026076708A1PCT designated stage Publication Date: 2026-04-16PKU HKUST SHENZHEN HONGKONG INSTITUTION
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-12
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

Existing treatment strategies for ischemic heart failure have failed to effectively improve cardiac function, leading to a vicious cycle of myocardial fibrosis, and lack precise treatment targets.

Method used

By screening the myocardium of healthy mice and IHF mice using single-cell sequencing technology, we revealed the regulatory modules and pathways related to malignant fibrosis and screened the Pdgfb and Tnfsf12 genes as targets for the treatment of myocardial fibrosis in ischemic heart failure.

Benefits of technology

Precise screening of Pdgfb and Tnfsf12 genes as targets for the treatment of myocardial fibrosis in ischemic heart failure provides a new treatment strategy to control and reverse malignant fibrosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024124440_16042026_PF_FP_ABST
    Figure CN2024124440_16042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a method for screening myocardial therapeutic targets for ischemic heart failure by using single-cell sequencing, which method comprises the following steps: S1, sample preparation; S2, construction of a single-cell expression matrix; S3, cell quality control; S4, cell type annotation; S5, cell communication analysis; and S6, co-expression network analysis. The provided method for screening myocardial therapeutic targets for ischemic heart failure by using single-cell sequencing comprises performing single-cell sequencing on hearts of healthy mice and IHF mice, screening for cell types with significant differences in cardiac transcriptional profiles of the healthy mice and IHF mice, then exploring interaction characteristics of various types of cells in malignant fibrotic IHF hearts, revealing potential regulatory modules and pathways related to malignant myocardial fibrosis in single-cell expression data of IHF hearts, and performing screening to obtain Pdgfb and Tnfsf12 genes which can be used as therapeutic targets for treating myocardial fibrosis in ischemic heart failure.
Need to check novelty before this filing date? Find Prior Art

Description

A novel method for screening myocardial therapeutic targets in ischemic heart failure using single-cell sequencing Technical Field

[0001] This invention belongs to the field of bioinformatics technology, specifically relating to a novel method for screening myocardial therapeutic targets in ischemic heart failure using single-cell sequencing. Background Technology

[0002] Heart failure, a disease that seriously affects human health, is a major concern in the global public health field. Epidemiological surveys show that the prevalence of heart failure in my country is 1.3%, meaning that there are currently more than 10 million people suffering from heart failure, with 3 million new cases added each year. According to the Heart Failure Medical Quality Report (2022) of the China Heart Failure Center Alliance, ischemic heart failure (IHF) is the most common type of heart failure, accounting for two-thirds of all heart failure cases, and its prevalence is showing an increasing trend year by year. Commonly used clinical treatments for ischemic heart failure (IHF) (nitrates, nicorandil, neprilysin inhibitor sacubitril + angiotensin II receptor blocker valsartan, Qili Qiangxin capsules, aspirin) are not very effective in treating IHF patients. Currently, the treatment strategy for ischemic heart failure mainly focuses on controlling symptoms and improving quality of life, such as drug therapy and revascularization, as well as monitoring natriuretic peptides to guide treatment. However, these treatments have failed to fundamentally improve cardiac function and reduce mortality in ischemic heart failure, demonstrating the limitations of current treatment strategies. Therefore, clarifying the pathogenesis of ischemic heart failure (IHF) and accurately screening key therapeutic targets for IHF are urgently needed.

[0003] Myocardial ischemia is a major cause of ischemic heart failure. Prolonged myocardial ischemia can lead to cardiomyocyte hypertrophy and fibroblast proliferation and differentiation, resulting in malignant fibrosis of the myocardial tissue. This gradually causes cardiac enlargement, decreased myocardial contractility, and ultimately heart failure. In fact, appropriate fibroblast proliferation and extracellular matrix deposition are necessary to maintain the structural integrity of the ventricles. However, excessive fibroblast activation (malignant fibrosis) leads to pathological remodeling of the heart, exacerbating myocardial ischemia and creating a vicious cycle of myocardial ischemia-heart failure-myocardial ischemia. Specifically, fibrosis leads to excessive extracellular matrix deposition, causing changes in the microenvironment of various cell types, limiting the heart's expansion capacity, aggravating cardiac damage, and forming malignant fibrosis. Therefore, controlling and reversing this malignant fibrosis is one of the key strategies for treating ischemic heart failure (IHF).

[0004] Single-cell RNA sequencing (scRNA-seq) technology can comprehensively reveal the transcriptional heterogeneity of various cells in the cardiac myocardium, identify potential therapeutic targets for malignant myocardial fibrosis, and is also an effective method for systematically analyzing the heterogeneity and molecular regulatory mechanisms of cardiac myocardial tissue. This invention stems from this.

[0005] Summary of the Invention

[0006] To address at least one of the aforementioned technical problems, this invention provides a novel method for screening therapeutic targets for myocardial fibrosis in ischemic heart failure using single-cell sequencing. The method involves performing single-cell sequencing on the myocardial cells of healthy mice and IHF mice, screening for cell types with significantly different transcriptional profiles between the hearts of healthy mice and IHF mice, exploring the cell interaction characteristics of different cell types in malignant fibrotic IHF hearts, and revealing potential regulatory modules and pathways related to malignant myocardial fibrosis in the single-cell expression data of IHF hearts. The Pdgfb and Tnfsf12 genes were identified as potential therapeutic targets for treating myocardial fibrosis in ischemic heart failure.

[0007] To solve the above technical problems, the technical solution of the present invention is as follows:

[0008] The purpose of this invention is to provide a novel method for screening myocardial therapeutic targets in ischemic heart failure using single-cell sequencing, comprising the following steps:

[0009] S1. Sample preparation; S2. Single-cell expression matrix construction; S3. Cell quality control; S4. Cell type annotation; S5. Cell communication analysis; S6. Co-expression network analysis.

[0010] Specifically, in S1, sample preparation, six 8-week-old SPF (Specific Pathogen Free) male C57BL / 6 mice were randomly divided into a sham-operated group (Sham group) and an IHF group. In the IHF group, the left anterior descending coronary artery (LAD) ligation was used to induce myocardial infarction and fibrosis in the heart of the C57BL / 6 male mice, simulating the pathological process of ischemic heart failure (IHF) in humans. Both groups of mice were fed standard feed and additive-free drinking water.

[0011] S2. Single-cell expression matrix construction: Three biological replicates were taken from each of the collected heart tissue groups (Sham and IHF) for single-cell sequencing (scRNA-seq) using an Illumina NovaSeq6000 platform. The sequencing depth was approximately 56×, resulting in a single sample data volume of approximately 150Gb. The data was then processed using Syngenomics software. A reference genome alignment was performed to obtain the original expression matrix.

[0012] S3. Cell quality control: Cells with a median UMI of less than 1000 or greater than 5000 and a median gene count of less than 500 are defined as low-quality cells. After filtering out these low-quality cells, high-quality data is obtained for downstream analysis.

[0013] S4. Cell type annotation: After using classic cell marker gene annotation, a total of 8 cell types were identified, including endothelial cells (Flt1, Cdh5, Egfl7, Pecam1), fibroblasts (Pdgfra, Gsn, Dcn, Timp1, Col1a2, Col3a1), B cells (Ighm, Iglc2, Cd79a, Cd79b), T cells (Trbc1, Cd3e, Cd3d), macrophages (Adgre1, C1qa, C1qb, Csf1r, Mrc1, C3ar1), smooth muscle cells (Acta2, Tagln, Myh11, Lmod1), perivascular cells (Kcnj8, Steap4, Abcc9, Vtn, Colec11), and neutrophils (Csf3r, Retnlg, S100a8, S100a9, Hdc, Slpi). The marker genes required for cell type identification exhibit good cell type specificity, and the cell annotation and clustering are accurate and reasonable, meeting the requirements of data analysis. Based on the above data, it was found that the number of fibroblasts in the myocardium of IHF mice was 1.18 times that of the Sham group, and the number of endothelial cells in the myocardium of IHF mice was 1.30 times that of the Sham group. The Piezo1 gene, a mechanosensitive cation channel, was specifically expressed in endothelial cells of fibrotic hearts. In conclusion, it is preliminarily determined that the expression of the Piezo1 gene in endothelial cells of mouse heart myocardium is closely related to the occurrence and development of myocardial fibrosis.

[0014] S5. Cell communication analysis: Using the Cellchat function, communication signals between different cell types were detected, and the differences in communication patterns between endothelial cells and fibroblasts in the myocardium of the Sham group and IHF heart were compared. Intercellular communication is crucial for the coordination of organismal functions and the development of disease. When cells cannot interact correctly or decode molecular information incorrectly, disease can result. To investigate the interaction characteristics of different cell types in the heart of malignant fibrotic IHF, intercellular communication analysis was performed on different cell types in the myocardium based on the scRNA-seq cell type annotation information obtained in section "S4" and combined with the CellChat ligand receptor database. The results showed that, by statistically analyzing the receptor-ligand logarithms between different cell groups, the interaction frequency between endothelial cells and fibroblasts in the myocardium of IHF heart was reduced. By comparing the communication patterns between endothelial cells and fibroblasts in the myocardium of the Sham group and IHF heart, it was found that the Pdgfb-Pdgfrb signaling pathway emitted by endothelial cells and the Tnfsf12-Tnfrsf12a signaling pathway emitted by fibroblasts in the myocardium of IHF heart were activated. It is speculated that the abnormal expression of Pdgfb and Tnfsf12 genes is the main cause of malignant fibrosis in IHF myocardium, and Pdgfb and Tnfsf12 genes can be used as therapeutic targets for the treatment of myocardial fibrosis in ischemic heart failure.

[0015] S6. Co-expression Network Analysis: To reveal potential regulatory modules and pathways associated with myocardial fibrosis in IHF cardiac single-cell expression data, we used hierarchical clustering analysis (HCA) and high-dimensional weighted gene co-expression network analysis (hdWGCNA) algorithms based on the scRNA-seq cell type annotation information obtained in step S4 to classify genes with similar expression patterns into different modules. We then delved into the functions of genes within each module to reveal the role of gene sets within those modules in IHF myocardial fibrosis. Results showed that with a soft threshold of 10 (selection criteria: scale-free topology model fit greater than or equal to 0.8), the gene co-expression network more closely resembled the characteristics of a scale-free network. Nine gene co-expression modules were identified using the ConstructNetwork function. By comparing endothelial cell-related co-expression modules in the myocardium of the Sham and IHF groups, it was found that the average expression level and percentage of genes in the co-expression module M2 gene set were significantly enriched in the myocardium of the Sham group, while the average expression level and percentage of genes in the co-expression module M8 gene set were significantly enriched in the myocardium of the IHF group. This suggests that genes in the co-expression modules M2 and M8 play important roles in the development and progression of myocardial fibrosis in IHF. Further analysis of the functions of genes in the co-expression modules M2 and M8 revealed that genes in the co-expression module M2 are mainly related to cell morphology regulation processes, while genes in the co-expression module M8 are mainly related to calcium...2+ It is related to the transport process of ions.

[0016] Compared with the prior art, the advantages of the present invention are:

[0017] This invention involves single-cell sequencing of the myocardium of healthy mice and IHF mice to screen for cell types with significant differences in the transcriptional profiles of the hearts of healthy mice and IHF mice. This allows for the exploration of the cell interaction characteristics of different cell types in the heart of malignant fibrotic IHF mice, and the revelation of potential regulatory modules and pathways related to malignant myocardial fibrosis in the single-cell expression data of IHF hearts. The Pdgfb and Tnfsf12 genes were screened and found to be therapeutic targets for the treatment of myocardial fibrosis in ischemic heart failure. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. The accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 shows the single-cell transcriptional patterns of mouse hearts in the Sham and IHF groups. In the figure, A is a bubble chart of mouse heart tissue cell annotations; B is a UMAP map of cell distribution; C is a heatmap of the top 20 highly variable genes (HVGs) in mouse heart tissue cells; D shows the proportions of different cell types in the Sham and IHF groups; and E is a violin plot of the Piezo1 gene in the Sham and IHF groups.

[0020] Figure 2 shows the cardiac cell communication analysis of mice in the Sham and IHF groups. In the figure, A and C represent the statistical frequency of cardiac cell interactions between mice in the Sham and IHF groups, respectively; B and D represent the intensity of cardiac signaling pathway interactions between mice in the Sham and IHF groups, respectively. The size of the circles in the bubble chart indicates the number of receptor-ligand pairs for that cell type; the larger the circle diameter, the more receptor-ligand pairs between cells, and the darker the circle color, the greater the probability of receptor-ligand communication between cells.

[0021] Figure 3 shows the single-cell co-expression network analysis (hdWGCNA) of the heart cells of mice in the Sham and IHF groups. In the figure, A represents soft threshold selection; B represents the construction of the co-expression network dendrogram (each leaf in the dendrogram represents a gene, and the color at the bottom indicates the allocation of co-expression modules; the "gray" modules consist of genes not grouped into any co-expression module); C represents the expression level of co-expression module genes in different cell types; D represents the functional enrichment analysis of co-expression module genes (Top 5); E represents the hub genes and protein-protein interaction network (PPI) of the co-expression module M2; and F represents the hub genes and protein-protein interaction network (PPI) of the co-expression module M8.

[0022] Figure 4 shows the hub genes and protein-protein interaction network (PPI) of the nine co-expressed gene modules (represented by M1 to M9). Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0024] The present invention provides a novel method for screening myocardial therapeutic targets for ischemic heart failure using single-cell sequencing, comprising the following steps:

[0025] I. Sample Preparation

[0026] To clarify the pathogenesis of ischemic heart failure (IHF) and accurately screen key therapeutic targets for IHF, C57BL / 6 mice were used as experimental animals. The IHF mouse model reported in the literature (Golforoush P, 2020 Basic Res Cardiol, 30; 115(6):73) was used for the experiment. The specific details are as follows: The experimental mice were fasted for 12 hours before the operation. The mice were anesthetized, fixed, shaved, and disinfected. A longitudinal incision of about 1.5 cm was made on the skin about 1-2 mm from the left sternal border. The chest wall muscles were bluntly dissected layer by layer. The thoracic cavity was quickly entered from the 4th intercostal space. The intercostal space was opened with hemostatic forceps. The heart was gently squeezed with the left hand in coordination with the heartbeat to make the heart pop out of the opening. The left atrial appendage was ligated with an 8-0 suture needle through the left anterior descending coronary artery 1-2 mm below the left atrial appendage and 0.5 mm next to the pulmonary artery conus. When the outer surface of the anterior wall of the left ventricle turned pale, it proved that the ischemia was successful. Heart tissue is a typical heterogeneous tissue, and the results of single-cell dissociation directly affect the quality of single-cell transcriptome sequencing. We optimized the single-cell dissociation technique for heart tissue, achieving a cell concentration of 8.0 × 10⁵ / mL, a viability of 90.5%, and no obvious impurities in the microscopic field of view (see Figure 1A). In summary, the number and viability of captured cells meet the requirements for single-cell transcriptome sequencing.

[0027] II. Construction of Single-Cell Expression Matrix

[0028] Fastp software was used to perform basic quality checks and filtering on the raw reads after the data was processed, resulting in high-quality, clean reads suitable for analysis. The filtering criteria were: (1) removing reads containing polyA; (2) removing reads with more than 3 N bases; and (3) removing low-quality reads (reads with a quality value Q less than or equal to 5 accounting for more than 20% of the total reads). Subsequently, Syngenomics software was used... (v1.17.0) Performs cell quality control, detection, and alignment with the reference genome (mm9) to obtain the raw expression matrix. The specific steps are as follows: First, construct the reference genome directory using the `mkref` command, requiring the genome sequence (FASTA file) and gene annotation information (GTF file). Once the reference genome is constructed, use the Celescope RNA command to analyze the samples. Specific steps include: providing sample information (sample), such as sample ID, analysis type, and reagent version; filtering and correcting the barcode based on the read1 sequence information, adding the corrected barcode and the original UMI sequence to the read2 ID (barcode); performing quality control on the data to ensure sequence quality (cutadapt); calling STAR to align the read2 sequences to the reference genome, generating an alignment file (BAM file) (star); using the FeatureCounts tool to count the alignment results in each cell according to the BAM file (featureCounts); performing UMI counting and cell number assessment, finally generating the expression matrix (count).

[0029] III. Cell quality control

[0030] Cells with a median UMI of less than 1000 or greater than 5000, and a median gene count of less than 500, were defined as low-quality cells. After filtering out these low-quality cells, high-quality data were obtained for downstream analysis.

[0031] IV. Cell Type Annotation

[0032] After quality control, high-quality cells were retained for downstream analysis. The Sctransform function was then used to standardize the data and screen for hypervariable genes. The specific steps are as follows: First, a generalized linear model (GLM) was used to fit the relationship between the expression level of each gene and sequencing depth, calculating the regression coefficients (β0 and β1) and the negative binomial distribution parameter θ for each gene using a quasi-Poisson distribution. This step used the R function glm for fitting and estimated the θ value using the theta.ml function in the MASS package. Second, a smoothing function (ksmooth) was used to correct the fitting results to prevent overfitting of low-expression genes. This process included calculating the logarithmic mean and variance of each gene and filtering outliers, optimizing the smoothing effect using the bandwidth adjustment parameter (bw_adjust). Finally, the Pearson residuals for each gene were calculated based on the fitted mean μ and standard deviation σ, ensuring that the residual value did not exceed the square root of the total cell count. The corrected read count matrix was generated using the Pearson residuals and the median sequencing depth. The 2000 genes with the largest expression fluctuations among cells were selected as highly variable genes (HVGs). Further, we used the Harmony function (https: / / github.com / immunogenomics / harmony) to remove potential batch effects between samples. The specific steps are as follows: First, a Seurat object is created and standard PCA analysis is performed, while setting condition variables. Then, the RunHarmony function is called, passing the Seurat object and specifying the variables to be integrated. Key parameters include group.by.vars (specifying which variable to group by; this application uses group to distinguish between the Sham and IHF groups), max.iter.harmony (setting the maximum number of iterations, default is 10), lambda (integration strength parameter, default value is 1, generally between 0.5 and 2), theta (cluster diversity penalty parameter, default value is 2, larger values ​​lead to higher diversity), dims.use (PCA dimensions used, default is all dimensions), and sigma (width of soft k-means clustering, default value is 0.1). Finally, the Harmony-corrected embedding results are obtained and viewed. Subsequently, we performed dimensionality reduction on the integrated data using Principal Component Analysis (PCA). The steps of PCA dimensionality reduction are as follows: First, the VariableFeatures function is used to extract highly variable genes identified after standardization from the dataset. Then, when calling RunPCA, these highly variable genes are used as input features for PCA calculation.The RunPCA function, by default, calculates and stores the first 50 principal components. The ElbowPlot function from the Seurat package is used to visualize the decrease in PC variance, and the first 15 principal components are selected for downstream analysis. The FindNeighbors and FindClusters functions are then used to cluster all cells, with the latter set to a resolution of 0.2. The specific steps are as follows: First, the default Annoy algorithm is called using the NNHelper function to calculate the 20 nearest neighbors for each cell. The Annoy algorithm accelerates the search by building a tree-structured index and uses Euclidean distance as the metric. Next, a Shared Nearest Neighbor (SNN) is calculated based on Jaccard similarity, recording the nearest neighbor overlap for each cell and pruning according to a default threshold of 1 / 15. The calculated nearest neighbor matrix (NN) and shared nearest neighbor matrix (SNN) are converted into Graph objects and stored in the graphs slot of the Seurat object, named RNA_nn and RNA_snn. The FindClusters function optimizes cell group identification through the shared nearest neighbor (SNN) module. Then, the function initialization parameters are set, such as `graph.name` (defaults to `RNA_snn`), `resolution` (0.5 in this application), and `algorithm` (defaults to the original Louvain algorithm). Next, the function checks the validity of the graph object and calls the corresponding clustering function according to the specified algorithm: the Louvain algorithm uses `RunModularityClustering` implemented in C++, and the Leiden algorithm uses `RunLeiden` implemented in Python. During clustering, the function optimizes the modularization function to determine cell clusters and handles outliers, reassigning individual cells to the densest populations. Finally, the function adds the clustering results to the metadata of the `Seurat` object, updates the `Idents` and `seurat_clusters` columns, and logs the results. Two-dimensional data visualization of the principal component analysis data is performed using the dimensionality reduction method of Uniform Manifold Approximation and Projection (UMAP). The specific computation is as follows: the R package `uwot` is called, and UMAP dimensionality reduction is implemented using the specified dimensionality reduction result (the first 10 dimensions of PCA in this application). When executing UMAP, the `uwot` method calls the `uwot::umap` function, passing in the data matrix and parameters. The function converts the UMAP output to a specified format, sets column names such as `UMAP_1` and `UMAP_2`, creates a `DimReduc` object to store the results, and saves the UMAP model as needed. Finally, the function updates the dimensionality reduction result of the `Seurat` object, adds a new `DimReduc` object, and logs the changes.Using classic cell marker gene annotation, a total of endothelial cells (Flt1, Cdh5, Egfl7, Pecam1), fibroblasts (Pdgfra, Gsn, Dcn, Timp1, Col1a2, Col3a1), B cells (Ighm, Iglc2, Cd79a, Cd79b), T cells (Trbc1, Cd3e, Cd3d), and macrophages (Adgre1, C1) were identified. Eight cell types were identified, including qa, C1qb, Csf1r, Mrc1, C3ar1, smooth muscle cells (Acta2, Tagln, Myh11, Lmod1), perivascular cells (Kcnj8, Steap4, Abcc9, Vtn, Colec11), and neutrophils (Csf3r, Retnlg, S100a8, S100a9, Hdc, Slpi) (see Figure 1, A to C). The marker genes required for cell type identification showed good cell type specificity, and the cell annotation clusters were accurate and reasonable, meeting the requirements of data analysis. Based on the above data, it was found that the number of fibroblasts in the myocardium of IHF mice was 1.18 times that of the Sham group, and the number of endothelial cells in the myocardium of IHF mice was 1.30 times that of the Sham group (see Figure 1, D). The Piezo1 gene, a mechanosensitive cation channel, was specifically expressed in endothelial cells of fibrotic hearts (see Figure 1, E).

[0033] V. Cell Communication Analysis

[0034] Intercellular communication is crucial for the coordination of biological functions and the development of disease. Disease arises when cells fail to interact correctly or decode molecular information inappropriately. To investigate the interaction characteristics of different cell types in malignant fibrotic heart disease (IHF), this application used CellChat software to infer, analyze, and visualize intercellular communication networks from single-cell RNA sequencing (scRNA-seq) data. First, the necessary R libraries, including CellChat and patchwork, were loaded. A CellChat object was created from the digital gene expression matrix and cell tag information. The CellChat object was initialized by calling the createCellChat function and passing in gene expression data and cell metadata. Specifically, the createCellChat function was used with cell-annotated Seurat data as input. Next, group.by was specified to define the cell group, with RNA as the default data type. Then, using these parameters, the createCellChat function created a CellChat object for subsequent intercellular communication analysis. Next, a database for ligand-receptor interactions was set up, selecting the human database CellChatDB.mouse, and intercellular communication analysis was performed using only a subset of secreted signals. The subset of the CellChat database was set as the default database for the object by calling the subsetDB function. To infer intercellular communication networks, we used a "trimean" statistical method to calculate the average gene expression level for each cell type, yielding fewer but stronger interactions. Specifically, for each cell type, the first quartile (Q1), median, and third quartile (Q3) were calculated, and Tukey's trimean formula (Q1 + 2 * Median + Q3) / 4 was applied. This process was performed on all cell types in the list, resulting in a trimean value for each cell type, reflecting the central trend in gene expression levels across cell types. The computeCommunProb function was called to calculate the communication probability, with the default parameter type set to "triMean," to calculate the average gene expression level for each cell group. Furthermore, cells in each cell group with a minimum number of cells less than 10 for intercellular communication were filtered out. Next, we use the computeCommunProbPathway function to calculate the communication probability at the signaling pathway level and aggregate the intercellular communication network. The specific process is as follows: set "triMean" as the default parameter and adjust several parameters (including whether to use the original data, whether to consider the size of the cell population, and whether to introduce spatial distance constraints, etc.) to adapt to the data characteristics.Based on the set parameters, `computeCommunProbPathway` updates the `CellChat` object after execution, calculating the communication probability and statistical significance (p-value) of each signaling pathway. By calling the `netVisual_aggregate` function, we generated communication network diagrams for different signaling pathways and displayed the interactions between different cell populations using a bubble chart. Specifically, we used `thresh` (p-value) to filter interactions with lower significance. We adjusted the bubble size in the diagram using `vertex.weight` and `vertex.size.max` to reflect the importance and number of cell populations, while the edge width was set using `edge.weight.max` and `edge.width.max` to represent the communication intensity. The results showed that, by counting the number of receptor-ligand pairs between different cell groups, the interaction frequency between endothelial cells and fibroblasts in the myocardium of IHF hearts was reduced (see Figure 2A and Figure 2C). By comparing the communication patterns between endothelial cells and fibroblasts in the myocardium of Sham group and IHF heart disease, it was found that the Pdgfb-Pdgfrb signaling pathway emitted by endothelial cells and the Tnfsf12-Tnfrsf12a signaling pathway emitted by fibroblasts in IHF heart disease were activated. It is speculated that abnormal expression of Pdgfb and Tnfsf12 genes is the main cause of malignant fibrosis in IHF heart disease (see Figure 2B and Figure 2D), and that Pdgfb and Tnfsf12 genes could serve as therapeutic targets for treating myocardial fibrosis in ischemic heart failure.

[0035] VI. Co-expression Network Analysis

[0036] To reveal potential regulatory modules and pathways associated with myocardial fibrosis in IHF cardiac single-cell expression data, we first imported a preprocessed Seurat object containing Harmony-integrated sample data using the `readRDS` function. Then, we used the `DimPlot` function to visualize the distribution of cell types and confirm the basic characteristics of the data. Specifically, we loaded the Seurat object, specifying cell type as the plotting dimension (parameter set to `celltype`). We used the `cols` parameter to set the cell color and the `group.by` parameter to group and color cells according to their type. During visualization, we could also adjust the point size (`pt.size` parameter) and transparency (`alpha` parameter). We used `cols.highlight` and `sizes.highlight` to adjust the color and size of highlighted cells. Finally, the `DimPlot` function returned a `ggplot` object. By adjusting these parameters, we could visually represent the distribution of cell types in the reduced-dimensional space. Before performing hdWGCNA analysis, we needed to configure the Seurat object using the `SetupForWGCNA` function. Specifically, when selecting genes, we used the "fraction" method, specifying genes expressed in at least 5% of the cells. The process of constructing Metacells used the MetacellsByGroups function to group cells by "celltype" and "Sample" to identify similar cells from the Sham and IHF groups. We chose k=25, indicating that each Metacell aggregates 25 similar cells, and limited the overlap between Metacells (max_shared=10). Next, the Metacell expression matrix was normalized using the NormalizeMetacells function and then subjected to subsequent Seurat processing, including ScaleMetacells, RunPCAMetacells, RunHarmonyMetacells, and RunUMAPMetacells, to visualize the aggregated expression characteristics. Specifically, using a Seurat object as the loading object, the NormalizeData function of Seurat was called. When normalizing the Metacell expression matrix using the LogNormalize method, the total gene expression level of each cell was first calculated, and then the expression count of each gene was divided by the total count of the corresponding cells to obtain the relative expression level of that gene in the cell. Next, this relative expression level is multiplied by a predetermined scaling factor (10000 in this application), and the result is transformed by the natural logarithm to balance the differences in sequencing depth between different cells and reduce the impact of extreme values.Then, the ScaleMetacells function was used to standardize the expression data for each gene, eliminating absolute differences in expression levels by calculating the Z-score (gene expression value in each cell minus the mean and divided by the standard deviation). Next, the RunPCAMetacells function reduced the dimensionality of the cellular gene expression data using principal component analysis (PCA), preserving the principal components that captured major variations and thus simplifying the data structure. To correct for batch effects, the RunHarmonyMetacells function applied the Harmony algorithm to optimize the data, aligning similar cells under different conditions in a multidimensional space for better integration. Finally, the RunUMAPMetacells function further reduced the dimensionality using the UMAP algorithm, creating a low-dimensional space where local and global data structures were preserved, allowing similar cells to be close to each other in the graph, thus forming a clear visualization of cell subpopulations. Before performing co-expression network analysis, we used the SetDatExpr function to set the expression matrix for network analysis. We selected endothelial cells as the analysis object. Subsequently, the soft threshold was selected using the TestSoftPowers function. With the signed network type, 10 was chosen as the soft threshold, which showed good fit (R) in the scale-free topology model. 2(≥0.8). When constructing the co-expression network, the `ConstructNetwork` function was used, specifying `soft_power = 10`. The network results can be visualized using the `PlotDendrogram` function, obtaining the gene composition information of different modules (see Figure 3A). To calculate module feature genes (MEs), we first standardized the data (ScaleData), and then used the `ModuleEigengenes` function to calculate the corresponding module feature genes. The specific steps are as follows: First, specify the `VariableFeatures` in the data as the objects to be standardized, and select to use a linear model for regression (`model.use="linear"`). Then, centering and scaling operations are performed. The maximum value of scaling the data is limited to 10 by the parameter `scale.max` to reduce the influence of features expressed only in a very small number of cells. After the calculation is completed, we extracted the harmonized module feature genes (hMEs) and the uncorrected module feature genes (MEs). Next, we calculated the kME based on the feature genes. The `ModuleConnectivity` function was used to obtain the kME value for each gene. The specific procedure is as follows: First, the `group.by` parameter was used to select the column containing grouping information in `seurat_obj@meta.data` (here, the cell type column "cell_type" was selected), and the `group_name` parameter was used to specify the specific group to be used for kME calculation (in this application, "endothelial cells" was selected). The `ModuleConnectivity` function calculates the correlation between each gene in the specified endothelial cell and the feature genes (MEs) of each module, thus obtaining the kME value for each gene. A high kME value indicates that the gene has high network connectivity within its module, which helps to identify the central genes within the module. To identify hub genes within the module, we used the `GetHubGenes` function to extract the top 10 hub genes. The results show the name of each hub gene, its module, and its kME value. To better illustrate the results of the module feature genes, we used the `ModuleFeaturePlot` function to construct a feature map for each module, displaying hMEs and hub gene scores. Using Seurat's DotPlot function, we can intuitively analyze the expression of each module in different cell types. The specific operation is as follows: By calling the DotPlot function and passing in the Seurat object seurat_obj, the module list mods, and the grouping criterion group.by parameter (in this application, cell_type is used as the setting object), a DotPlot plot is generated. The plot shows the expression of each module in different cell types, where the size of the dot represents the proportion of cells expressing that module in that cell type, and the color of the dot represents the average expression level.The results showed that nine gene co-expression modules were identified using the ConstructNetwork function (see Figure 4, where M1, M2, M3, M4, M5, M6, M7, M8, and M9 are used to represent these nine gene co-expression modules). Comparison of endothelial cell-related co-expression modules in the myocardium of the Sham and IHF groups revealed that the average expression level and percentage of genes in the M2 gene set were significantly enriched in the myocardium of the Sham group, while the average expression level and percentage of genes in the M8 gene set were significantly enriched in the myocardium of the IHF group. This suggests that genes in the M2 and M8 co-expression modules play an important role in the development and progression of myocardial fibrosis in IHF (see Figure 3C). Further analysis of the functions of genes in co-expression modules M2 and M8 revealed that genes in co-expression module M2 (see Figure 3E and Figure 3G) are mainly related to cellular morphology regulation processes (see Figure 3D), while genes in co-expression module M8 (see Figure 3F and Figure 3H) are mainly related to Ca. 2+ The transport process of ions is related (see D in Figure 3).

[0037] It should be understood that the specific embodiments described above are merely illustrative or explanatory of the principles of the invention and do not constitute a limitation thereof. Therefore, any modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and scope of the invention should be included within the protection scope of the invention. Furthermore, the appended claims are intended to cover all variations and modifications falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.

Claims

1. A novel method for screening myocardial therapeutic targets in ischemic heart failure using single-cell sequencing, characterized in that, Includes the following steps: S1. Sample preparation: SPF-grade mice were used as experimental animals. The sham surgery group (Sham group) and the ischemic heart failure (IHF group) group were designed. The IHF group was induced by ligation of the left anterior descending coronary artery to induce myocardial infarction and fibrosis pathological changes to simulate the pathological process of human ischemic heart failure. Mice in both the Sham and IHF groups were given normal feed and additive-free drinking water. Heart tissues were collected from mice in the Sham and IHF groups respectively. S2. Single-cell expression matrix construction: Single-cell sequencing was performed on collected heart tissue using an Illumina NovaSeq 6000 platform with a sequencing depth of approximately 56×. The sequencing was then performed using Syngenomics software. Perform a reference genome alignment to obtain the original expression matrix; S3. Cell quality control: Cells with a median UMI of less than 1000 or greater than 5000 and a median gene count of less than 500 are defined as low-quality cells. After filtering out these low-quality cells, high-quality cell data is obtained for downstream analysis. S4. Cell type annotation: After using the classic cell marker gene annotation, a total of 8 cell types were identified, including endothelial cells, fibroblasts, B cells, T cells, macrophages, smooth muscle cells, perivascular cells and neutrophils, and scRNA-seq cell type annotation information was obtained. S5. Cell communication analysis: Using the Cellchat function, we detected the communication signals between different types of cells and compared the differences in the communication patterns between endothelial cells and fibroblasts in the cardiac myocardium of the Sham group and the IHF group. S6. Co-expression network analysis: Based on the scRNA-seq cell type annotation information obtained in step S4, hierarchical clustering analysis and high-dimensional weighted gene co-expression network analysis algorithms are used to classify genes with similar expression patterns into different modules, and the functions of genes in each module are explored in depth to reveal the role of gene sets in IHF myocardial malignant fibrosis.

2. The novel method according to claim 1, characterized in that, In step S1, the collected cardiac tissue underwent single-cell dissociation, and the cell concentration after dissociation was 8.0 × 10⁻⁶. 5 / mL, activity 90.5%.

3. The new method according to claim 1, characterized in that, In step S2, Fastp software is used to perform basic quality checks and filtering on the raw data after it is taken off the machine, so as to obtain high-quality data that can be used for analysis. The filtering criteria are as follows: (1) Remove Reads containing polyA; (2) Remove Reads with more than 3 N; (3) Remove low-quality Reads, wherein the low-quality Reads are Reads in which the number of bases with a quality value Q less than or equal to 5 accounts for more than 20% of the total Reads.

4. The new method according to claim 1, characterized in that, In step S2, the steps for operating the original expression matrix include: constructing a reference genome catalog using the mkref command, which requires genome sequence and gene annotation information; once the reference genome catalog is constructed, using the celescope rna command to analyze the samples; Provide sample information, which includes sample ID, analysis type, and reagent version; Based on the information filtering and barcode correction of read1 sequence, the corrected barcode and the original UMI sequence are added to the ID of read2; Perform quality control on the data to ensure the quality of the sequences; Use STAR to align the reads2 sequence to the reference genome and generate an alignment file; The FeatureCounts tool is used to count the alignment results in each cell based on the alignment file. UMI counting and cell number assessment were performed, and finally, an expression matrix was generated.

5. The novel method according to claim 1, characterized in that, Step S4 includes the following specific operations: The Sctransform function was used to standardize and screen for hypervariable genes in quality-controlled cell data. Among them, the 2000 genes with the largest expression fluctuations among cells were selected as hypervariable genes. Use the Harmony function to remove potential batch effects between samples; Principal component analysis was used to reduce the dimensionality of the integrated data. The ElbowPlot function in the Seurat package was used to visualize the decrease in PC variance. The first 15 principal components were selected for downstream analysis. The FindNeighbors and FindClusters functions were used to cluster all cells in sequence, with the latter having a resolution of 0.

2. Two-dimensional data visualization is performed on the principal component analysis data by using uniform manifold approximation and projection dimensionality reduction methods.

6. The novel method according to claim 1, characterized in that, In step S5, CellChat software is used to infer, analyze, and visualize intercellular communication networks from single-cell RNA sequencing data; A database for ligand-receptor interactions was set up, using the human database CellChatDB.mouse, and intercellular communication analysis was performed using only a subset of secreted signals. The subset of the CellChat database was set as the default database for the object by calling the subsetDB function. The trimean statistical method was used to calculate the average gene expression level for each cell type to generate fewer but stronger interactions. The computeCommunProb function was called when calculating communication probabilities, with the default type parameter set to triMean, to calculate the average gene expression level for each cell group. Cells in each cell group with a minimum number of cells less than 10 for intercellular communication were filtered out. The computeCommunProbPathway function was used to calculate the communication probability at the signaling pathway level and aggregate the intercellular communication network. The netVisual_aggregate function was used to generate communication network diagrams for different signaling pathways, and bubble charts were used to illustrate the interactions between different cell populations.

7. The novel method according to claim 1, characterized in that, In step S6, the readRDS function was used to import the preprocessed Seurat object, which contained the sample data integrated by Harmony; the DimPlot function was used to visualize the distribution of cell types and confirm the basic characteristics of the data.

8. The novel method according to claim 7, characterized in that, In step S6, prior to the high-dimensional weighted gene co-expression network analysis... The Seurat object needs to be set up using the SetupForWGCNA function. Specific operations include: when selecting genes, the fraction method is used to specify genes expressed in at least 5% of cells; the MetacellsByGroups function is used to group cells by cell type and sample to identify similar cells from the Sham and IHF groups; the NormalizeMetacells function is used to normalize the Metacell expression matrix, followed by subsequent Seurat processing, including ScaleMetacells, RunPCAMetacells, RunHarmonyMetacells, and RunUMAPMetacells, to visualize the aggregated expression features. Specifically, using the Seurat object as the loading object, the NormalizeData function of Seurat is called. When normalizing the Metacell expression matrix using the LogNormalize method, the total gene expression level of each cell is first calculated, and then the expression level of each gene is... The relative expression level of a gene within a cell is obtained by dividing the expression count by the total number of corresponding cells. This relative expression level is then multiplied by a predetermined scaling factor, and the result is transformed using the natural logarithm to balance the differences in sequencing depth between different cells and reduce the impact of extreme values. The expression data of each gene is standardized using the ScaleMetacells function, and the absolute magnitude difference of expression levels is eliminated by calculating the Z-score, which is the gene expression value in each cell minus the mean and divided by the standard deviation. The RunPCAMetacells function reduces the dimensionality of the cell's gene expression data through principal component analysis, retaining the principal components that capture major variations, thereby simplifying the data structure. To correct for batch effects, the RunHarmonyMetacells function applies the Harmony algorithm to optimize the data, aligning similar cells under different conditions in a multidimensional space for better integration. The RunUMAPMetacells function further reduces the dimensionality using the UMAP algorithm, creating a low-dimensional space in which local and global data structures are preserved, allowing similar cells to be close to each other in the graph, thus forming a clear visualization of cell subpopulations. Use the SetDatExpr function to set the expression matrix for network analysis, select the analysis object, and use the TestSoftPowers function to select the soft threshold. Use the signed network type and select 10 as the soft threshold.

9. The novel method according to claim 8, characterized in that, In step S6, when constructing a high-dimensional weighted gene co-expression network, the ConstructNetwork function is used with soft_power=10 specified. The network results can be visualized using the PlotDendrogram function, which provides information on the gene composition of different modules.

10. The novel method according to claim 1, characterized in that, The therapeutic targets for ischemic heart failure and myocardial fibrosis screened using the novel method are the Pdgfb and Tnfsf12 genes.

Citation Information

Cited By

  • Disease target network construction method fusing pathological image and space transcriptome

    CN122117003A