A method for screening transcription regulatory elements and use thereof

By integrating single-cell transcriptome and chromatin accessibility data, we screened and validated transcriptional regulatory elements, solved the systemic problem of constructing transcriptional regulatory networks, identified key regulatory nodes, and applied them to the identification of therapeutic targets for diseases such as COVID-19.

CN120636556BActive Publication Date: 2026-06-16WUHAN UNIV OF SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN UNIV OF SCI & TECH
Filing Date
2025-05-07
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing technologies lack systematic methods to predict the synergistic effects of transcriptional regulatory elements and cross-level regulatory events, which affects the identification of key regulatory nodes in diseases and the in-depth analysis of pathological mechanisms. In particular, in the immune dysregulation caused by SARS-CoV-2 infection, the key regulatory nodes in the process of excessive activation of immune cells such as monocytes have not been fully elucidated.

Method used

By integrating single-cell transcriptome (scRNA-seq) and chromatin accessibility (scATAC-seq) data, noise reduction, dimensionality reduction clustering, cell annotation, and differential analysis were performed to screen differentially expressed genes and transcriptional elements involved in chromatin accessibility regulation. The regulatory specificity of transcriptional regulatory elements was verified by combining immunoprecipitation and chromatin immunoprecipitation qPCR, and a monocyte-specific epigenetic-transcriptional regulatory network was constructed.

Benefits of technology

This study reveals synergistic interactions among transcriptional regulatory elements and chromatin remodeling complex interactions, identifies key transcription factors and epigenetic regulatory elements, and provides new approaches to disease therapeutic targets, particularly offering novel targets for the prevention and treatment of inflammatory responses in COVID-19.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636556B_ABST
    Figure CN120636556B_ABST
Patent Text Reader

Abstract

The application discloses a method for screening transcriptional regulatory elements and application thereof, relates to the technical field of genes, and provides a method for predicting synergistic transcriptional regulatory elements based on single-cell multiomics and experimental exploration, which can reveal potential synergistic effects between transcriptional regulatory elements, chromatin remodeling complex interaction and other cross-level regulation events, and provides a method for identifying key regulation nodes of diseases. On the other hand, the application discloses key transcriptional regulatory elements involved in disease response based on single-cell multiomics sequencing technology. By integrating single-cell chromatin accessibility analysis and transcriptome sequencing data, the interaction between core transcription factors, epigenetic regulatory elements and signal transduction pathways driving disease generation is clarified, and a monocyte-specific epigenetic-transcriptional regulatory network is successfully constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical technology, and in particular to a method for screening transcriptional regulatory elements and its application. Background Technology

[0002] Gene expression regulation is a multi-level, precise regulatory process encompassing multiple key stages, from transcription initiation and post-transcriptional processing to translation and post-translational modifications. This multi-level regulatory network influences cellular function and state by precisely controlling gene expression. As core executors of this regulatory network, trans-acting factors, represented by transcription factors, chromatin remodeling complexes, and RNA polymerases, form dynamic regulatory modules through complex synergistic interactions. Although existing research has revealed the important functions of these regulatory elements, the construction of their regulatory networks and the patterns of their interactions remain unclear. Currently, there is a lack of systematic methods for predicting their synergistic effects. Analyzing the hierarchical structure and operational patterns of gene regulatory networks can not only elucidate the potential mechanisms of transcriptional regulation in eukaryotes but also provide methods for identifying key regulatory nodes in diseases and for in-depth analysis of pathological mechanisms, which is of great significance for the screening of therapeutic targets. Summary of the Invention

[0003] This application provides a method for screening transcriptional regulatory elements and its application, thereby providing a method for screening transcriptional regulatory elements.

[0004] In a first aspect, this application provides a method for screening transcriptional regulatory elements, comprising the following steps:

[0005] Denoising is performed on the scRNA-seq and scATAC-seq data of the samples to be analyzed to obtain preprocessed scRNA-seq and scATAC-seq data;

[0006] Dimensionality reduction clustering and cell annotation were performed on the preprocessed scRNA-seq and scATAC-seq data to obtain at least one cell type of the sample to be analyzed.

[0007] Perform dimensionality reduction clustering and cell annotation on at least one type of cell in the sample to be analyzed to obtain a cell subpopulation of that type of cell;

[0008] Differentially expressed genes and differentially expressed chromatin peaks were analyzed in the cell subpopulations of the sample to be analyzed, and differentially expressed gene data and differentially expressed chromatin peak data of the cell subpopulations of the sample to be analyzed were obtained.

[0009] scRNA-seq and scATAC-seq were used to analyze the differentially expressed gene data and differential chromatin accessibility peak data of the cell subpopulations of the sample to be analyzed, in order to identify the differentially expressed genes regulated by chromatin accessibility in the cell subpopulations of the sample to be analyzed.

[0010] Enrichment analysis was performed on differential chromatin accessibility peak data of cell subpopulations of the sample to be analyzed to identify transcriptional elements involved in the regulation of gene expression related to chromatin accessibility peaks.

[0011] By analyzing the potential regulatory regions and target genes of transcriptional regulatory elements, transcriptional elements with highly overlapping regulatory regions and target genes are screened, which are the target regulatory elements.

[0012] This application provides a method for predicting synergistic transcriptional regulatory elements based on a combination of single-cell multi-omics and experimental investigation. This method can reveal potential synergistic effects among transcriptional regulatory elements, interactions of chromatin remodeling complexes, and other cross-level regulatory events, providing a method for identifying key regulatory nodes in diseases. On the other hand, based on single-cell multi-omics sequencing technology, the application systematically reveals key transcriptional regulatory elements involved in disease responses. By integrating single-cell chromatin accessibility analysis and transcriptome sequencing data, the core transcription factors driving disease, epigenetic regulatory elements, and interactions between signal transduction pathways are clarified, successfully constructing a monocyte-specific epigenetic-transcriptional regulatory network.

[0013] In some embodiments, the samples to be analyzed include public sample data. Typically, the samples to be analyzed consist of a large amount of scRNA-seq and scATAC-seq sequencing data. Since scRNA-seq and scATAC-seq sequencing work requires collaboration among multiple teams, publicly available data is usually obtained from public sample libraries for analysis.

[0014] In some embodiments, the public sample data includes scRNA-seq and scATAC-seq sequencing data. Typically, both scRNA-seq and scATAC-seq sequencing data are required for analysis in this application.

[0015] In some embodiments, the cell type includes at least one of monocytes, T cells, dendritic cells, natural killer cells, and B cells.

[0016] In some embodiments, the subpopulation of monocytes includes at least one of CD14+ monocytes and CD16+ monocytes; and / or,

[0017] The T cell subset includes at least one of CD8+ T cells and CD4+ T cells; and / or,

[0018] The subgroup of dendritic cells includes at least one of plasmacytoid dendritic cells and classical dendritic cells.

[0019] In some embodiments, the differentially expressed gene and differentially expressed chromatin accessibility peak analysis of the cell subpopulations of the sample to be analyzed, to obtain differentially expressed gene data and differentially expressed chromatin accessibility peak data of the cell subpopulations of the sample to be analyzed, includes:

[0020] Group the cell subpopulations of the sample to be analyzed;

[0021] Differentially expressed genes and differentially expressed chromatin peaks were analyzed in the cell subpopulations of the grouped samples to obtain differentially expressed gene data and differentially expressed chromatin peak data of the cell subpopulations of the samples to be analyzed.

[0022] By grouping the cell subpopulations of the sample to be analyzed, and obtaining differentially expressed gene data and differentially expressed chromatin peak data of the cell subpopulations based on the grouping, the differences in differentially expressed genes and differentially expressed chromatin peaks between different groups can be more accurately identified, thus improving the accuracy of the analysis.

[0023] It should be noted that the analysis can be performed by grouping cell subpopulations of all cell types to achieve full data analysis, or by selecting cell subpopulations of some or specific cell types based on empirical values ​​to achieve partial data analysis.

[0024] In some embodiments, grouping the cell subpopulations of the sample to be analyzed includes:

[0025] Cell subpopulations of the samples to be analyzed were grouped according to their health status.

[0026] Typically, the sample data to be analyzed includes two health conditions: diseased and healthy. Grouping the cell subpopulations of the sample to be analyzed according to the health condition can reveal the differences in differentially expressed genes and differentially accessible chromatin peaks between diseased and healthy individuals.

[0027] In some embodiments, the analysis of potential regulatory regions and target genes of transcriptional regulatory elements, and the screening of transcriptional elements with high overlap between regulatory regions and target genes, are identified as target regulatory elements:

[0028] The core nucleotide sequence of the motif motif that the target transcriptional regulatory element SMARCC1 binds to includes SEQ ID NO. 5.

[0029] In some embodiments, after analyzing the potential regulatory regions and target genes of transcriptional regulatory elements and screening for transcriptional elements with highly overlapping regulatory regions and target genes (i.e., the target regulatory elements), the method further includes:

[0030] The interactions between multiple target transcriptional regulatory elements were verified using co-immunoprecipitation; and / or,

[0031] The regulation of target genes by target transcriptional regulatory elements was verified using chromatin immunoprecipitation qPCR.

[0032] Verifying the interactions between multiple target transcriptional regulatory elements using immunoprecipitation can validate the accuracy of the analytical methods used in this application.

[0033] Verifying the regulation of target genes by target transcriptional regulatory elements using chromatin immunoprecipitation qPCR can validate the accuracy of the analytical methods used in this application.

[0034] In some embodiments, the verification of the regulation of the target transcriptional regulatory element on the target gene by chromatin immunoprecipitation qPCR includes verifying the regulatory specificity of JUND on the TIMP1 gene promoter region. The primers used to verify the regulatory specificity of JUND on the TIMP1 gene promoter region include an upstream primer and a downstream primer, wherein:

[0035] The upstream primer sequence is shown in SEQ ID NO.1;

[0036] The downstream primer sequence is shown in SEQ ID NO.2.

[0037] Secondly, this application provides an application of a method for screening transcriptional regulatory elements in identifying therapeutic targets for diseases. When a target transcriptional regulatory element is identified, it is considered an upstream regulatory element related to the disease being analyzed. Acting on this upstream regulatory element becomes a therapeutic target for the disease, thus improving drug development progress. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 This is a flowchart of a method for screening transcriptional regulatory elements according to an embodiment of this application.

[0040] Figure 2 This is a flowchart of a method for screening transcriptional regulatory elements according to an embodiment of this application.

[0041] Figure 3 This is a flowchart of a method for screening transcriptional regulatory elements according to an embodiment of this application.

[0042] Figure 4 This is a flowchart of a method for screening transcriptional regulatory elements according to an embodiment of this application.

[0043] Figure 5This is a dimensionality-reduced clustering diagram of peripheral blood mononuclear cell scRNA-seq data in the embodiments of this application.

[0044] Figure 6 This is a dimensionality-reduced clustering diagram of peripheral blood mononuclear cell scATAC-seq data in the embodiments of this application.

[0045] Figure 7 This is a dimensionality-reduced clustering diagram of scATAC-seq data of mononuclear cells in the embodiments of this application.

[0046] Figure 8 This is a dimensionality-reduced clustering diagram of monocyte scRNA-seq data in the embodiments of this application.

[0047] Figure 9 This is a peak-volcano plot showing the differences in monocytes between COVID-19 patients and healthy volunteers in the embodiments of this application.

[0048] Figure 10 This is a volcano diagram of differentially expressed genes in monocytes of COVID-19 patients and healthy volunteers, as shown in the embodiments of this application.

[0049] Figure 11 This is a correlation diagram between genes and chromatin accessibility peaks in the embodiments of this application.

[0050] Figure 12 This is a motif enrichment map of a highly accessible region of monocytes in a COVID-19 patient, as shown in the embodiments of this application.

[0051] Figure 13 This is a diagram showing the activity of SMARCC1 and JUND in monocyte subclusters in the embodiments of this application.

[0052] Figure 14 This is a diagram of the motif nucleotide sequence of SMARCC1 and JUND binding in the embodiments of this application.

[0053] Figure 15 This represents the intersection of the potential regulatory regions of SMARCC1 and JUND in the embodiments of this application.

[0054] Figure 16 This is a bar chart showing the functional enrichment analysis of target genes in the potential regulatory regions of SMARCC1 and JUND in the embodiments of this application.

[0055] Figure 17 These are potential binding sites for SMARCC1 and JUND in the genomic region of the pro-inflammatory cytokine TIMP1 in the embodiments of this application.

[0056] Figure 18 This is a Co-IP diagram showing the combination of SMARCC1 and JUND in an embodiment of this application.

[0057] Figure 19 This is a ChIP-qPCR diagram of JUND regulating TIMP1 in an embodiment of this application. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with the embodiments of this application. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0059] Gene expression regulation is a multi-level, precise regulatory process encompassing multiple key stages, from transcription initiation and post-transcriptional processing to translation and post-translational modifications. This multi-level regulatory network influences cellular function and state by precisely controlling gene expression. As core executors of this regulatory network, trans-acting factors, represented by transcription factors, chromatin remodeling complexes, and RNA polymerases, form dynamic regulatory modules through complex synergistic interactions. Although existing research has revealed the important functions of these regulatory elements, the construction of their regulatory networks and the patterns of their interactions remain unclear. Currently, there is a lack of systematic methods for predicting their synergistic effects. Analyzing the hierarchical structure and operational patterns of gene regulatory networks can not only elucidate the potential mechanisms of transcriptional regulation in eukaryotes but also provide methods for identifying key regulatory nodes in diseases and for in-depth analysis of pathological mechanisms, which is of great significance for the screening of therapeutic targets.

[0060] Breakthroughs in single-cell multi-omics technologies have provided unprecedented resolution for analyzing gene expression regulation. SCENIC et al. constructed inference models of transcriptional regulation based on single-cell scRNA-seq co-expression networks, enabling systematic prediction of core transcription factors driving gene expression and their targets in cell subpopulations. Li Hao et al. developed a method and device for predicting single-cell transcription factor regulatory networks, which can solve the problem of predicting transcription factor regulatory networks at the single-cell level using scATAC-seq data. Furthermore, Zhang Yongqing et al. invented a method for predicting single-cell transcription factors based on deep learning and attention mechanisms. Although existing research has explored single-omics data or specific methods, effectively integrating multiple omics data to comprehensively capture the spatiotemporal characteristics of transcription factor regulation remains a key challenge. By integrating single-cell transcriptomics (scRNA-seq), single-cell chromatin accessibility (scATAC-seq), and epigenetic modification data, we can better capture the spatiotemporal coupling relationship between transcription factor binding activity, chromatin remodeling state, and target gene expression at the single-cell level, revealing the specific combination rules of regulatory elements in a cellular heterogeneous context. However, such methods are currently scarce, and there are also few methods for systematically analyzing cross-level regulatory events such as transcription factor synergy and chromatin remodeling complex interactions.

[0061] Cells precisely regulate chromatin topological state through ATP-driven chromatin remodeling complexes to overcome the physical limitations of higher-order three-dimensional structures on the reading of genetic information. To date, four distinct families of chromatin remodeling complexes have been identified in humans: the SWI / SNF, ISWI, CHD, and INO80 families. The SWI / SNF subfamily (pBAF, cBAF, and ncBAF) consists of shared subunits such as SMARCC1, SMARCA2 / 4, and SMARCD1, as well as subfamily-specific subunits such as BRD9, BRD7, and ARID1A / B. SMARCC1 is the core subunit of the SWI / SNF complex, containing three domains: SWIRM, SANT, and CC. The CC domain belongs to DNA-binding chromatin targeting regions. The SWI / SNF complex dynamically regulates gene transcription through nucleosome remodeling (sliding / expulsion / histone variant exchange). The SWI / SNF complex not only synergizes with transcription factors to achieve specific regulation of target genes (such as activation of inflammation-related genes) but also forms an epigenetic synergistic network with histone modifications. Targeting the dynamic assembly or functional modules of the SWI / SNF complex may become a novel therapeutic strategy for regulating pathological inflammation.

[0062] Immune dysregulation triggered by SARS-CoV-2 infection has become a key feature of the pathological process of COVID-19. Studies have shown that SARS-CoV-2 can induce an abnormal cytokine storm, characterized by the excessive release of pro-inflammatory cytokines such as IL-6, TNF-α, and CCL2. This uncontrolled inflammatory response can lead to lung tissue damage, systemic multi-organ dysfunction, and failure. Notably, this persistent disruption of immune homeostasis may, through long-term activation of immune response mechanisms, become an important pathological basis for long-term sequelae of pneumonia caused by SARS-CoV-2 (Long COVID). Immune cells such as alveolar macrophages and monocytes produce and release inflammatory cytokines such as IL-1β, IL-6, and TNF-α, thereby actively participating in the inflammatory response and promoting the occurrence and development of inflammatory responses in COVID-19 patients. Currently, there are still gaps in our understanding of the key regulatory nodes in the overactivation process of immune cells such as monocytes, particularly regarding the core transcription factors driving the production of inflammatory genes, epigenetic regulatory elements, and the interaction networks between signal transduction pathways, which have not been fully elucidated. In-depth analysis of transcriptional regulatory elements involved in the production of inflammatory genes in monocytes during SARS-CoV-2 infection may provide new targets for the prevention and treatment of inflammatory responses.

[0063] In view of this, this application provides a method for screening transcriptional regulatory elements and its application, so as to provide a method for screening transcriptional regulatory elements.

[0064] Firstly, such as Figure 1 As shown, this application provides a method for screening transcriptional regulatory elements, comprising the following steps:

[0065] S100. Denoise the scRNA-seq and scATAC-seq data of the sample to be analyzed to obtain preprocessed scRNA-seq and scATAC-seq data.

[0066] S200. Perform dimensionality reduction clustering and cell annotation on the preprocessed scRNA-seq and scATAC-seq data to obtain at least one cell type of the sample to be analyzed.

[0067] S300. Perform dimensionality reduction clustering and cell annotation on at least one type of cell in the sample to be analyzed to obtain a cell subpopulation of that type of cell.

[0068] S400. Perform differentially expressed gene and differentially expressed chromatin accessibility peak analysis on the cell subpopulation of the sample to be analyzed to obtain differentially expressed gene data and differentially expressed chromatin accessibility peak data of the cell subpopulation of the sample to be analyzed.

[0069] S500. Perform scRNA-seq and scATAC-seq combined analysis on differentially expressed gene data and differential chromatin accessibility peak data of cell subpopulations of the sample to be analyzed to identify differentially expressed genes regulated by chromatin accessibility in cell subpopulations of the sample to be analyzed.

[0070] S600. Perform enrichment analysis on differential chromatin accessibility peak data of cell subpopulations of the sample to be analyzed to identify transcriptional elements involved in the regulation of gene expression related to chromatin accessibility peaks.

[0071] S700: Analyze the potential regulatory regions and target genes of transcriptional regulatory elements, and screen for transcriptional elements whose regulatory regions and target genes highly overlap, which are the target regulatory elements.

[0072] This application provides a method for predicting synergistic transcriptional regulatory elements based on a combination of single-cell multi-omics and experimental investigation. This method can reveal potential synergistic effects among transcriptional regulatory elements, interactions of chromatin remodeling complexes, and other cross-level regulatory events, providing a method for identifying key regulatory nodes in diseases. On the other hand, based on single-cell multi-omics sequencing technology, the application systematically reveals key transcriptional regulatory elements involved in disease responses. By integrating single-cell chromatin accessibility analysis and transcriptome sequencing data, the core transcription factors driving disease, epigenetic regulatory elements, and interactions between signal transduction pathways are clarified, successfully constructing a monocyte-specific epigenetic-transcriptional regulatory network.

[0073] In conjunction with the first aspect, in some embodiments provided in this application, the sample to be analyzed includes public sample data. Typically, the sample to be analyzed consists of a large amount of scRNA-seq and scATAC-seq sequencing data. Since scRNA-seq and scATAC-seq sequencing work requires collaboration among multiple teams, publicly available data is usually obtained from public sample libraries for analysis.

[0074] In conjunction with the first aspect, in some embodiments provided in this application, the public sample data includes scRNA-seq and scATAC-seq sequencing data. Typically, scRNA-seq and scATAC-seq sequencing data are required for the analysis in this application within the public sample data.

[0075] In conjunction with the first aspect, in some embodiments provided in this application, the cell type includes at least one of monocytes, T cells, dendritic cells, natural killer cells, and B cells.

[0076] In conjunction with the first aspect, in some embodiments provided in this application, the subpopulation of monocytes includes at least one of CD14+ monocytes and CD16+ monocytes.

[0077] In conjunction with the first aspect, in some embodiments provided in this application, the T cell subpopulation includes at least one of CD8+ T cells and CD4+ T cells.

[0078] In conjunction with the first aspect, in some embodiments provided in this application, the subpopulation of dendritic cells includes at least one of plasmacytoid dendritic cells and classical dendritic cells.

[0079] In conjunction with the first aspect, such as Figure 2 As shown, in some embodiments provided in this application, the step of performing differentially expressed gene and differentially expressed chromatin accessibility peak analysis on the cell subpopulation of the sample to be analyzed, to obtain differentially expressed gene data and differentially expressed chromatin accessibility peak data of the cell subpopulation of the sample to be analyzed, includes:

[0080] S401. Group the cell subpopulations of the sample to be analyzed;

[0081] S402. Perform differentially expressed gene and differentially expressed chromatin accessibility peak analysis on the cell subpopulations of the grouped samples to be analyzed, and obtain differentially expressed gene data and differentially expressed chromatin accessibility peak data of the cell subpopulations of the samples to be analyzed.

[0082] By grouping the cell subpopulations of the sample to be analyzed, and obtaining differentially expressed gene data and differentially expressed chromatin peak data of the cell subpopulations based on the grouping, the differences in differentially expressed genes and differentially expressed chromatin peaks between different groups can be more accurately identified, thus improving the accuracy of the analysis.

[0083] It should be noted that the analysis can be performed by grouping cell subpopulations of all cell types to achieve full data analysis, or by selecting cell subpopulations of some or specific cell types based on empirical values ​​to achieve partial data analysis.

[0084] Typically, the sample data to be analyzed includes two health conditions: diseased and healthy. Grouping the cell subpopulations of the sample to be analyzed according to the health condition can reveal the differences in differentially expressed genes and differentially accessible chromatin peaks between diseased and healthy individuals.

[0085] In conjunction with the first aspect, in some embodiments provided in this application, the analysis of potential regulatory regions and target genes of transcriptional regulatory elements, and the screening of transcriptional elements in which the regulatory regions and target genes highly overlap, are the target regulatory elements. Among them, the core nucleotide sequence of the motif motif bound by the target transcriptional regulatory element SMARCC1 includes SEQ ID NO.5, and the base composition of the nucleotide sequence of SEQ ID NO.5 is specifically TGACTCA.

[0086] In conjunction with the first aspect, such as Figure 3As shown in some embodiments provided in this application, after analyzing the potential regulatory regions and target genes of transcriptional regulatory elements and screening for transcriptional elements with highly overlapping regulatory regions and target genes, which are the target regulatory elements, the method further includes:

[0087] S800, using immunoprecipitation to verify the interactions between multiple target transcriptional regulatory elements.

[0088] Verifying the interactions between multiple target transcriptional regulatory elements using immunoprecipitation can validate the accuracy of the analytical methods used in this application.

[0089] In conjunction with the first aspect, such as Figure 4 As shown in some embodiments provided in this application, after analyzing the potential regulatory regions and target genes of transcriptional regulatory elements and screening for transcriptional elements with highly overlapping regulatory regions and target genes, which are the target regulatory elements, the method further includes:

[0090] S900. The regulation of target genes by target transcriptional regulatory elements was verified by chromatin immunoprecipitation qPCR.

[0091] Verifying the regulation of target genes by target transcriptional regulatory elements using chromatin immunoprecipitation qPCR can validate the accuracy of the analytical methods used in this application.

[0092] In conjunction with the first aspect, in some embodiments provided in this application, the verification of the regulation of the target transcriptional regulatory element on the target gene by chromatin immunoprecipitation qPCR includes verifying the regulatory specificity of JUND on the TIMP1 gene promoter region. The primers used to verify the regulatory specificity of JUND on the TIMP1 gene promoter region include an upstream primer and a downstream primer, wherein:

[0093] The upstream primer sequence is shown in SEQ ID NO.1, and the nucleotide sequence of SEQ ID NO.1 has the base composition of CAGGGAGAGGGAGAGGAGG;

[0094] The downstream primer sequence is shown in SEQ ID NO.2, and the specific base composition of the nucleotide sequence of SEQ ID NO.2 is CGGAAACCACAGGCCTCC.

[0095] Secondly, this application provides an application of a method for screening transcriptional regulatory elements in identifying therapeutic targets for diseases. When a target transcriptional regulatory element is identified, it is considered an upstream regulatory element associated with the disease to be analyzed. Acting on this upstream regulatory element becomes a therapeutic target for the disease, thus improving drug development progress. The diseases to be analyzed include, but are not limited to, infectious diseases (such as COVID-19, sepsis) and inflammatory diseases.

[0096] The technical solution provided in this application will be described in detail below with reference to embodiments. Taking peripheral blood mononuclear cell data from healthy and infected individuals with COVID-19 as an example, the data source is the public announcement datasets GSE206283 and GSE206455.

[0097] Example 1

[0098] scRNA-seq and scATAC-seq were used to identify cell types in peripheral blood mononuclear cells.

[0099] (1) Cell quality control of scRNA-seq data

[0100] The scRNA-seq data were processed using the Seurat software package. The integration of scRNA-seq data was performed using an anchor-based integration method. To remove low-quality cells, cells were selected based on criteria including a mitochondrial expression rate of less than 10%, a gene expression count between 200 and 4000, and a UMI count of less than 30,000. Additionally, cells with fewer than three cells expressing a gene were excluded.

[0101] (2) Dimensionality reduction and cell annotation of scRNA-seq data

[0102] The scRNA-seq data were normalized and scaled using the Seurat function "Normalization". Principal component analysis (PCA) was performed based on the top 2000 most variable features identified by the VST method. Appropriate PCA dimensions were selected, and cell populations were identified using the "FindNeighbors" and "FindClusters" functions. Cell subpopulations were visualized using a uniform manifold approximation and projection. A dual-strategy approach was employed to annotate scRNA-seq cell subpopulations: first, automatic annotation of cell subpopulations was performed using the SingleR package; then, the specific subpopulation type was further determined based on the highly expressed genes in each subpopulation. Cell annotation results are shown below. Figure 5 .

[0103] (3) Cell quality control of scATAC-seq data

[0104] The scATAC-seq data was processed using the ArchR software package. To remove low-quality cells and duplicate cells, the fragment files from the scATAC-seq data were first imported into R Studio. Cells were then screened based on the criteria of having more than 1000 unique fragments per cell and an enrichment level of transcription start sites greater than 4. After screening, Arrow files were created and merged to obtain the ArchR Project for downstream analysis.

[0105] (4) Dimensionality reduction of scATAC-seq data cells

[0106] The peak matrix is ​​created using the functions "addGroupCoverages", "addReproduciblePeakSet", and "addPeakMatrix" and added to the ArchRProject project. The peak matrix and gene expression matrix are extracted for subsequent analysis. Seurat objects are created based on the peak matrix. Dimensionality reduction clustering is performed using scopen. Batch effects between samples are corrected using the harmony method. After removing batch effects, cell clusters are identified using the functions "FindNeighbors" and "FindClusters". Subsequently, the cell clusters, uniform manifold approximation, and projection results are added to the ArchRProject project.

[0107] (5) Cell annotation and cell subpopulation extraction from scATAC-seq data

[0108] The `getMarkerFeatures` function was used to retrieve the specific genes expressed by each cell population. Next, cell type annotations were performed based on the specific genes expressed by each cell population; for example, cell populations highly expressing MS4A1, CD79A, and CD79B were annotated as B cells, and cell populations highly expressing FCGR3A, LYZ, CD14, and MS4A7 were annotated as monocytes, etc. The cell annotation results are shown below. Figure 6 .

[0109] Example 2

[0110] Combined analysis of differentially expressed chromatin peaks and differentially expressed genes identified genes regulated by chromatin accessibility (using monocytes from COVID-19 patients as an example).

[0111] (1) Dimensionality reduction clustering and annotation of monocyte subsets in scATAC-seq and scRNA-seq

[0112] For scATAC-seq data, mononuclear cells were extracted from the ArchRProject project using the "subset" function and then subjected to dimensionality reduction clustering. First, dimensionality reduction was performed using the "addIterativeLSI" function based on the first 25,000 variable features and the first 30 dimensions. Subsequently, cells were clustered at an appropriate resolution (e.g., 0.8), and low-dimensional visualization of the cells was achieved using a uniform manifold approximation and projection. Results are shown below. Figure 7 .

[0113] For scRNA-seq data, the "subset" function was used to extract monocyte-related information from Seurat objects, and then the data was re-clustered with reduced dimensionality. First, the data was re-standardized and scaled. Then, principal component analysis was performed on the 2000 most variable genes identified by the VST method. Appropriate principal component analysis dimensions were selected, and the "FindNeighbors" and "FindClusters" functions were used to identify cell populations. The uniform manifold approximation and projection were used to visualize the cell subpopulations. Results are shown below. Figure 8 .

[0114] (2) scATAC-seq differential chromatin accessibility peak analysis

[0115] Group labels were added to the ArchRProject project to divide the data into COVID-19 and Healthy groups for differential chromatin accessibility peak analysis. To enable chromatin accessibility analysis between different groups, the "addGroupCoverages" function was used to retrieve reproducible peaks based on the group, and the "addReproduciblePeakSet" function of the Macs2 algorithm was used to create a reproducible peak matrix by group. After the peak matrix was created, the Wilcoxon rank-sum test was used as the validation method, and the "getMarkerFeatures" function was used to filter differentially accessible peaks that met the criteria (e.g., FDR < 0.1 and |Log2FC| > 0.5). Among these, inflammatory cytokines such as TIMP1, IL21R, and CCRL2 showed increased expression levels in monocytes of COVID-19 patients. Results are shown below. Figure 9 .

[0116] (3) scRNA-seq differential gene expression analysis

[0117] Add grouping labels to the monocyte seruat objects to divide the data into COVID-19 and Healthy groups for differentially expressed gene analysis. Use the grouping results as cell labels and the "FindAllMarkers" function to retrieve differentially expressed genes between the COVID-19 and Healthy groups. See results below. Figure 10 Among them, inflammatory cytokines such as TIMP1, IL15, IL21R, CCRL2, and IFI30 were expressed at increased levels in monocytes of COVID-19 patients.

[0118] (4) Joint analysis of scATAC and scRNA-seq data

[0119] To perform joint analysis of scATAC-seq and scRNA-seq data, the "FindTransferAnchor" function (Seurat function) and the "addGeneIntegrationMatrix" function (ArchR function) were used to match the cell-independent gene score matrix (from scATAC-seq data) and the gene expression matrix (from scRNA-seq data), and all monocytes from scATAC-seq were integrated with those from scRNA-seq. The gene integration scores were then matched to each cell to create a gene integration matrix.

[0120] (5) Analysis of peak accessibility and determination of peak-to-gene links

[0121] The co-accessibility correlation between peaks in scATAC-seq monocyte data was calculated using the "addCoAccessibility" function. Furthermore, the correlation between peak accessibility and gene expression in monocytes was calculated using the "addPeak2GeneLinks" function. Peaks with CoAccessibility > 0.2 and P < 0.05 were considered to have potential regulatory relationships with genes. Results are shown below. Figure 11 .

[0122] Example 3

[0123] Motif enrichment analysis screens transcriptional elements that are synergistically involved in gene expression regulation.

[0124] (1) Motif enrichment analysis of highly accessible chromatin regions

[0125] To analyze potential transcriptional regulatory elements involved in gene expression in monocytes of COVID-19 patients, after identifying differentially accessible chromatin peaks, we first used the "addmotifAnnotations" function of the ArchR software package to systematically annotate transcription factor binding motifs in highly open chromatin regions of COVID-19 patient monocytes based on the "CIS-BP" regulatory element database, generating motif-Matches-In-Peaks objects. Subsequently, we used the "peakAnnoEnrichment" function to quantitatively analyze significantly enriched regulatory motifs in differentially open regions, calculated the bias of all motifs in each cell using the "addDeviationsMatrix" function of the HOMER software, and calculated the motif accessibility bias value in each sample using the ChromVAR algorithm to measure transcription element activity. The enrichment patterns of key transcription factors were visualized using ggplot2. Results are shown below. Figure 12 and Figure 13Motif enrichment analysis showed that motifs binding to transcriptional regulatory elements such as SMARCC1 and JUND were significantly enriched in highly accessible chromatin regions of monocytes from COVID-19 patients. Furthermore, the activities of SMARCC1 and JUND showed significant consistency across monocyte clusters, particularly in the C4 and C7 clusters, where their activity was markedly upregulated.

[0126] (2) Analysis of Transcription Element Targeting Regions

[0127] To visualize the specific nucleotide sequences of motifs binding to regulatory elements such as SMARCC1 and JUND, PPM matrix information corresponding to these elements was extracted from the peakAnnotation object in the ArchRProject project and converted into PWM matrices. The seqLogo package was used for visualization, and transcriptional regulatory elements with highly consistent motif nucleotide sequences were screened. Results are shown below. Figure 14 The nucleotide sequence identification map of the motif shows that the core nucleotide sequence of the motif motif bound by JUND and SMARCC1 is “TGACTCA” (SEQ ID NO.5).

[0128] To identify potential regulatory regions controlled by regulatory elements such as JUND and SMARCC1, chromatin regions containing motifs binding to transcriptional regulatory elements such as SMARCC1 and JUND were extracted from the generated motif-Matches-In-Peaks objects. Intersections of these regions were then performed to screen for transcriptional regulatory elements with highly overlapping regulatory regions. Results are shown below. Figure 15 Simultaneously, target genes of potential regulatory regions of transcriptional regulatory elements were analyzed, and their functions were analyzed using GO and KEGG enrichment analyses. Results are shown below. Figure 16 The results showed that SMARCC1 and JUND potentially regulate 68.1% of their regions, and their target genes are highly correlated with cytokine production. Furthermore, SMARCC1 and JUND have binding sites in the promoter regions of pro-inflammatory cytokines such as TIMP1, IL21R, and CCRL2. See below for further details. Figure 17 These results suggest that JUND and SMARCC1 may synergistically participate in the expression of inflammatory genes such as TIMP1, IL21R, and CCRL2.

[0129] Example 4

[0130] Co-IP and ChIP-qPCR were used to verify the interactions between transcriptional elements and the regulation of potential target genes by these elements.

[0131] Co-IP validates interactions between transcriptional elements

[0132] (1) Take 1×10⁻⁶ of plants in good growth condition 7 THP-1 cells were added to pre-chilled IP lysis buffer containing 1× PMSF protease inhibitor. After cell lysis, the cell suspension was collected in 1.5 mL centrifuge tubes and centrifuged at 12000–16000×g in a low-temperature ultracentrifuge at 4°C for 15 min. The supernatant was collected. Protein concentration was determined by the BCA method.

[0133] (2) Take 100 μL of the extracted protein solution (supernatant) as the forward and reverse input controls (50 μL each). Add 3-5 μL of SMARCC1 and JUND antibodies to the remaining protein solution and incubate overnight at 4°C. After overnight antibody incubation, add 0.25 mg of Protein A / G Magnetic Beads (the magnetic beads have been washed and resuspended with lysis buffer) and incubate on a shaker at room temperature for 60 min.

[0134] (3) Place the centrifuge tube in a magnetic rack to allow the magnetic beads to gather to one side of the tube, and aspirate the supernatant. Add lysis buffer to the centrifuge tube, vortex to mix, collect the magnetic beads through the magnetic rack, discard the supernatant, and repeat the washing process 2-3 times. To perform protein denaturation elution, add 100 μL of loading buffer (1×) to the centrifuge tubes containing the input and the sample, respectively. Then, heat the sample in a 100°C metal bath for 10 min. Separate the magnetic beads using a magnetic rack. Finally, aspirate the supernatant containing the target protein and store it at -20°C for later use.

[0135] (4) Preparation of SDS-PAGE protein gel: First, check the sealing of the gel casting plate. Add distilled water to the assembled plate to ensure it does not leak. Then, discard the water and let the plate dry. Take 2 mL of sterile water and place it in a 15 mL centrifuge tube. Add 1.25 mL of 1.5 M Tris-HCl (pH 8.8), 100 μL of 10% ammonium persulfate solution, 100 μL of 10% SDS, and 1.65 mL of 30% acrylamide. Then, add 5 μL of TEMED solution in a fume hood, mix well, and gently and evenly add it to the electrophoresis tank, avoiding the formation of air bubbles, to prepare a 10% separating gel. Add 4 mL of separating gel to each plate, add water for sealing, and let it stand at room temperature for 20–30 min. After the separating gel solidifies, discard the water. Take 1.1 mL of pure water and place it in a 15 mL centrifuge tube. Add 250 μL of 1M Tris-HCl solution (pH 6.8), 330 μL of 30% acrylamide solution, 20 μL of 10% SDS solution, and 20 μL of 10% APS solution. In a fume hood, add 7 μL of TEMED solution to the centrifuge tube. Finally, add about 1.5 mL of stacking gel to the gel casting plate and insert the comb into the stacking gel. Let it stand at room temperature for 20–30 minutes. After the gel has solidified, vertically remove the comb for use.

[0136] (5) Assemble the prepared gel plate with the electrophoresis tank and pour in 1× electrophoresis buffer. Slowly add approximately 3–5 μL of the prepared protein sample to the corresponding wells, and add 2–3 μL of the reference protein marker to the wells on both sides to facilitate subsequent determination of the target protein size. Load 1× loading buffer into the blank wells. The electrophoresis program consists of two stages: the first stage is at a constant voltage of 80V for 30 min; the second stage is at a constant voltage of 120V for 60 min.

[0137] (6) Cut the PVDF membrane to the appropriate size (0.22μm or 0.45μm) according to requirements and activate it by soaking in methanol. After electrophoresis, remove the separating gel and trim off the excess. Then assemble and place it into the transfer system in the following order: negative electrode (black clip) — sponge — filter paper (three layers) — separating gel — PVDF membrane — filter paper (three layers) — sponge — positive electrode (transparent clip). Add pre-cooled transfer buffer and transfer the membrane under ice bath at a constant current of 200mA for 60min (the time needs to be adjusted according to the size of the target protein).

[0138] (7) Blocking and Antibody Incubation: Immerse the transferred PVDF membrane in a container filled with 5% skim milk and place the container on a shaker at room temperature for 1 hour. Dilute the SMARCC1 and JUND antibodies according to the reagent manufacturer's instructions beforehand. After membrane blocking and washing, cut the bands as needed based on the size and position of the protein bands (provided by the protein marker) and the molecular size of the target protein. After cutting, place the bands in an antibody incubation box, add the corresponding antibody diluent, and incubate overnight at 4°C. After 16 hours of incubation, recover the antibody diluent from the antibody incubation box and add 1×PBST to wash the membrane three times, 5 minutes each time. Dilute the secondary antibody with the antibody diluent according to the reagent manufacturer's instructions, then place the washed bands back into the antibody incubation box, add the corresponding secondary antibody diluent, and incubate on a shaker at room temperature for 60–90 minutes. After incubation, collect the secondary antibody diluent and wash the membrane again with 1×PBST for the same time and number of washes as above. Finally, add 1×PBST to soak the strips.

[0139] (8) Prepare ECL solution according to the reagent supplier's instructions, and use immediately after preparation. Before placing the band into the chemiluminescence imaging system for development and imaging, soak it in ECL solution for 1–2 minutes. Set different exposure times to take pictures and export the corresponding image results. Co-IP results show that there is an interaction between SMARCC1 and JUND, see [link to Co-IP results]. Figure 18 .

[0140] Regulation of potential target genes by ChIP-qPCR transcriptional elements

[0141] (1) Take a bottle of THP-1 cells in good growth condition, transfer the cell suspension to a 1.5 mL centrifuge tube, centrifuge at 500×g for 5 min, discard the culture medium, add 1 mL of 1% formaldehyde solution, and let stand at room temperature for 10 min to allow the cell DNA and protein to crosslink.

[0142] (2) Discard the solution, add 100 μL of 10×glycine solution to the centrifuge, rotate it slightly to mix it, and incubate at room temperature for 5 min.

[0143] (3) Collect cells by centrifuging at 1200×g for 5 min at 4℃, and wash twice with pre-cooled 1mL 1×PBS Buffer. Collect cells by centrifuging at 1200×g for 5 min at 4℃, and remove the supernatant.

[0144] (4) Resuspend the cells in pre-chilled 1 mL of 1×Cell Swelling Buffer and incubate on ice for 10 min, inverting and mixing every 3 min. Centrifuge at 4°C, 5000×g for 5 min to precipitate the cell nuclei. Discard the solution and resuspend the precipitate in 1 mL of ChIP sonication buffer. Incubate on ice for 10 min. Use a 3 mm diameter microprobe of a Xinzhi probe-type sonicator with a power of 150W. Sonicate for 5 seconds, let stand for 5 seconds, and sonicate for about 20 min. Centrifuge at 4°C, 12000×g for 10 min to remove cell nucleus fragments and other impurities. Transfer the supernatant to obtain the cross-linked chromatin sample and determine the sample concentration.

[0145] (5) Add an equal amount of cross-linked chromatin sample to each sample tube and replenish the total volume to 500 μL with 1× sonication buffer. Transfer 25 μL of sample from each sample tube to a new centrifuge tube as a 5% sample input control. Do not perform immunoprecipitation reaction. Store temporarily at -20℃.

[0146] (6) Add the corresponding antibody (2-5 μg) to the target protein sample tube, add normal isotype IgG to the negative control tube, incubate overnight at 4°C, add protein A / G magnetic beads, reseal the sample tube, incubate at 4°C for 2 h, place the centrifuge tube on a magnetic separator rack, and adsorb the protein A / G magnetic beads onto the tube wall. After the solution becomes clear in 1-2 min, aspirate the supernatant.

[0147] (7) Add 1 mL of low-salt wash buffer, incubate at 4°C for 5 min, place the sample back on the magnetic separator, wait 2 min at room temperature until the solution is clear, carefully aspirate the solution, add 1 mL of high-salt wash buffer and wash again. Add 1 mL of TE buffer, vortex and place the sample back on the magnetic separator, wait for the solution to clear, aspirate the supernatant, and repeat the TE buffer wash once.

[0148] (8) Add 125 μL of ChIP elution buffer to a 5% Input tube and 150 μL of ChIP elution buffer to each sample tube. Incubate in a 65°C metal bath for 30 min. Gently vortex the tube every 5 min to elute chromatin from the antibody-protein A / G magnetic beads. Adsorb the protein A / G magnetic beads using a magnetic separator. Once the solution is clear, carefully transfer the supernatant and label it. Add 6 μL of 5M NaCl solution and 2 μL of RNase A solution to all tubes. Incubate at 37°C for 30 min. Then add 2 μL of proteinase K solution and incubate at 65°C for 2 h. After returning to room temperature, proceed to the next step.

[0149] (9) Add the obtained solution vertically into the adsorption column (place the adsorption column in the collection tube), let it stand for 2 minutes, then centrifuge at 12000×g for 30-60 seconds. Discard the waste liquid in the tube and aspirate the liquid along the wall. Place the adsorption column back into the collection tube, add 600 μL of rinsing buffer containing anhydrous ethanol, centrifuge and discard the waste liquid. Place the adsorption column back into the collection tube and wash twice. After drying the adsorption column, place it in a clean centrifuge tube, add 30-60 μL of elution buffer dropwise to the center of the adsorption membrane, let it stand for 2 minutes, then centrifuge at 12000×g for 2 minutes to obtain the DNA solution.

[0150] (10) The concentration and purity of the DNA solution were determined using an ultra-micro spectrophotometer.

[0151] (11) Locate the promoter sequence of the target gene from the NCBI database. It is generally believed that the 2kb upstream region of the gene is the promoter region. Download the 2kb upstream promoter sequence. Locate the binding site of the target transcription factor from the JASPAR database. The prototype is human. Input the 2kb promoter sequence of the target gene to obtain the predicted binding result of the transcription factor to the target gene. Based on the predicted binding site, use the Primer3plus online tool to design primers for each site. The designed product length is approximately 300bp, and each site should be positioned in the middle of its corresponding product. The ChIP primers designed in this invention for the target gene TIMP1 are P1 (upstream primer sequence as shown in SEQ ID NO.1, downstream primer as shown in SEQ ID NO.2) and P2 (upstream primer sequence as shown in SEQ ID NO.3, downstream primer as shown in SEQ ID NO.4).

[0152] SEQ ID NO.1: CAGGGAGAGGGAGGAGG;

[0153] SEQ ID NO.2: CGGAAACCACAGGCCTCC;

[0154] SEQ ID NO.3: GGCACACAGTAGATGCACAATA;

[0155] SEQ ID NO.4: CCAGTCCCCTAAAGTCTCCAG;

[0156] (12) Add the samples, including the corresponding antibody sample group, 0.5% Input group, IgG sample group, and blank control without DNA template, to PCR tubes and label them accordingly. Add 2 μL of the corresponding DNA sample as a template to each tube. The solution system in each tube is 20 μL (10 μL 2×PCR Mix + 2 μL primer + 1 μL DNA + 5 μL ddH2O). Set the program according to the standard two-step qPCR procedure (including melting curve reaction time). After the program runs, calculate the ChIP enrichment efficiency based on the experimental results using the following formula:

[0157] Percentage of Input=0.5%×2 (C[T]Input Sample-C[T]IP Sample)

[0158] In the formula, C[T]InputSample — the copy number of the target gene in the Input sample group;

[0159] In the formula, C[T]IP Sample corresponds to the copy number of the target gene in the antibody sample group.

[0160] See results Figure 19 ChIP-qPCR showed that the transcription factor JUND can regulate the expression of the pro-inflammatory cytokine TIMP1.

[0161] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or devices. The terms "first," "second," and "third," etc., are used to distinguish different objects, etc., and do not indicate a sequence, nor do they limit "first," "second," and "third" to different types.

[0162] In the description of the embodiments of this application, terms such as "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a concrete manner.

[0163] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.

[0164] In some processes described in the embodiments of this application, multiple operations or steps are included in a specific order. However, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or they may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.

[0165] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.

[0166] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for screening transcriptional regulatory elements in COVID-19, characterized in that, Includes the following steps: Denoising is performed on the scRNA-seq and scATAC-seq data of the samples to be analyzed to obtain preprocessed scRNA-seq and scATAC-seq data; Dimensionality reduction clustering and cell annotation were performed on the preprocessed scRNA-seq and scATAC-seq data to obtain at least one cell type of the sample to be analyzed. Perform dimensionality reduction clustering and cell annotation on at least one type of cell in the sample to be analyzed to obtain a cell subpopulation of that type of cell; Differentially expressed genes and differentially expressed chromatin peaks were analyzed in the cell subpopulations of the sample to be analyzed, and differentially expressed gene data and differentially expressed chromatin peak data of the cell subpopulations of the sample to be analyzed were obtained. scRNA-seq and scATAC-seq were used to analyze the differentially expressed gene data and differential chromatin accessibility peak data of the cell subpopulations of the sample to be analyzed, in order to identify the differentially expressed genes regulated by chromatin accessibility in the cell subpopulations of the sample to be analyzed. Enrichment analysis was performed on differential chromatin accessibility peak data of cell subpopulations of the sample to be analyzed to identify transcriptional elements involved in the regulation of gene expression related to chromatin accessibility peaks. Analyze the potential regulatory regions and target genes of transcriptional regulatory elements, and screen for transcriptional elements whose regulatory regions and target genes highly overlap, which are the target regulatory elements. The analysis of potential regulatory regions and target genes of the transcriptional regulatory elements and the screening of transcriptional elements with high overlap of regulatory regions and target genes are the target regulatory elements. Among them, the core nucleotide sequence of the motif motif of the target transcriptional regulatory element JUND and SMARCC1 is identified as SEQ ID NO. 5, and the specific base composition of the nucleotide sequence of SEQ ID NO. 5 is TGACTCA.

2. The method for screening COVID-19 transcriptional regulatory elements as described in claim 1, characterized in that, The samples to be analyzed include public sample data.

3. The method for screening COVID-19 transcriptional regulatory elements as described in claim 2, characterized in that, The public sample data includes scRNA-seq and scATAC-seq sequencing data.

4. The method for screening COVID-19 transcriptional regulatory elements as described in claim 1, characterized in that, The cell types include at least one of monocytes, T cells, dendritic cells, natural killer cells, and B cells.

5. The method for screening COVID-19 transcriptional regulatory elements as described in claim 4, characterized in that: The monocyte subset includes at least one of CD14+ monocytes and CD16+ monocytes; and / or, The T cell subset includes at least one of CD8+ T cells and CD4+ T cells; and / or, The subgroup of dendritic cells includes at least one of plasmacytoid dendritic cells and classical dendritic cells.

6. The method for screening COVID-19 transcriptional regulatory elements as described in claim 1, characterized in that, The differentially expressed genes and differentially expressed chromatin peaks of the cell subpopulations of the sample to be analyzed are analyzed to obtain differentially expressed gene data and differentially expressed chromatin peak data of the cell subpopulations of the sample to be analyzed, including: Group the cell subpopulations of the sample to be analyzed; Differentially expressed genes and differentially expressed chromatin peaks were analyzed in the cell subpopulations of the grouped samples to obtain differentially expressed gene data and differentially expressed chromatin peak data of the cell subpopulations of the samples to be analyzed.

7. The method for screening COVID-19 transcriptional regulatory elements as described in claim 1, characterized in that, The process of analyzing potential regulatory regions and target genes of transcriptional regulatory elements, and screening for transcriptional elements with highly overlapping regulatory regions and target genes (i.e., the target regulatory elements), further includes: The interactions between multiple target transcriptional regulatory elements were verified using co-immunoprecipitation; and / or, The regulation of target genes by target transcriptional regulatory elements was verified using chromatin immunoprecipitation qPCR.

8. The method for screening COVID-19 transcriptional regulatory elements as described in claim 7, characterized in that, The verification of the regulation of the target transcriptional regulatory element on the target gene by chromatin immunoprecipitation qPCR includes verifying the regulatory specificity of JUND on the TIMP1 gene promoter region. The primers used to verify the regulatory specificity of JUND on the TIMP1 gene promoter region include an upstream primer and a downstream primer, wherein: The upstream primer sequence is shown in SEQ ID NO.1; The downstream primer sequence is shown in SEQ ID NO.2.