A method for constructing a three-dimensional gene immunoregulation network of fish
By combining in situ Hi-C technology and immunotranscriptomics, a directed weighted three-dimensional gene immune regulatory network for fish was constructed. This solved the problem that existing technologies could not capture the long-range spatial interaction between enhancers and promoters, and enabled the construction of a high-confidence regulatory network, providing target information for molecular breeding of fish to resist diseases.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF AQUATIC LIFE ACAD SINICA
- Filing Date
- 2026-06-17
- Publication Date
- 2026-07-31
AI Technical Summary
Existing methods for constructing fish immune regulatory networks cannot capture the long-range spatial interactions between enhancers and promoters, resulting in the omission of key information on the regulation of immune gene expression by distant regulatory elements.
In situ Hi-C technology was used to obtain chromatin spatial interaction information of the whole genome. Combined with immunotranscriptomics, chromatin compartmentalization, topologically related domain identification, and significant chromatin loop detection were performed. By using the intersection analysis of promoter anchors and enhancer feature markers, a directed weighted three-dimensional gene immune regulatory network of fish was constructed, and the regulatory direction was inferred by transcription factor binding motif enrichment analysis.
This study enables the analysis of immune regulatory relationships from a three-dimensional spatial conformation perspective, accurately extracts promoter-enhancer interaction pairs with functional regulatory potential, constructs a high-confidence directed weighted regulatory network, and provides direct three-dimensional target information, thus supporting molecular breeding for disease resistance in fish fry and fingerlings.
Smart Images

Figure CN122493959A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fish immune three-dimensional genomics technology, specifically a method for constructing a fish three-dimensional gene immune regulatory network. Background Technology
[0002] This invention relates to the fields of aquatic animal molecular immunology and three-dimensional genomics, specifically to a method for constructing a three-dimensional gene immune regulatory network in fish. Systematic analysis of fish immune regulatory networks is a crucial foundation for molecular breeding of disease-resistant fish fry and fingerlings. This method integrates three-dimensional spatial interaction information of whole-genome chromatin obtained under immune stimulation conditions with dynamic transcriptome expression data to construct a directed weighted regulatory network centered on promoter and enhancer spatial interactions. This network is used to reveal the long-range gene regulatory mechanisms in the fish immune response process, providing technical support for molecular target screening of disease resistance traits.
[0003] Existing methods for constructing fish immune regulatory networks primarily rely on transcriptome sequencing technology, inferring regulatory relationships by calculating co-expression correlations between differentially expressed genes after immune stimulation or based on linear associations of neighboring genes. However, in eukaryotes, distal regulatory elements such as enhancers typically require spatial proximity to target gene promoters through three-dimensional chromatin folding to perform their regulatory functions. Transcriptome-based linear co-expression techniques cannot capture the long-range spatial interactions between enhancers and promoters, resulting in the omission of crucial information regarding the regulation of immune gene expression by numerous distal regulatory elements. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a method for constructing a three-dimensional gene immune regulatory network for fish, which solves the problems of its inability to capture the long-range spatial interaction between enhancers and promoters and the omission of key information on the regulation of immune gene expression by distant regulatory elements.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for constructing a three-dimensional gene immune regulatory network in fish, comprising the following steps:
[0006] S1, Sample collection and multi-omics library preparation: Selected fish fry or fingerlings that have been immune-stimulated, collected immune tissues and performed in situ Hi-C cross-linking treatment and RNA extraction, respectively, to construct a three-dimensional gene interaction DNA library and a transcriptome sequencing library.
[0007] S2, Multi-omics Data Acquisition and Immune Gene Screening: High-throughput sequencing was performed on the library obtained in S1 to obtain a standardized chromatin interaction frequency matrix and gene expression data; using control samples without immune stimulation as a reference, differentially expressed genes were screened and enriched through immune pathways to obtain a set of immune-related differentially expressed genes.
[0008] S3, Three-dimensional genome interaction identification: Based on the chromatin interaction frequency matrix of S2, chromatin compartmentalization, topologically related domain identification, and significant chromatin loop detection are performed to extract loop anchor point coordinates; the intersection analysis of anchor points with transcription start sites and enhancer feature markers of immune-related differentially expressed genes is performed to obtain the set of interaction relationships between immune gene promoters and enhancers;
[0009] S4, Three-dimensional immune regulation network modeling: Based on the promoter and enhancer interaction relationship set of S3, with immune genes as nodes and three-dimensional interactions as edges, an initial network is constructed by combining gene expression level and interaction strength; the regulatory direction is inferred by enhancing sequence transcription factor combined with motif enrichment analysis, forming a directed weighted fish three-dimensional gene immune regulation network.
[0010] S5, Network Validation and Iterative Optimization: Repeated validation was performed using individuals with different immune phenotypes, and functional validation was performed on key interactions. False positives were eliminated and missing associations were added based on the validation results. The final network was obtained after iterative optimization.
[0011] By employing the above-mentioned technical solutions, and organically integrating three-dimensional genomics analysis with immunotranscriptomics, the in situ Hi-C technology captures whole-genome chromatin spatial interaction information. Combined with differentially expressed gene screening and pathway enrichment filtering under immune stimulation conditions, it achieves the analysis of immune regulatory relationships from a three-dimensional spatial conformation perspective. Through hierarchical structural analysis of chromatin compartmentalization, TAD identification, and significant chromatin loop detection, and by utilizing precise intersection analysis of promoter anchors and enhancer feature markers, promoter-enhancer interaction pairs with functional regulatory potential are accurately extracted from massive three-dimensional interaction data. The regulatory direction is inferred through transcription factor-motif enrichment analysis, and edge weights are assigned by combining gene expression dynamic data, upgrading the undirected interaction network into a directed weighted regulatory network. Through a complementary verification system of repeated validation in individuals with different immune phenotypes and molecular functional experiments, false positives are eliminated and missing associations are supplemented. Therefore, a high-confidence directed weighted regulatory network that accurately reflects the three-dimensional spatial regulatory relationships of fish immune genes is obtained, providing direct three-dimensional target information for molecular breeding of disease resistance in fish fry and fingerlings.
[0012] Preferably, in S1, the immune challenge is performed by pathogen immersion infection or injection of immune stimulants, and samples are taken at multiple time points at 24 hours, 48 hours and 72 hours after challenge; the immune tissue is selected from at least one of the head kidney, spleen, liver, gills and intestinal mucosa; the fish fry include larvae and juveniles, and the fish species are juveniles.
[0013] By adopting the above technical solution, sampling at three time points (24 hours, 48 hours, and 72 hours) covers the early identification and signal activation stages, the stage of large-scale expression of effector molecules, and the regulation and decline stages of the innate immune response in fish. The selection of immune-related tissues such as the head kidney, spleen, liver, gills, and intestinal mucosa encompasses central immune organs, peripheral immune organs, and mucosal immune barriers. Therefore, it is possible to comprehensively capture the dynamic evolution of the immune regulatory network and the spatial regulatory characteristics of different immune tissues, providing initial samples with complete spatiotemporal information on the immune response for subsequent multi-omics analysis.
[0014] Preferably, in step S2, the in-situ Hi-C crosslinking process includes: formaldehyde crosslinking fixation, sucrose gradient centrifugation to separate cell nuclei, chromatin digestion using tetrabasic restriction endonucleases MboⅠ or DpnⅡ, biotin-labeled blunt ends, adjacent ligation, decrosslinking, and DNA purification; sequencing depth is not less than 200× genome coverage; the raw data is subjected to sequence alignment, duplicate filtering, and iterative correction using the HiC-Pro workflow; and a standardized chromatin interaction frequency matrix is generated after KR normalization.
[0015] By adopting the above technical solution, the use of tetrabasic restriction endonucleases for chromatin digestion results in a moderate cleavage frequency in the fish genome, and the resulting fragment size distribution is suitable for high-resolution Hi-C analysis. The sequencing depth of no less than 200× genome coverage ensures the statistical power required for chromatin loop detection. The HiC-Pro workflow combined with KR iterative correction normalization eliminates systematic errors such as enzyme digestion efficiency, GC content, and alignment rate. Therefore, a high-resolution, low-noise whole-genome chromatin interaction frequency matrix can be obtained, providing a reliable data foundation for subsequent saliency detection of chromatin loops.
[0016] Preferably, in S2, the screening criteria for differentially expressed genes are |log2FoldChange|>1 and the false discovery rate (FDR)<0.05; the immune pathway enrichment filtering involves performing immune-related GO function and KEGG pathway enrichment analysis on the differentially expressed genes, retaining only genes that are significantly enriched in the Toll-like receptor signaling pathway, NOD-like receptor signaling pathway, JAK-STAT signaling pathway, or immune effector processes, thus constituting the immune-related differentially expressed gene set;
[0017] The transcriptome sequencing used strand-specific RNA sequencing, and the gene expression data were TPM or FPKM normalized expression matrices.
[0018] By employing the above technical solutions, the dual threshold screening method (|log2FoldChange|>1 and FDR<0.05) balances the statistical significance of differential expression with the magnitude of biological effects. Secondary screening using immune-related GO and KEGG pathway enrichment excludes genes exhibiting non-specific changes such as stress responses and metabolic fluctuations. Strand-specific RNA sequencing preserves transcript direction information, improving the accuracy of gene quantification. Therefore, a high-quality set of differentially expressed genes with a clear immunological functional background can be obtained, providing an accurate core gene list for three-dimensional interaction identification and avoiding interference from non-immune genes in network construction.
[0019] Preferably, in step S3, chromatin compartment division is performed using principal component analysis (PCA) for A / B compartment division; topologically relevant domains are identified using TAD boundary identification, calculated using the directionality index method; significant chromatin loop detection uses the HiCCUPS or FitHiC algorithm, with a false detection rate (FDR) < 0.01 as the significance threshold; the criteria for determining promoter anchor points are: one anchor point of the chromatin loop falls within ±2kb upstream and downstream of the transcription start site of immune-related differentially expressed genes; the enhancer feature markers are DNase I hypersensitive sites, H3K27ac signal peaks, or H3K4me1 signal peaks; after intersection analysis of the enhancer feature markers, anchor points that meet the criteria are selected as candidate enhancer anchor points, with the criteria including: another anchor point overlaps with the enhancer feature marker and is not annotated as a promoter region; the intersection analysis is performed using the BEDTools tool.
[0020] By employing the above technical solutions, hierarchical three-dimensional structure analysis is achieved by using A / B compartment division at the large chromosome scale to provide a background for genomic transcriptional activity, identifying TAD boundaries at the medium scale to define the effective spatial range of functional interactions, and detecting significant chromatin loops at the fine scale using HiCCUPS or FitHiC algorithms with a strict threshold of FDR < 0.01. Promoter anchor point determination uses a ±2kb range covering the core promoter region, and enhancer candidate anchor point determination utilizes a combination of DNaseI hypersensitive sites, H3K27ac, and H3K4me1 epigenetic features for dual screening. Therefore, it is possible to accurately extract promoter-enhancer interaction pairs with transcriptional regulatory functions from massive three-dimensional interactions, significantly reducing the false positive rate and providing high-confidence spatial interaction relationships for network modeling.
[0021] Preferably, in step S4, the transcription factor binding motif enrichment analysis uses HOMER or MEME tools to screen for P-values <1×10⁻⁶. -5Significant enrichment of motifs and their corresponding transcription factors; the inference of the regulatory direction is performed by weighted gene co-expression network analysis or Bayesian network method, combined with gene expression dynamic data at different time points in S2 to perform directional inference and assign edge weights; the initial network is visualized and constructed using Cytoscape, with node size corresponding to gene expression level and edge thickness corresponding to three-dimensional interaction strength.
[0022] By employing the above technical solutions, the relationship between transcription factors enriched in the enhancer region and their target genes was identified through transcription factor binding motif enrichment analysis of enhancer sequences, revealing the molecular mechanism by which enhancers exert their regulatory functions. Weighted gene co-expression network analysis or Bayesian network methods combined with multi-timepoint expression dynamic data were used to infer the direction of regulation, distinguishing interaction relationships into activating and repressive types and assigning edge weights. Visualization was constructed using Cytoscape, with node and edge attributes corresponding to expression levels and interaction strengths. Therefore, the undirected interaction network can be upgraded into a directed weighted three-dimensional regulatory network containing regulatory direction, strength, and functional mechanisms, facilitating intuitive identification of core regulatory nodes and key regulatory pathways.
[0023] Preferably, in step S4, the directed weighted network output is in adjacency matrix format.
[0024] By adopting the above technical solution, since the network is output in the adjacency matrix format, the matrix elements contain complete information on the regulatory direction and weights. Therefore, it is convenient for downstream bioinformatics analysis and batch screening of disease-fighting molecular targets, and the reusability and computational compatibility of network data are improved.
[0025] Preferably, in step S5, the different immune phenotypes of fish fry or fish species include individuals with disease-resistant phenotypes and individuals with disease-susceptible phenotypes, obtained through challenge experiments or natural disease screening; the functional verification is performed by using a CRISPR / dCas9-KRAB interference system to transcribe and repress candidate enhancer regions, and then detecting changes in target gene expression by RT-qPCR; or by constructing a luciferase reporter vector carrying enhancer sequences and detecting enhancer activity in fish cell lines.
[0026] By employing the above technical solutions, and through repeated validation using individuals with both disease-resistant and susceptible phenotypes, the intersection analysis of edge sets preserves stable interaction relationships that occur under different immune response backgrounds, eliminating false positives caused by experimental batch effects or random noise. The CRISPR / dCas9-KRAB interference system and luciferase reporter vector experiments are used to directly verify the molecular function of key promoter-enhancer interactions, confirming the functional authenticity of the interaction relationships from two levels: enhancer transcriptional regulatory activity and target gene expression response. Therefore, through the complementary system of population stability validation and molecular function validation, false positive relationships can be effectively eliminated, missing associations discovered during the validation process can be supplemented, and ultimately a highly reliable three-dimensional gene immune regulatory network for fish can be obtained.
[0027] Preferably, in S1, the control sample is a fish fry or fingerling that is at the same developmental stage as the immune stimulation group, was raised in the same batch, was not subjected to immune stimulation treatment, and was only given an equal amount of solvent as the control.
[0028] By adopting the above technical solution, since the control group and the experimental group are at the same developmental stage and are cultivated in the same batch, there is only one variable: immune stimulation. Giving an equal amount of solvent eliminates the potential influence of the solvent itself. Therefore, it can effectively eliminate the interference of non-immune factors such as developmental time sequence changes, environmental factors and solvent stress on gene expression, and ensure that the differentially expressed genes screened out truly and specifically reflect the immune response process.
[0029] Preferably, in step S3, the obtained set of promoter and enhancer interaction relationships is further annotated with gene function and immune signaling pathway to generate a gene function annotation file; the gene function annotation is GO function annotation, and the immune signaling pathway annotation is KEGG pathway annotation; the annotation results are used as network node attribute information in step S4 for modeling.
[0030] By adopting the above technical solution, GO functional annotation and KEGG pathway annotation are performed simultaneously on genes in the interaction relationship set. The molecular function, biological process and signaling pathway information of the genes are assigned to the network as node attributes. Therefore, when modeling the S4 network, nodes can be classified and screened based on functional attributes. This makes it easier to identify the regulatory modules within specific immune pathways and the interaction nodes between different pathways, thereby enhancing the biological interpretability of the network.
[0031] This invention provides a method for constructing a three-dimensional gene-based immune regulatory network in fish. It has the following beneficial effects:
[0032] 1. This invention employs a technical solution that deeply integrates in situ Hi-C chromatin conformation capture with immunotranscriptomics. By hierarchically deconstructing the three-dimensional spatial organization of chromatin and simultaneously acquiring dynamic expression data of immune responses, it achieves the technical effect of revealing the remote regulatory relationship of immune genes from the perspective of three-dimensional spatial conformation. Compared with the existing technical solution that only constructs an immune regulatory network based on linear co-expression of the transcriptome, this invention solves the problems that it cannot capture the remote spatial interaction relationship between enhancers and promoters and misses the key information on the regulation of immune gene expression by distant regulatory elements.
[0033] 2. This invention employs a dual-screening approach combining precise promoter anchor location and multi-dimensional enhancer epigenetic markers. It performs intersection analysis between significant chromatin loop anchors and transcription start sites of immunodifferential genes and active enhancer markers, achieving the technical effect of accurately extracting promoter-enhancer interaction pairs with transcriptional regulatory functions from massive three-dimensional interactions. Compared to existing technologies that rely solely on interaction frequency or a single distance threshold to screen regulatory relationships, this invention solves the problems of high false positive rates and the inability to distinguish between structural and functional regulatory interactions.
[0034] 3. This invention employs a technical solution that combines enhancer sequence transcription factor binding motif enrichment analysis with multi-time point expression dynamic data for joint inference. By identifying transcription factor binding motifs enriched in enhancer regions and combining them with weighted gene co-expression network analysis to infer regulatory direction and assign edge weights, it achieves the technical effect of constructing a directed weighted three-dimensional immune regulatory network. Compared with the existing technologies that construct mostly undirected and unweighted static interaction graphs, this invention solves the problems that it cannot provide information on regulatory direction, is difficult to distinguish between activation and inhibition relationships, and cannot directly guide the priority screening of anti-disease molecular targets. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of the method steps of the present invention;
[0036] Figure 2 This is a schematic diagram of the three-dimensional genome interaction identification process of the present invention;
[0037] Figure 3 This is a schematic diagram illustrating the construction of the directed weighted three-dimensional immune regulatory network of the present invention;
[0038] Figure 4 This is a schematic diagram illustrating the network verification and iterative optimization of the present invention;
[0039] Figure 5 This is a schematic diagram of the final network structure of the present invention. Detailed Implementation
[0040] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0041] Please see the appendix Figure 1 - Appendix Figure 5 This invention provides a method for constructing a three-dimensional gene immune regulatory network in fish, comprising the following steps:
[0042] S1, Sample collection and multi-omics library preparation: Selected fish fry or fingerlings that have been immune-stimulated, collected immune tissues and performed in situ Hi-C cross-linking treatment and RNA extraction, respectively, to construct a three-dimensional gene interaction DNA library and a transcriptome sequencing library.
[0043] Furthermore, immune challenge was performed by pathogen immersion infection or injection of immune stimulants. Samples were taken at multiple time points after challenge, including 24 hours, 48 hours, and 72 hours. Immune tissues were selected from at least one of the head kidney, spleen, liver, gills, and intestinal mucosa. Fish fry included larvae and juveniles, while fingerlings were juveniles.
[0044] The control samples were fish fry or fingerlings that were at the same developmental stage as the immune stimulation group, were raised in the same batch, had not undergone immune stimulation treatment, and were given only an equal amount of solvent as a control.
[0045] Specifically, three time points—24 hours, 48 hours, and 72 hours—were set after immune challenge, based on the kinetic characteristics of the innate immune response in fish: 24 hours represents the early recognition and signal activation phase, 48 hours the phase of high expression of effector molecules, and 72 hours the phase of regulation and decline of the immune response. Sampling at multiple time points captures the dynamic evolution of the immune regulatory network. During tissue collection, the fish were anesthetized on ice and rapidly dissected. Immediately after ex vivo, the tissue was divided into two parts: one part was cross-linked and fixed in PBS solution containing 1% formaldehyde for subsequent Hi-C experiments; the other part was placed in RNAlater preservation solution or flash-frozen in liquid nitrogen for RNA extraction. In this protocol, the formaldehyde cross-linking fixation time was preferably 15-20 minutes. This condition ensures sufficient fixation of the chromatin spatial conformation while avoiding excessive cross-linking that could reduce the efficiency of subsequent enzyme digestion. The control group was set up to eliminate interference from non-immune factors such as developmental time variations and solvent stress on gene expression, ensuring that the differentially expressed genes screened accurately reflect the immune response process. The two types of samples obtained in this step connect the Hi-C library construction and transcriptome library construction processes in S2, respectively.
[0046] S2, Multi-omics Data Acquisition and Immune Gene Screening: High-throughput sequencing was performed on the library obtained in S1 to obtain a standardized chromatin interaction frequency matrix and gene expression data; using control samples without immune stimulation as a reference, differentially expressed genes were screened and enriched through immune pathways to obtain a set of immune-related differentially expressed genes.
[0047] Furthermore, the in situ Hi-C cross-linking process includes: formaldehyde cross-linking fixation, sucrose gradient centrifugation to separate cell nuclei, chromatin digestion using tetrabasic restriction endonucleases MboⅠ or DpnⅡ, biotin-labeled blunt ends, adjacent ligation, decross-linking, and DNA purification; sequencing depth is no less than 200× genome coverage, and the raw data is subjected to sequence alignment, duplicate filtering, and iterative correction using the HiC-Pro workflow, and a standardized chromatin interaction frequency matrix is generated after KR normalization;
[0048] The screening criteria for differentially expressed genes were |log2FoldChange|>1 and false discovery rate (FDR)<0.05; the immune pathway enrichment filter was performed on differentially expressed genes by immune-related GO function and KEGG pathway enrichment analysis, and only genes significantly enriched in the Toll-like receptor signaling pathway, NOD-like receptor signaling pathway, JAK-STAT signaling pathway or immune effector process were retained to form an immune-related differentially expressed gene set.
[0049] Transcriptome sequencing uses strand-specific RNA sequencing, and gene expression data are TPM or FPKM normalized expression matrices;
[0050] Specifically, the cross-linked nuclear samples obtained in S1 are cross-linked with formaldehyde. Spatially adjacent chromatin fragments within the cell are covalently fixed. Through steps such as in-situ restriction endonuclease digestion, biotin-labeled blunt ends, and proximity ligation, spatially close DNA ends are joined to form chimeric fragments. The abundance of these chimeric fragments reflects the contact frequency of corresponding genomic loci in the three-dimensional space of the cell nucleus. In this scheme, the preferred tetrabasic restriction endonuclease is MboI, whose recognition site is GATC. Its cleavage frequency in fish genomes is moderate, and the resulting fragment size distribution is suitable for Hi-C analysis. The sequencing depth is preferably 250× to 300× genome coverage to ensure the statistical power of chromatin loop detection. After data processing using the HiC-Pro workflow, normalization is performed using the KR iterative correction algorithm. The principle is to eliminate experimental biases, such as systematic errors in enzyme digestion efficiency, GC content, and alignment rate, through matrix balancing, making the chromatin interaction frequency matrix comparable. This normalized matrix is the direct input data for the S3 three-dimensional genome structure analysis.
[0051] Simultaneously, transcriptome analysis was performed on homologous tissue RNA samples obtained in S1. Strand-specific RNA sequencing can preserve the directional information of transcripts, improving the accuracy of gene quantification. Differential expression analysis was performed using DESeq2, which is based on a negative binomial distribution model to model gene count data and uses a contraction estimation method to handle the high dispersion of low-expression genes. The screening threshold |log2FoldChange|>1 and FDR<0.05 were used to comprehensively consider statistical significance and the magnitude of biological effect. Subsequently, the hypergeometric test in GO and KEGG pathway enrichment analysis was used to calculate the enrichment significance of differentially expressed genes in specific immune pathways, retaining only genes significantly enriched in immune-related pathways or functional terms. The significance of this filtering step is that not all genes whose expression changes after immune stimulation directly participate in immune regulation. Through secondary screening of pathway enrichment, genes with non-specific changes such as stress response and metabolic fluctuations are excluded, ensuring that the gene set entering subsequent steps has a clear immunological functional background, providing a high-quality core gene list for the three-dimensional interaction identification in S3.
[0052] S3, Three-dimensional genome interaction identification: Based on the chromatin interaction frequency matrix of S2, chromatin compartmentalization, topologically related domain identification, and significant chromatin loop detection are performed to extract loop anchor point coordinates; the intersection analysis of anchor points with transcription start sites and enhancer feature markers of immune-related differentially expressed genes is performed to obtain the set of interaction relationships between immune gene promoters and enhancers;
[0053] Furthermore, chromatin compartmentalization was performed using principal component analysis (PCA) for A / B compartmentalization, and topologically relevant domains were identified as TAD boundaries using the directionality index method. Significant chromatin loops were detected using the HiCCUPS or FitHiC algorithm, with a false discovery rate (FDR) < 0.01 as the significance threshold. The criteria for determining promoter anchors were: one anchor point of a chromatin loop must fall within ±2kb upstream and downstream of the transcription start site of immune-related differentially expressed genes. Enhancer markers were identified as DNase I hypersensitive sites, H3K27ac signal peaks, or H3K4me1 signal peaks. After intersection analysis of the enhancer markers, anchors meeting the criteria were selected as candidate enhancer anchors. The criteria included: another anchor point overlapping with the enhancer marker and not being annotated as a promoter region. Intersection analysis was performed using BEDTools.
[0054] The obtained set of promoter and enhancer interaction relationships was further annotated with gene function and immune signaling pathways to generate gene function annotation files; the gene function annotation was GO function annotation and the immune signaling pathway annotation was KEGG pathway annotation; the annotation results were used as network node attribute information in S4 for modeling.
[0055] Specifically, following the standardized chromatin interaction frequency matrix output by S2, the process begins with A / B compartmentalization at a large chromosome scale. This is achieved by performing principal component analysis on the standardized interaction matrix, with the positive and negative values of the first principal component corresponding to compartments A and B, respectively. Compartment A represents transcriptionally active, open chromatin regions, while compartment B represents transcriptionally silent, compressed chromatin regions. This step provides a genomic activity background for subsequent interaction analysis. Secondly, TAD boundaries are identified at a medium scale. The directional index method is used to calculate the upstream and downstream interaction deviations of each genomic region, and the sites where the interaction direction changes significantly are identified as TAD boundaries. The biological significance of TADs lies in the fact that the chromatin interaction frequency within the same TAD is much higher than between TADs. Functional interactions between genes and their regulatory elements mainly occur within TADs; therefore, TAD identification defines a meaningful spatial range for subsequent chromatin loop detection.
[0056] Within the aforementioned structural framework, the HiCCUPS algorithm is employed to detect significant chromatin loops within the TAD. This algorithm identifies locally enriched point-like signals in the interaction matrix as anchor pairs of loop structures and evaluates their significance through statistical tests. In this scheme, the FDR threshold is preferably <0.01, and the window size is set to 5kb or 10kb, which can be flexibly adjusted according to the genome size.
[0057] The key innovation lies in the functional annotation of the anchor points. The determination of promoter anchor points involves intersecting the coordinates of the loop anchor point with the TSS region of the immune-related differentially expressed genes obtained in S2. The TSS region is a range of ±2kb upstream and downstream of the transcription start site, which can cover typical core promoter regions. The determination of enhancer candidate anchor points comprehensively utilizes multiple epigenomic features: DNase I hypersensitive sites indicate open chromatin regions, H3K27ac signal peaks mark active enhancers and promoters, and H3K4me1 signal peaks specifically mark enhancer regions. If an anchor point simultaneously overlaps with these active enhancer features and is not located in a known promoter region, it is determined to be a candidate enhancer anchor point. The working principle of this dual-screening strategy is that not all distal sequences spatially contacting the promoter have regulatory activity; only regions that simultaneously possess the epigenomic features of active enhancers can exercise regulatory functions, thereby accurately extracting functional regulatory relationships from massive three-dimensional interactions.
[0058] The final set of promoter and enhancer interaction relationships contains two parts: spatial interaction pairs and functional annotations of their associated genes. The latter was completed through GO and KEGG analysis, providing complete input data with attribute labels for network modeling of S4.
[0059] S4, Three-dimensional immune regulation network modeling: Based on the promoter and enhancer interaction relationship set of S3, with immune genes as nodes and three-dimensional interactions as edges, an initial network is constructed by combining gene expression level and interaction strength; the regulatory direction is inferred by enhancing sequence transcription factor combined with motif enrichment analysis, forming a directed weighted fish three-dimensional gene immune regulation network.
[0060] Furthermore, transcription factor-motif enrichment analysis was performed using HOMER or MEME tools to screen for significantly enriched motifs and their corresponding transcription factors with a p-value <1×10-5; the direction of regulation was inferred using weighted gene co-expression network analysis or Bayesian network methods, combining gene expression dynamic data at different time points in S2 to infer directionality and assign edge weights; the initial network was visualized and constructed using Cytoscape, with node size corresponding to gene expression level and edge thickness corresponding to three-dimensional interaction strength.
[0061] The output of a directed weighted network is in the form of an adjacency matrix;
[0062] Specifically, building upon the promoter-enhancer interaction set and gene function annotation file from S3, the initial network construction uses immune-related genes as nodes and the three-dimensional interaction relationships between promoters and enhancers as edges. The initial weights of the edges are determined by the Hi-C interaction frequency, which is defined as the normalized contact count of the corresponding anchor pair in the normalized interaction matrix. This value reflects the tightness of chromatin spatial contact.
[0063] Enhancers exert their regulatory functions by recruiting transcription factors; therefore, the types and abundance of transcription factor binding motifs contained in the enhancer region determine its regulatory potential and direction. For all candidate enhancer anchor sequences identified by S3, known motif enrichment analysis was performed using HOMER software. The principle is to scan for overexpressed short sequence patterns in the enhancer sequences and match them with a database of known transcription factor binding motifs. The significance of enrichment was calculated using a cumulative binomial distribution test. In this scheme, the P-value threshold is preferably <1×10⁻⁻⁻⁶. 5 To strictly control false positives, significantly enriched transcription factors and their target gene relationships are incorporated into the network.
[0064] The inference of regulatory direction is based on dynamic gene expression data. The logic of weighted gene co-expression network analysis is adopted: the expression correlation between genes at each time point in S2 is calculated. If the expression of enhancer-associated genes and target genes shows a positive correlation and a temporal sequence, it is inferred to be activation-type regulation; if they show a negative correlation, it is inferred to be repression-type regulation. Simultaneously, the following regulatory direction score is constructed:
[0065] Enhancer E-associated gene is provided With target genes At a certain point in time The expression levels were respectively and Calculate the Pearson correlation coefficient. :
[0066] ;
[0067] in, For genes Average expression levels at three time points For genes Average expression levels at three time points The values 1, 2, and 3 correspond to the three sampling time points. When... and At that time, it was inferred to be activation-type regulation, with the direction being... The edge weight is set to ;when and At that time, it was inferred to be inhibitory regulation, with the direction being... The edge weight is set to ;when At that time, the interaction relationship is retained but marked as the direction of regulation to be determined.
[0068] The constructed network is stored as an adjacency matrix A, where the matrix elements are... This represents the weight of the regulatory edge from gene i to gene j, with positive values indicating activation, negative values indicating inhibition, and zero values indicating no direct regulatory relationship. The direct output of this directed weighted network serves as the object to be verified in the S5 verification step.
[0069] S5, Network Validation and Iterative Optimization: Repeated validation was performed using individuals with different immune phenotypes, and functional validation was performed on key interactions. False positives were eliminated and missing associations were added based on the validation results. The final network was obtained after iterative optimization.
[0070] Furthermore, fish fry or fingerlings with different immune phenotypes, including disease-resistant individuals and disease-susceptible individuals, were obtained through challenge experiments or natural disease screening; functional verification was performed by using a CRISPR / dCas9-KRAB interference system to transcribe candidate enhancer regions and detecting changes in target gene expression by RT-qPCR; or by constructing a luciferase reporter vector carrying enhancer sequences and detecting enhancer activity in fish cell lines.
[0071] Specifically, the directed weighted network that receives the output from S4 undergoes biological validation and iterative optimization in this step. Validation is divided into two levels: first, population stability validation, and second, molecular function validation.
[0072] In the population stability verification, individuals with resistant and susceptible phenotypes were selected separately, and the complete process from S1 to S4 was independently repeated to obtain the regulatory networks under the two phenotypes. An intersection analysis of the edge sets of the two networks was performed: the network was defined. For disease-resistant phenotype networks, For a susceptible phenotype network, if a certain regulatory edge simultaneously appears and In this strategy, if the nodes at both ends of an edge are identical and the regulatory direction is consistent, it is considered a stable interaction and retained. If an edge appears only in a single phenotype network, it is considered a candidate false positive and eliminated. The principle behind this strategy is that the true immune regulatory relationship should exist in the immune response process of different phenotypes, while false positive interactions are often caused by experimental batch effects or random noise, making them difficult to reproduce in different samples.
[0073] In molecular function verification, the edge with the highest weight in the network is selected, i.e. The interaction pairs between the top 5 to 10 promoters and enhancers with the highest expression values were experimentally verified. The verification principle of the CRISPR / dCas9-KRAB system is as follows: sgRNAs targeting candidate enhancer regions are designed to guide the dCas9-KRAB fusion protein to bind to the enhancer region. The KRAB domain recruits heterochromatin formation-related factors, causing histone H3K9me3 modification and deposition in the enhancer region, thereby specifically inhibiting the transcriptional regulatory activity of the enhancer. Changes in target gene mRNA expression were detected by RT-qPCR, and fold changes in expression were calculated using the Ct method.
[0074] ;
[0075] in, The results were measured in the dCas9-KRAB+sgRNA treatment group; The assay was performed in the dCas9-KRAB+ non-targeted sgRNA control group; The preferred internal reference gene is β-actin or EF1α, which are stably expressed genes in fish. If the expression of the target gene significantly decreases after enhancer interference, i.e. and This confirms that the enhancer has a positive regulatory function on the target gene.
[0076] The complementarity principle for validating luciferase reporter vectors is as follows: A candidate enhancer sequence is cloned upstream of a pGL4 luciferase reporter vector containing a minimal promoter to construct a luciferase reporter vector carrying the enhancer sequence. After transfection into fish cell lines, luciferase activity is detected. Using an empty vector as a control, the enhancer activity fold is calculated.
[0077] ;
[0078] Where Luc represents the luciferase activity value of firefly and Renilla represents the luciferase activity value of kidney lycopersicum, used as an internal control to correct transfection efficiency. and If so, the sequence is determined to have enhancer activity.
[0079] Based on the above two-layer validation results, false positive interaction edges that failed functional validation were removed, and new regulatory target gene relationships of enhancers discovered during functional validation were added to the network. After at least one round of iterative optimization, the network's edge set tended to stabilize, ultimately yielding a high-confidence three-dimensional gene immune regulation network for fish. This network can be directly used for molecular target screening of disease resistance traits in fish fry breeding and for systematic analysis of fish immune mechanisms.
[0080] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a three-dimensional gene-immune regulatory network in fish, characterized in that, Includes the following steps: S1, Sample collection and multi-omics library preparation: Selected fish fry or fingerlings that have been immune-stimulated, collected immune tissues and performed in situ Hi-C cross-linking treatment and RNA extraction, respectively, to construct a three-dimensional gene interaction DNA library and a transcriptome sequencing library. S2, Multi-omics Data Acquisition and Immune Gene Screening: High-throughput sequencing was performed on the library obtained in S1 to obtain a standardized chromatin interaction frequency matrix and gene expression data; using control samples without immune stimulation as a reference, differentially expressed genes were screened and enriched through immune pathways to obtain a set of immune-related differentially expressed genes. S3, Three-dimensional genome interaction identification: Based on the chromatin interaction frequency matrix of S2, chromatin compartmentalization, topologically related domain identification, and significant chromatin loop detection are performed to extract loop anchor point coordinates; the intersection analysis of anchor points with transcription start sites and enhancer feature markers of immune-related differentially expressed genes is performed to obtain the set of interaction relationships between immune gene promoters and enhancers; S4, Three-dimensional immune regulation network modeling: Based on the promoter and enhancer interaction relationship set of S3, with immune genes as nodes and three-dimensional interactions as edges, the initial network is constructed by combining gene expression level and interaction strength; By enhancing the binding of transcription factors to motif enrichment analysis, the direction of regulation is inferred, and a directed weighted three-dimensional gene immune regulation network for fish is formed. S5, Network Validation and Iterative Optimization: Repeated validation was performed using individuals with different immune phenotypes, and functional validation was performed on key interactions. False positives were eliminated and missing associations were added based on the validation results. The final network was obtained after iterative optimization.
2. The method for constructing a three-dimensional gene immune regulatory network in fish according to claim 1, characterized in that, In S1, immune stimulation is performed by pathogen immersion infection or injection of immune stimulants. Samples are taken at multiple time points after stimulation, including 24 hours, 48 hours, and 72 hours. The immune tissue is selected from at least one of the head kidney, spleen, liver, gills, and intestinal mucosa. The fish fry include larvae and juveniles, and the fish species are juveniles.
3. The method for constructing a three-dimensional gene immune regulatory network in fish according to claim 1, characterized in that, In S2, the in-situ Hi-C crosslinking process includes: formaldehyde crosslinking fixation, sucrose gradient centrifugation to separate cell nuclei, chromatin digestion using tetrabasic restriction endonucleases MboⅠ or DpnⅡ, biotin-labeled blunt ends, adjacent ligation, decrosslinking, and DNA purification; sequencing depth is not less than 200× genome coverage; the raw data is subjected to sequence alignment, duplicate filtering, and iterative correction using the HiC-Pro workflow; and a standardized chromatin interaction frequency matrix is generated after KR normalization.
4. The method for constructing a three-dimensional gene immune regulatory network in fish according to claim 1, characterized in that, In S2, the screening criteria for differentially expressed genes are |log2FoldChange|>1 and false discovery rate (FDR)<0.05; the immune pathway enrichment filtering involves performing immune-related GO function and KEGG pathway enrichment analysis on differentially expressed genes, retaining only genes significantly enriched in the Toll-like receptor signaling pathway, NOD-like receptor signaling pathway, JAK-STAT signaling pathway, or immune effector processes, thus constituting the immune-related differentially expressed gene set. The transcriptome sequencing used strand-specific RNA sequencing, and the gene expression data were TPM or FPKM normalized expression matrices.
5. The method for constructing a three-dimensional gene immune regulatory network in fish according to claim 1, characterized in that, In S3, chromatin compartment division is performed using principal component analysis (PCA) for A / B compartment division. Topologically relevant domains are identified using TAD boundary identification, calculated using the directionality index method. Significant chromatin loop detection employs the HiCCUPS or FitHiC algorithm, with a false detection rate (FDR) < 0.01 as the significance threshold. The promoter anchor point is determined when one chromatin loop anchor point falls within ±2 kb upstream and downstream of the transcription start site of an immune-related differentially expressed gene. Enhancer feature markers include DNase I hypersensitive sites, H3K27ac signal peaks, or H3K4me1 signal peaks. After the intersection analysis of the enhancer feature markers, anchor points that meet the conditions are selected as candidate enhancer anchor points. The determination conditions include: another anchor point overlaps with the enhancer feature marker and is not annotated as a starter region. The intersection analysis was performed using the BEDTools tool.
6. The method for constructing a three-dimensional gene immune regulatory network in fish according to claim 1, characterized in that, In the S4, transcription factor binding motif enrichment analysis uses HOMER or MEME tools to screen significant enrichment motifs and their corresponding transcription factors with P value < 1x10 -5 ; the inference of the regulation direction adopts weighted gene co-expression network analysis or Bayesian network method to infer the direction and assign edge weights in combination with the dynamic gene expression data at different time points in S2; and the initial network is visualized and constructed using Cytoscape, with the node size corresponding to the gene expression amount and the edge thickness corresponding to the three-dimensional interaction intensity.
7. The method for constructing a three-dimensional gene immune regulatory network in fish according to claim 1, characterized in that, In S4, the directed weighted network output is in the adjacency matrix format.
8. The method for constructing a three-dimensional gene immune regulatory network in fish according to claim 1, characterized in that, In S5, the different immune phenotypes of fish fry or fish species include disease-resistant individuals and disease-susceptible individuals, obtained through challenge experiments or natural disease screening; the functional verification is to use the CRISPR / dCas9-KRAB interference system to transcribe and repress candidate enhancer regions, and detect changes in target gene expression by RT-qPCR; or to construct a luciferase reporter vector carrying enhancer sequences and detect enhancer activity in fish cell lines.
9. The method for constructing a three-dimensional gene immune regulatory network in fish according to claim 1, characterized in that, In S1, the control sample is a fish fry or fingerling that is at the same developmental stage as the immune stimulation group, was raised in the same batch, was not subjected to immune stimulation treatment, and was only given an equal amount of solvent as the control.
10. The method for constructing a three-dimensional gene immune regulatory network in fish according to claim 1, characterized in that, In step S3, gene function annotation and immune signaling pathway annotation are performed on the obtained set of promoter and enhancer interaction relationships to generate a gene function annotation file; the gene function annotation is GO function annotation, and the immune signaling pathway annotation is KEGG pathway annotation; the annotation results are used as network node attribute information in step S4 for modeling.