Method, device and equipment for screening high-expression integration sites of saccharomyces cerevisiae

By calculating the chromatin interaction direction preference and local interaction intensity in the Saccharomyces cerevisiae genome and screening for highly expressed integration sites based on gene expression data, the problem of expression differences caused by the genomic position effect of Saccharomyces cerevisiae was solved, and efficient and stable exogenous gene expression was achieved.

CN121747692APending Publication Date: 2026-03-27ACADEMY OF MILITARY MEDICAL SCIENCES
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-02
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

The lack of systematic theoretical guidance in existing technologies leads to improper selection of the integration site of exogenous genes in the Saccharomyces cerevisiae genome, resulting in low construction efficiency and unstable product yield. Existing three-dimensional genome analysis algorithms cannot accurately capture the fine chromatin structure features of Saccharomyces cerevisiae and cannot effectively predict transcriptional potential.

Method used

By acquiring high-throughput chromatin conformation capture data and gene expression data from the Saccharomyces cerevisiae genome, we calculated chromatin interaction direction preference and local interaction intensity, identified chromatin interaction domains and their boundaries, and screened out highly expressed integration sites by combining gene expression data.

Benefits of technology

Precisely capturing functional units with high transcriptional potential in the genome can improve the expression level and stability of exogenous genes in Saccharomyces cerevisiae, providing a theoretical basis for efficient microbial cell factories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747692A_ABST
    Figure CN121747692A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device and equipment for screening high-expression integration sites of saccharomyces cerevisiae, and relates to the technical field of bioinformatics. The screening method comprises the following steps: acquiring high-throughput chromatin conformation capture data and gene expression data; calculating chromatin interaction direction preference and determining a chromatin interaction structural domain; calculating local interaction intensity and identifying a frequent interaction area; and screening to obtain the high-expression integration site. According to the method, quantitative association of a fine three-dimensional structure and transcriptional activity of the saccharomyces cerevisiae is realized by analyzing interaction direction preference and local interaction strength; by jointly screening the structural domain boundary and the frequent interaction region, an integration site with high expression potential is accurately positioned, the position effect is effectively overcome, and the expression efficiency and stability of an exogenous gene are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and particularly relates to a screening method for high-expression integration sites of Saccharomyces cerevisiae, and a device and equipment thereof. BACKGROUND

[0002] Saccharomyces cerevisiae is an important chassis cell in the field of synthetic biology and biological manufacturing, and is widely used in the production of various high-value chemicals. In the process of constructing an efficient microbial cell factory, the integration of exogenous genes into the host genome is a key step to realize the stable synthesis of target products. However, the transcriptional activity of different regions of the genome is significantly different, and the expression level of the exogenous gene is often severely affected by the position effect of the integration site. At present, the selection of the integration position of the exogenous gene often lacks systematic theoretical guidance, resulting in problems such as low construction efficiency and unstable product yield in the process of constructing a cell factory.

[0003] With the development of bioinformatics technology, it has been found that gene expression is closely related to the three-dimensional spatial organization of chromatin. In the research of higher eukaryotic (such as mammal) cells, a relatively mature framework for analyzing chromosome organization has been established, for example, by analyzing high-throughput chromatin conformation capture data to identify topologically associated domains and other structures to explore the spatial characteristics of the genome. These analysis methods based on three-dimensional genomics provide a new perspective for understanding the mechanism of gene transcriptional regulation and also provide a potential research path for solving the problem of genomic position effect.

[0004] However, the genomic structural characteristics of Saccharomyces cerevisiae are significantly different from those of mammals. Its chromosomes are relatively short, and the chromatin interaction mode is more delicate and complex. Most of the existing three-dimensional genome analysis algorithms are designed for large-scale characteristics of mammalian genomes, and when they are directly applied to data analysis of Saccharomyces cerevisiae, they often face the challenge of algorithm inapplicability. The existing technology cannot accurately capture the unique local structural characteristics of Saccharomyces cerevisiae, and the identified domains often have fuzzy boundaries or improper scales, which cannot truly reflect the delicate organization mode of its chromatin.

[0005] In addition, the existing analysis methods have not established a close and quantitative correlation model between the three-dimensional chromatin structural characteristics of Saccharomyces cerevisiae and the transcriptional activity of genes, and cannot effectively predict the transcriptional potential of a specific chromatin region. In summary, there is a lack of an effective solution in the existing technology that can adapt to the genomic characteristics of Saccharomyces cerevisiae and systematically screen gene integration sites with high confidence and high expression potential by deeply mining three-dimensional chromatin organization information. SUMMARY

[0006] In a first aspect, the present application provides a screening method for high-expression integration sites of Saccharomyces cerevisiae, comprising: obtaining high-throughput chromatin conformation capture data and gene expression data of a Saccharomyces cerevisiae genome; dividing the genome into a plurality of consecutive bins, calculating a chromatin interaction direction preference of each of the bins based on the high-throughput chromatin conformation capture data, and determining chromatin interaction domains on the genome according to the chromatin interaction direction preference; calculating a local interaction strength of each of the bins based on the high-throughput chromatin conformation capture data, and identifying frequent interaction regions with high local interaction strength and positive correlation with gene expression levels shown by the gene expression data in combination with the gene expression data; screening genome sites located at boundaries of the chromatin interaction domains and overlapping or adjacent to the frequent interaction regions as the high-expression integration sites.

[0007] In an optional embodiment, the calculation expression of the chromatin interaction direction preference of each of the bins is: ; wherein CDS(bin i ) represents a chromatin interaction direction preference score of the i-th bin; represents an average value of interaction scores between the i-th bin and the upstream k bins, and the calculation method is: represents an interaction score between the i-th bin and the j-th upstream bin; represents an average value of interaction scores between the i-th bin and the downstream k bins, and the calculation method is: represents an interaction score between the i-th bin and the m-th downstream bin; and k represents the number of bins used for calculating the interaction range according to a preset resolution.

[0008] In an optional embodiment, the determination of the chromatin interaction domains on the genome according to the chromatin interaction direction preference comprises: identifying positions of adjacent two bins on the genome with opposite signs of direction preference scores; defining the positions with opposite signs as boundaries of the chromatin interaction domains.

[0009] In an optional embodiment, before the determination of the chromatin interaction domains on the genome according to the chromatin interaction direction preference, further comprising: setting an interference threshold; excluding bins with absolute values of the calculated direction preference scores lower than the interference threshold from being used as a basis for identifying boundaries of the chromatin interaction domains.

[0010] In an optional embodiment, before determining the chromatin interaction domains on the genome according to the chromatin interaction direction preference, further comprising: obtaining high-throughput chromatin conformation capture data of at least two biological replicates of the same Saccharomyces cerevisiae strain; setting different interval length parameters to identify corresponding chromatin interaction domain boundaries respectively; comparing the overlap degrees of the chromatin interaction domain boundaries between different biological replicate samples, and selecting the interval length parameter with the highest overlap degree as the final basis for genome division.

[0011] In an optional embodiment, the calculation of the local interaction intensity of each interval based on the high-throughput chromatin conformation capture data comprises: for each interval, calculating the sum of interaction frequencies with all other intervals within a preset genomic distance range to obtain the local interaction intensity of the interval.

[0012] In an optional embodiment, the identification of the frequently interacting region with high local interaction intensity and positive correlation with the gene expression level shown by the gene expression data comprises: using a sliding window to slide along the genome, calculating the Pearson correlation coefficient between the local interaction intensity and the gene expression level within each window; selecting a region with a Pearson correlation coefficient greater than a preset correlation threshold and a local interaction intensity higher than a preset intensity threshold as the frequently interacting region.

[0013] In an optional embodiment, after determining the chromatin interaction domains on the genome, further comprising: calculating the multi-dimensional structure features of each chromatin interaction domain; performing cluster analysis on all the chromatin interaction domains based on the multi-dimensional structure features to divide the chromatin interaction domains into different categories; combining the gene expression data to evaluate the average transcription level of the chromatin interaction domains in each category; determining the category with the highest average transcription level as the target category; and the screening of the genomic sites located at the boundary of the chromatin interaction domain and overlapping or adjacent to the frequently interacting region comprises: screening the genomic sites located at the boundary of the chromatin interaction domain of the target category and overlapping or adjacent to the frequently interacting region.

[0014] In an optional embodiment, the multi-dimensional structure features comprise at least one of the following: Average distance between domains: the interaction density within a domain calculated by converting the contact matrix into a distance matrix; Domain insulation score: the absolute value of the chromatin interaction direction preference at the domain boundary; Average contact score of structural domain: the total interaction strength within the structural domain; Domain contact variation coefficient: the degree of variation in the interaction strength within a domain.

[0015] In a second aspect, the present invention provides a screening device for highly expressed integration sites in Saccharomyces cerevisiae, comprising: The acquisition module is used to acquire high-throughput chromatin conformation capture data and gene expression data of the Saccharomyces cerevisiae genome; The computation module is used to divide the genome into multiple continuous regions, calculate the chromatin interaction direction preference of each region based on the high-throughput chromatin conformation capture data, and determine the chromatin interaction domains on the genome according to the chromatin interaction direction preference. The identification module is used to calculate the local interaction intensity of each interval based on the high-throughput chromatin conformation capture data, and in combination with the gene expression data, identify frequently interacting regions with high local interaction intensity and positive correlation with the gene expression level shown by the gene expression data. The screening module is used to screen out genomic sites located at the boundary of the chromatin interaction domain and overlapping with or adjacent to the frequently interacting regions, as the highly expressed integration sites.

[0016] Thirdly, the present invention provides a computer device comprising a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the screening method for highly expressed integration sites of *Saccharomyces cerevisiae* as described in any of the foregoing embodiments.

[0017] Fourthly, the present invention provides a computer storage medium storing a computer program, which, when executed on a processor, implements a method for screening highly expressed integration sites of *Saccharomyces cerevisiae* according to any one of the foregoing embodiments.

[0018] The embodiments of this application have the following beneficial effects: This invention, through in-depth mining of high-throughput chromatin conformation capture data from the entire genome of *Saccharomyces cerevisiae*, introduces two key dimensions: chromatin interaction direction preference and local interaction intensity. This effectively overcomes the inapplicability of existing three-dimensional genome analysis algorithms when processing the fine structure of short chromosomes in *Saccharomyces cerevisiae*. By accurately identifying chromatin interaction domains and their boundaries, and further combining this with gene expression data to locate frequently interacting regions, this method can precisely capture functional units with specific regulatory properties in the genome at the three-dimensional spatial structure level.

[0019] More importantly, this method innovatively combines the boundary features of chromatin interaction domains with the high-activity features of frequently interacting regions, establishing a systematic set of screening criteria. This dual screening strategy not only quantifies the intrinsic relationship between chromatin spatial structure and transcriptional activity but also significantly improves the confidence level of the screening results. The sites ultimately identified, located at the domain boundaries and overlapping with or adjacent to frequently interacting regions, possess naturally high transcriptional potential and stable expression environments, effectively avoiding expression differences caused by genomic position effects. Integration sites screened using this method can significantly enhance the expression level and stability of exogenous genes in *Saccharomyces cerevisiae*, providing a reliable theoretical basis and data support for constructing efficient microbial cell factories. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and therefore should not be considered as a limitation on the scope of protection of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the hardware operating environment involved in an embodiment of the method for screening highly expressed integration sites in Saccharomyces cerevisiae according to the present invention; Figure 2 This is a flowchart illustrating Example 1 of the method for screening highly expressed integration sites in Saccharomyces cerevisiae according to the present invention. Figure 3 This is a schematic diagram illustrating the identification principle of chromatin interaction domains in Example 1 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention. Figure 4 This is a flowchart illustrating the refinement of step S200 (S210~S220) in Example 2 of the screening method for high-expression integration sites in Saccharomyces cerevisiae of the present invention. Figure 5 This is a flowchart illustrating the supplementary steps (S200-1~S200-2) following step S200 in Example 2 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention. Figure 6 This is a flowchart illustrating the supplementary steps (S200-3~S200-5) following step S200 in Example 3 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention. Figure 7 This is a schematic diagram illustrating the influence of the length (k value) of the upstream and downstream interaction interval on the conservation of the CID boundary in Example 3 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention. Figure 8 This is a schematic diagram illustrating the trend of the influence of the interference threshold (p-value) on the conservation of CID in Example 3 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention. Figure 9 This is a detailed flowchart of step S300 in Example 4 of the method for screening integration sites highly expressed in Saccharomyces cerevisiae according to the present invention. Figure 10 This is a schematic diagram illustrating the detailed process of identifying frequently interacting regions in Example 4 of the method for screening integration sites highly expressed in Saccharomyces cerevisiae according to the present invention. Figure 11 This is a scatter plot showing the correlation of CDS values ​​among different biological replicates in Example 6 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention. Figure 12 Pearson correlation histogram of the results of downsampling data at different ratios and the original data in Example 6 of the screening method for high expression integration sites in Saccharomyces cerevisiae of the present invention; Figure 13 This is a statistical diagram showing the distribution and enrichment of highly expressed genes near the CID boundary in Example 7 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention. Figure 14 A violin diagram comparing gene expression levels in the CID boundary region and non-boundary region in Example 7 of the screening method for high-expression integration sites in Saccharomyces cerevisiae of the present invention; Figure 15 This is a graph showing the correlation between the average transcriptional activity of CID and its chromatin accessibility and length in Example 7 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention. Figure 16 The consensus cumulative distribution function diagram used to determine the optimal cluster number k value in Example 8 of the screening method for high-expression integration sites in Saccharomyces cerevisiae of the present invention; Figure 17 Box plot comparing the average transcriptional activity of each category of CID obtained by clustering in Example 8 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention; Figure 18 Box plot comparing chromatin accessibility of each category of CID obtained by clustering in Example 8 of the screening method for high-expression integration sites in Saccharomyces cerevisiae of the present invention; Figure 19 Box plot comparing the insulation scores of each category of CID obtained by clustering in Example 8 of the screening method for high-expression integration sites in Saccharomyces cerevisiae of the present invention; Figure 20 Box plot comparing the internal average distances of each category of CID obtained by clustering in Example 8 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention; Figure 21 Box plot comparing the contact coefficient of variation of each category of CID obtained by clustering in Example 8 of the screening method for high expression integration sites in Saccharomyces cerevisiae of the present invention; Figure 22 Box plot comparing contact scores of CIDs of each cluster obtained in Example 8 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention; Figure 23 This is a genome-wide frequent interaction score map in Example 9 of the screening method for highly expressed integration sites in *Saccharomyces cerevisiae* according to the present invention; Figure 24 This is a graph showing the correlation between the intensity of local interactions across the entire genome and gene expression levels in Example 9 of the screening method for highly expressed integration sites in Saccharomyces cerevisiae of the present invention. Figure 25 Box plot comparing gene transcription levels between high-FIRE and low-FIRE regions in Example 9 of the screening method for high-expression integration sites in Saccharomyces cerevisiae of the present invention; Figure 26 This is a schematic diagram of the strategy for screening highly expressed integration sites in Saccharomyces cerevisiae using a combination of CID boundary and FIRE region in Example 9 of the screening method for high-expression integration sites in Saccharomyces cerevisiae of the present invention. Figure 27 This is a schematic diagram of the module connections of the screening device for high expression integration sites in Saccharomyces cerevisiae according to the present invention. Detailed Implementation

[0022] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0023] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0024] In the following, the terms “comprising,” “having,” and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as excluding, firstly, the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations thereof.

[0025] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.

[0026] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.

[0027] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0028] like Figure 1 The diagram shown is a structural schematic of the hardware operating environment of the terminal involved in an embodiment of the present invention.

[0029] The screening system (device) for high-expression integration sites in *Saccharomyces cerevisiae* according to embodiments of the present invention can be a PC, or a mobile terminal device such as a smartphone, tablet, or portable computer. This screening system for high-expression integration sites in *Saccharomyces cerevisiae* may include: a processor 1001, such as a CPU; a network interface 1004; a user interface 1003; a memory 1005; and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen, an input unit such as a keyboard, or a remote control; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory or a stable memory, such as a disk storage device. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001. Optionally, the screening system for high-expression integration sites in *Saccharomyces cerevisiae* may also include RF (Radio Frequency) circuitry, audio circuitry, a Wi-Fi module, etc. In addition, the screening system for highly expressed integration sites in Saccharomyces cerevisiae can also be equipped with other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, which will not be described in detail here.

[0030] Those skilled in the art will understand that Figure 1 The system (device) shown is not intended to limit it and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Figure 1 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a data interface control program, a network connection program, and a screening program for highly expressed integration sites of Saccharomyces cerevisiae.

[0031] In summary, the method provided by this invention overcomes the problem that existing algorithms are not applicable to the fine structure of Saccharomyces cerevisiae by calculating chromatin interaction direction preference and local interaction intensity, and accurately defines chromatin interaction domains and frequently interacting regions. At the same time, by establishing a quantitative correlation between spatial structure and transcriptional activity, and by using the superposition characteristics of domain boundaries and high-interaction regions for screening, it effectively avoids the genomic position effect, and provides a precise and efficient screening strategy for obtaining highly expressed and highly stable exogenous gene integration sites, solving the problems of lack of theoretical guidance and low expression efficiency in cell factory construction.

[0032] Example 1 Reference Figure 2 This embodiment provides a method for screening highly expressed integration sites in Saccharomyces cerevisiae, including: Step S100: Obtain high-throughput chromatin conformation capture data and gene expression data of the Saccharomyces cerevisiae genome.

[0033] This step is the data preparation phase of the entire method and aims to collect two key types of raw bioinformatics data.

[0034] The aforementioned high-throughput chromatin conformation capture data (Hi-C data) can reveal the folding and interaction of the genome in three-dimensional space, and can characterize the actual physical proximity of DNA fragments that are at a certain distance within the cell nucleus.

[0035] The gene expression data (RNA-seq data) mentioned above can quantify the activity level (i.e., transcription level) of each gene on the genome. It can indicate which genes are being heavily "read" and "transcribed" and which are in a silent state.

[0036] The fundamental premise of this method is that gene expression levels are not only related to their own sequence but also closely related to their surrounding three-dimensional chromatin structure. Therefore, data from both the "structure" and "function" dimensions must be obtained simultaneously to establish the correlation between the two.

[0037] By explicitly requesting the acquisition of these two types of data, a complete and necessary data foundation was provided for subsequent correlation analysis. This ensured that the analysis was not based on guesswork, but rather on solid data support.

[0038] Specifically, this can be achieved through at least the following two methods: (1) Download published, quality-compliant Saccharomyces cerevisiae Hi-C and RNA-seq datasets from public databases (such as NCBI GEO, ArrayExpress, etc.). (2) Conduct Hi-C and RNA-seq experiments on the target Saccharomyces cerevisiae strain to generate first-hand data.

[0039] Step S200: Divide the genome into multiple continuous regions, calculate the chromatin interaction direction preference of each region based on the high-throughput chromatin conformation capture data, and determine the chromatin interaction domains on the genome according to the chromatin interaction direction preference.

[0040] This step aims to use Hi-C data to divide the linear yeast genome into structurally independent but potentially functionally related "domains" and identify the "boundaries" of these domains.

[0041] Specifically, the entire genome can first be virtually divided into many small segments of equal length, called "bins." These are the basic units for computation. Then, the directional preference is calculated, which is the core of this step. For each bin, Hi-C data is used to calculate the total interaction strength between it and all bins "upstream" (along the direction of decreasing chromosome number), and the total interaction strength between it and all bins "downstream" (along the direction of increasing chromosome number). "Chromatin interaction directional preference" refers to the imbalance between upstream and downstream interactions. Finally, the domain is determined; when the preference of a bin suddenly changes from "upstream preference" to "downstream preference," this transition point is defined as the boundary of the domain. The region between two adjacent boundaries constitutes a chromatin interaction domain (CID).

[0042] For example, suppose the genome is divided into 1kb intervals. For interval 100, the total number of interactions with intervals 1-99 is confirmed to be 500, while the total number of interactions with intervals 101-200 is 50. This indicates a strong "upstream bias". For interval 101, the total number of interactions with intervals 1-100 is confirmed to be 60, while the total number of interactions with intervals 102-201 is 600, showing a strong "downstream bias". Therefore, there exists a CID boundary between intervals 100 and 101.

[0043] This step allows us to obtain the locations of all CIDs and their precise boundary sites across the entire genome. This method does not rely on models from other species; it starts directly from the data and objectively identifies the intrinsic, true structural organization patterns within the *Saccharomyces cerevisiae* genome. It endows the one-dimensional genome sequence with three-dimensional structural annotation.

[0044] For example, an algorithm can be designed to calculate a "Directional Index" (DI) along each interval of the chromosome, with the formula defined as: DI = (Downstream Interaction - Upstream Interaction) / (Downstream Interaction + Upstream Interaction). Then, the "zero-crossing points" where the DI value changes from positive to negative or vice versa are identified; these points are the CID boundaries.

[0045] refer to Figure 3 This demonstrates the identification principle of chromatin interaction domains in the embodiments of this application, which determines the orientation preference score (CDS) by calculating the logarithmic ratio of upstream and downstream interaction strengths.

[0046] Step S300: Calculate the local interaction intensity of each interval based on the high-throughput chromatin conformation capture data, and combine it with the gene expression data to identify frequently interacting regions with high local interaction intensity and positive correlation with the gene expression level shown in the gene expression data.

[0047] This step aims to utilize Hi-C data, but combine it with gene expression data, to identify "hotspot regions" on the genome that are both "social centers" (frequent interactions) and "functional centers" (active expression) from another dimension.

[0048] Specifically, local interaction strength can be calculated, where for each interval, the total number of interactions between it and other intervals within a certain range (e.g., 50 intervals before and after it) is calculated. This value reflects the "local social activity" of the interval. Then, association analysis is performed by combining gene expression data, that is, comparing the "local interaction strength" calculated in the previous step with the "gene expression level" of the interval and its vicinity; FIREs are identified, screening out regions that simultaneously meet two conditions: their local interaction strength is at a high level across the whole genome; and their local interaction strength shows a significant positive correlation with gene expression level (i.e., the stronger the interaction, the more active the gene expression). These regions are defined as Frequently Interacting Regions (FIREs).

[0049] This step yields a list of short, non-uniformly distributed "hotspot" regions on the genome. This step integrates structural information (interaction strength) and functional information (gene expression). The identified FIREs are not only structural "anchors" but also functional "engines," and their direct association with high expression makes them highly promising candidate regions.

[0050] For example, a sliding window method can be used. The entire genome can be scanned using a 20kb window and a 5kb step size. For each window, the total number of interactions (local interaction strength) and the average gene expression level within it are calculated. Then, statistical methods (such as calculating Spearman or Pearson correlation coefficients) are used to identify windows where both interaction strength and expression level are significantly higher than the genomic background, and both are strongly positively correlated; these are defined as FIREs.

[0051] Step S400: Select genomic sites located at the boundary of the chromatin interaction domain and overlapping or adjacent to the frequently interacting regions as the highly expressed integration sites.

[0052] This step is the final decision-making step. It integrates the results from the previous two steps and uses a "double standard" to identify the final, highest-confidence candidate site.

[0053] As mentioned above, CID boundaries (boundaries of chromatin interaction domains) represent structural "transition regions." These regions are typically open, easily accessible, and are the boundaries between different regulatory environments, thus possessing unique regulatory potential.

[0054] The aforementioned FIRE represents a functionally validated "highly active center".

[0055] The principle behind this step is that if a site is located at a critical node of structural transition (CID boundary) and has been confirmed as a functionally active center (FIRE), then that site is the most ideal target. Integrating a foreign gene at this site maximizes the chances of it benefiting from the inherent structural and regulatory environment of that region, which is conducive to efficient expression.

[0056] Specifically, this can be a simple set intersection or proximity search operation; take out the "CID boundary site" list obtained in the second step; take out the "FIRE region" list obtained in the third step; find all sites that belong to the CID boundary and fall inside (overlapping) or are adjacent to the FIRE region (nearby).

[0057] This step yields a highly purified, limited but extremely high-quality candidate list of "highly expressed integration sites," significantly improving the accuracy and efficiency of the screening process. Compared to screening using CID boundaries or FIRE alone, this dual standard filters out a large number of false positives, resulting in a higher "hit rate" for the final candidate sites and saving considerable time and cost for subsequent experimental validation.

[0058] For example, any tool that processes genomic regions (such as bedtools) can be used to perform this. Take the coordinate files of the CID boundaries and the coordinate files of the FIRE regions as input, and run the intersect (overlap) or closest (nearest) command to obtain the final result.

[0059] Furthermore, in steps S300 and S400, in addition to Hi-C and RNA-seq, chromatin accessibility data (ATAC-seq) can be further integrated. The screening criteria can be upgraded to sites "located at the CID boundary, overlapping with FIRE, and in a highly open chromatin region," further improving the success rate.

[0060] Furthermore, CIDs are classified: After the second step, a sub-step can be added. Based on characteristics such as the average gene expression level and interaction density within a CID, a clustering algorithm is used to classify CIDs into different categories such as "highly active," "moderately active," and "silent." Then, in step S400, sites located on the boundaries of "highly active CIDs" can be preferentially selected.

[0061] In addition, when calculating directional bias and local interaction strength, more complex normalization algorithms or a combination of multiple normalization algorithms can be used to eliminate the influence of technical biases such as sequencing depth and GC content, making the results more robust.

[0062] Example 2 This embodiment provides a method for screening highly expressed integration sites in Saccharomyces cerevisiae. In step S200, the expression for calculating the chromatin interaction direction preference of each interval (Formula 1) is as follows: ; Among them, CDS (bin) i ) represents the chromatin interaction direction preference score of the i-th interval; The average of the scores between the i-th interval and the k upstream intervals is calculated using Formula 2: C (i,j) This represents the interaction score between the i-th interval and the j-th upstream interval; The average of the interaction scores between the i-th interval and the k downstream intervals is calculated using Formula 3: C (i,m) The value represents the interaction score between the i-th interval and the m-th downstream interval; k represents the number of intervals used to calculate the interaction range, determined according to a preset resolution.

[0063] Formula 1 calculates the normalized difference. The numerator is B(bin). i )-A(bin i The denominator A(bin) represents the difference between the "downstream interaction average" and the "upstream interaction average". i )+B(bin i ) represents the total average interaction strength, used for normalization so that the CDS score is limited to a fixed range (from -1 to +1).

[0064] Based on Formula 1, the following logic can be deduced: (1) If the CDS score is greater than 0, it means that B(bi) ni )>A(bin i This means that the interaction between the i-th interval and the downstream is stronger.

[0065] (2) If the CDS score is less than 0, it means that A(bin) i B(bin) i That is, the interaction between the i-th interval and the upstream is stronger.

[0066] (3) If the CDS score is close to 0, it means that its interaction strength with upstream and downstream is roughly equal, and there is no obvious preference.

[0067] In Formula 1, k defines the size of the "window" considered when calculating the average interaction strength. It can be a positive integer, representing the number of intervals extending upstream or downstream. The value of k depends on the "preset resolution" of the data (i.e., the size of each interval) to ensure that the window of analysis covers a biologically meaningful physical distance.

[0068] The method described above transforms the qualitative concept of "directional preference" into a calculable and comparable quantitative score. By normalizing the denominator, the influence of differences in overall interaction strength across different regions is eliminated, allowing direct comparison of CDS scores at different genomic locations. Given the same Hi-C data and k-value, anyone can obtain identical results using this formula, ensuring the scientific validity and reproducibility of the method.

[0069] For example, if the resolution is 10kb (10kb per interval), and k is set to 20, then based on this method, the CDS score for interval 100 needs to be calculated. Calculate A(bin100): Add the interaction scores between interval 100 and the upstream intervals 80 to 99 (20 in total), then divide by 20. Calculate B(bin100): Add the interaction scores between interval 100 and the downstream intervals 101 to 120 (20 in total), then divide by 20. Assume A(bin100) = 50 and B(bin100) = 150. Then CDS(bin100) = (150 - 50) / (150 + 50) = 100 / 200 = +0.5. Based on this result, interval 100 has a strong preference for interacting with downstream intervals.

[0070] Reference Figure 4 In some embodiments, step S200, determining chromatin interaction domains on the genome based on the chromatin interaction direction preference, includes: Step S210: Identify positions on the genome where the directional preference score signs of two adjacent regions are opposite.

[0071] Step S220: The position opposite to the symbol is defined as the boundary of the chromatin interaction domain.

[0072] This method further specifies how to use the CDS scores calculated in the aforementioned embodiments to specifically and objectively delineate the boundaries of the “chromatin interaction domain (CID)”.

[0073] This method provides a boundary identification rule: the boundary is where the CDS score sign flips. The mechanism is that intervals within a CID generally have the same interaction preference (e.g., they tend to interact with members within the same domain). Therefore, within a CID, the CDS score sign will be relatively consistent. When transitioning from one CID to the next, the interaction "center of gravity" reverses, causing the CDS score sign to flip. This "flip point" is the boundary between the two domains.

[0074] Specifically, the CDS score sequence calculated in the aforementioned method and arranged along the genome can be obtained, for example, [CDS(bin1), CDS(bin2), ..., CDS(binN)]. Then, this sequence is traversed from the beginning, and each pair of adjacent scores CDS(bini) and CDS(bini+1) is checked. If the sign of CDS(bini) is opposite to that of CDS(bini+1) (one is positive and the other is negative), a CID boundary is marked between the i-th interval and the i+1-th interval, thereby obtaining a list of CID boundary sites that are accurate to specific intervals across the entire genome.

[0075] This method provides a boundary definition standard that is entirely based on data computation and requires no human judgment; the algorithm is very simple, has low computational complexity, and can quickly process whole genome data.

[0076] For example, suppose the CDS score sequence of a genome is: [..., +0.6, +0.4, -0.2, -0.5,...]. Between +0.4 and -0.2, the sign of the CDS score changes from positive to negative. Therefore, according to the rules of this method, a CID boundary is defined at this position (i.e., between the intervals corresponding to the second and third scores).

[0077] refer to Figure 5 In some embodiments, before determining the chromatin interaction domains on the genome based on the chromatin interaction direction preference in step S200, the method further includes: Step S200-1: Set the interference threshold; Step S200-2: Exclude the intervals where the absolute value of the calculated direction preference score is lower than the interference threshold, and do not use them as the basis for identifying the boundaries of chromatin interaction structural domains.

[0078] This method adds a "noise filtering" step to the aforementioned implementation, aiming to improve the reliability and robustness of CID boundary identification.

[0079] Specifically, before searching for sign reversal points, intervals with very small absolute values ​​of CDS scores (i.e., close to 0) can be filtered out. This means that only intervals exhibiting strong directional bias are eligible to participate in the boundary definition. The principle behind this is that biological signals are often accompanied by random noise. A CDS score very close to 0 may not necessarily mean that the region truly lacks directional bias, but rather that the actual signal is very weak and masked by random fluctuations during the experimental or computational process. A sign reversal from +0.01 to -0.01 is far less biologically significant than a reversal from +0.5 to -0.5. The principle behind this step is to set an "interference threshold" to retain only those clear signals with high signal-to-noise ratios, thereby avoiding misjudging noise as boundaries.

[0080] An "interference threshold" can be set, for example, 0.1; for each CDS (bin) calculated in the aforementioned implementation method... i ), calculate its absolute value |CDS(bin) i If |CDS(bini)| < 0.1, then the CDS score in this interval is considered invalid or "uninformative". Then, in the remaining valid score sequence, the aforementioned rules (steps S210 to S220) are applied to find the positions of adjacent score signs with opposite signs as boundaries, thereby obtaining a smaller, more reliable, and possibly more biologically meaningful list of CID boundaries.

[0081] This method effectively filters out random noise from interfering with boundary identification, making the final identified CID structure less sensitive to minor fluctuations in experimental data and resulting in more stable results. It can also help the method focus on chromatin domains that are structurally more stable and have clearer definitions.

[0082] For example, assuming the interference threshold is set to 0.2, a CDS score sequence is [..., +0.6, +0.15, -0.08, -0.5, ...]. The absolute values ​​of +0.15 and -0.08 are both less than 0.2, therefore they are considered invalid scores. A valid score sequence would be [..., +0.6, (invalid), (invalid), -0.5, ...]. Although there are two invalid scores between +0.6 and -0.5, they are "nearby" scores with opposite signs in the valid sequence. Therefore, a boundary can still be identified between the regions corresponding to +0.6 and -0.5. The sign reversal between +0.15 and -0.08 is successfully ignored.

[0083] Example 3 refer to Figure 6This embodiment provides a method for screening highly expressed integration sites in Saccharomyces cerevisiae. Before determining chromatin interaction domains on the genome based on the chromatin interaction direction preference in step S200, the method further includes: Step S200-3: Obtain high-throughput chromatin conformation capture data of at least two biological replicates of the same strain of Saccharomyces cerevisiae.

[0084] This step requires the preparation of at least two independent sets of experimental data.

[0085] The term "same strain" means that the experimental material (yeast strain) is exactly the same. "Biological replication" means that the two sets of data come from two or more independent experimental processes (e.g., extracting samples from different yeast cultures and independently completing the entire Hi-C library preparation and sequencing process).

[0086] It is important to note that this step is based on the principle of "reproducibility" in scientific research. A real, stable biological phenomenon (such as the chromatin domains described here) should be consistently observed in repeated, independent experiments. Conversely, findings that appear only in a single experiment and disappear in repeated experiments are more likely to be random noise or accidental products introduced by experimental procedures. Therefore, having at least two sets of biological replication data is a prerequisite for verifying the reliability of a finding, providing the necessary data foundation for subsequent stability assessments, and ensuring from the outset that the data upon which the method relies is not isolated evidence, thus increasing the credibility of the final conclusion.

[0087] Step S200-4: Set different interval length parameters to identify the corresponding chromatin interaction domain boundaries.

[0088] This step is an "exploratory testing" process. It does not presuppose a "correct" interval length, but actively explores multiple possibilities.

[0089] The aforementioned "setting different interval length parameters" can refer to creating a list of candidate parameters, such as [5kb, 10kb, 15kb, 20kb]. This length determines how many segments the genome will be cut into for analysis.

[0090] The phrase "identify the corresponding...boundaries" refers to performing the boundary identification process described in the preceding implementation (i.e., calculating directional preference, finding sign reversal points, etc.) completely for each candidate parameter in the list. Furthermore, this process needs to be run independently for each set of biological replicates.

[0091] The processing flow can be a nested loop process. For example, for each "interval length" (e.g., 10kb) in the list: (1) run the boundary recognition algorithm on the "first set of repeated data" to obtain "boundary list A"; (2) run the boundary recognition algorithm on the "second set of repeated data" to obtain "boundary list B"; repeat the above process for the next interval length; after this step, a series of boundary results can be obtained. For example, if 3 interval lengths are tested and there are 2 sets of repetitions, then 3×2=6 boundary lists will be obtained.

[0092] This method systematically explores the parameter space, avoiding biases or missing the optimal analytical scale that may result from subjective selection of a single parameter.

[0093] Step S200-5: Compare the overlap of the boundaries of the chromatin interaction domains among different biological replicate samples, and select the interval length parameter with the highest overlap as the final basis for genome division.

[0094] The step described above, "comparing the overlap of chromatin interaction domain boundaries between different biological replicates," is a "quantitative evaluation" of the exploration results from the preceding steps. The core objective is to compare how similar the lists of boundaries identified from different replicate experiments are for the same interval length.

[0095] The aforementioned "overlap" is a quantitative indicator for measuring the consistency between two sets of boundaries. A high overlap means that, within the current interval length, two independent experiments have consistently found very similar domain boundaries, indicating that these boundaries are real and reliable.

[0096] Specifically, "overlap" can be calculated in several ways. For example, a simple and intuitive algorithm is as follows: For a fixed interval length (e.g., 10kb), there are "boundary list A" and "boundary list B"; count the number of common boundaries (N) in list A and list B. common ); Count the total number of boundaries (N) of list A. A The total number of boundaries of list B (N) B The degree of overlap can be defined, for example, by the Jaccard index: Overlap = N common / (N A +N B -N common Or, to use a simpler ratio: Overlap = 2 × N common / (N A +N B ).

[0097] For example, to test a 10kb length: repeat A to find 100 boundaries; repeat B to find 105 boundaries; of which 90 boundaries are shared by both; the calculated overlap score is high (e.g., 2×90 / (100+105)≈0.88).

[0098] For example, when testing a 5kb length: repeating A finds 200 boundaries (higher resolution, more detail); repeating B finds 210 boundaries; of which only 120 are common (many details may be noise); the calculated overlap score is low (e.g., 2×120 / (200+210)≈0.58).

[0099] In the above method, the final "decision" step is to "select the interval length parameter with the highest overlap as the final basis for genome partitioning". This involves comparing all overlap scores calculated in the previous steps and selecting the interval length that produces the most stable and reproducible results.

[0100] Specific processing may include: summarizing all test interval lengths and their corresponding overlap scores, for example: {5kb: 0.58, 10kb: 0.88, 15kb: 0.85, ...}; finding the entry with the highest score; and determining the interval length corresponding to that entry (10kb in this example) as the standard division basis used for all subsequent analyses (including FIRE identification, final screening, etc.), thereby obtaining the optimal "interval length" parameter validated by experimental data.

[0101] This step provides a data-driven, objective basis for the selection of fundamental and critical parameters in the methodology. By selecting the analytical scale that best and most stably reproduces biological structures, the reliability and credibility of all subsequent steps and even the final screening results are greatly enhanced. It ensures that the entire analysis is built on a solid and robust foundation.

[0102] Furthermore, in addition to simple, precise boundary location matching, a more flexible definition of "overlap" can be used. For example, a certain range of boundary fluctuation (e.g., ±1 interval) can be allowed, and as long as the boundaries of two replicate experiments fall within this small window, they are considered to overlap. This can better tolerate minor localization errors introduced by sequencing and algorithms.

[0103] Furthermore, if three or more biological replicates are obtained, the strategy can be more robust. For example, one can select the interval length that produces the highest average overlap in the most replicated samples (e.g., ≥2 / 3 of the samples), or directly look for a "core boundary set" that appears consistently in all replicates.

[0104] In addition to overlap, other indicators can be considered when selecting the interval length, such as the average "insulation strength" (i.e., insulation score) of the CIDs identified at this scale. The goal is to select a "comprehensive optimal" interval length that ensures both high overlap and strong insulation signals.

[0105] Furthermore, to verify the effectiveness of the above parameter selection method, parameter sensitivity analysis was performed in the embodiments of this application. For example... Figure 7 As shown, the conservation trend of the CID boundary among biological replicates is illustrated by the variation in the length of the upstream and downstream interaction interval (k value); Figure 8 The diagram illustrates the impact of the interference threshold (p-value) on the conservatism of CID. By comparing the overlap (conservatism) under different parameters, the optimal calculation parameters can be determined (e.g., k=10 is preferred in this embodiment), thereby ensuring the stability of the recognition results.

[0106] Example 4 This embodiment provides a method for screening highly expressed integration sites in Saccharomyces cerevisiae. In step S300, the local interaction strength of each interval is calculated based on the high-throughput chromatin conformation capture data, including: Step S310: For each interval, calculate the sum of its interaction frequencies with all other intervals within a preset genomic distance range to obtain the local interaction strength of that interval.

[0107] The "local interaction strength" mentioned in this step is given a precise and operational definition. It refers to the fact that for each small "region" on the genome, the "local interaction strength" is not a vague concept, but a specific value that can be calculated by adding up the Hi-C interaction frequencies (or fractions) of all other regions within that region and its physical neighbor (defined by the "preset genomic distance range").

[0108] Essentially, local interaction strength measures the “local social activity” or “local centrality” of a genomic locus.

[0109] The basic principle is that chromatin is not randomly stacked within the cell nucleus, but rather folded in an organized manner. A genomic region that plays an important role in three-dimensional space (such as a structural anchor or regulatory center) tends to interact more frequently and strongly with its physically neighboring regions. Therefore, by calculating this "sum of local interactions," regions that may structurally be "hubs" can be preliminarily identified.

[0110] Specific algorithms may include: (1) Input: Normalized Hi-C interaction matrix; optimal interval length determined by the aforementioned method.

[0111] (2) Set parameters: Preset a “genome distance range”, for example, 200kb.

[0112] (3) Calculation: For each interval i on the genome: initialize “local interaction strength” LIS(i)=0; find all other intervals j whose linear distance from interval i is less than or equal to 200kb; for each found interval j, read the interaction frequency C from the Hi-C matrix. (i, j) ; will C (i, j) The values ​​are accumulated into LIS(i); the final LIS(i) is the local interaction strength of interval i.

[0113] This process generates a corresponding "local interaction strength" score for each region on the genome.

[0114] This method reduces the dimensionality of the complex, two-dimensional Hi-C interaction matrix information into a simple and intuitive one-dimensional fraction, which facilitates subsequent correlation analysis. By using a "preset distance range", it eliminates interference from long-distance interactions and focuses specifically on signals that reflect local chromatin compactness or interaction activity.

[0115] refer to Figure 9 In some embodiments, step S300 involves identifying frequently interacting regions with high local interaction strength that are positively correlated with the gene expression levels shown in the gene expression data, including: Step S320: Using a sliding window to slide along the genome, calculate the Pearson correlation coefficient between the local interaction strength and the gene expression level within each window.

[0116] The aforementioned "using a sliding window" is a standard computational method for systematically scanning the entire genome. It avoids the bias caused by manually dividing regions and can examine every corner of the genome without omission.

[0117] The above-mentioned "calculate...Pearson correlation coefficient" is the "association" step of the algorithm. The Pearson correlation coefficient is a standard statistical indicator that measures the strength of the linear relationship between two continuous variables, with a value between -1 and +1. Here, it is used to answer a key question: "Within a local region (window), does an increase in 'local interaction strength' accompany a synchronous increase in 'gene expression levels'?" A coefficient close to +1 provides a definitive answer to this question.

[0118] Step S330: Select the region where the Pearson correlation coefficient is greater than a preset correlation threshold and the local interaction strength is higher than a preset strength threshold as the frequent interaction region.

[0119] This step is the "filtering" step of the algorithm, and it is also the essence of the "double threshold" mechanism.

[0120] The aforementioned correlation thresholds ensure that the "structure-function" correlation in the selected region is real and significant, rather than accidental.

[0121] The above intensity thresholds ensure that the selected area itself is a structurally active "hotspot" and exclude "cold" areas that are related but have a low overall level of interaction.

[0122] The algorithm combines the "structural" information calculated in the aforementioned implementation with the obtained "functional" information.

[0123] Specifically, two data tracks aligned along genome coordinates can be prepared: one based on the calculated "local interaction strength" track, and the other on the "gene expression level" track. Parameter settings can include: sliding window size W (e.g., 40kb); sliding step size S (e.g., 10kb); and Pearson correlation coefficient threshold T. corr (e.g., 0.6); Local interaction strength threshold T strength (e.g., set to the 90th percentile of this value across the entire genome).

[0124] The algorithm execution may specifically include: starting from the chromosome origin, taking the first window of size W; extracting two sets of data within this window: "local interaction strength" and "gene expression level"; calculating the Pearson correlation coefficient R between these two sets of data; and calculating the average value Avg of the "local interaction strength" within this window. LIS If (R>T) corr And (Avg) LIS >T strength If the coordinates of this window region are recorded and marked as FIRE, then move the window forward by S and repeat the above process until the end of the chromosome.

[0125] The above method yields a high-confidence list of FIRE regions with clearly defined coordinates. Using the classic Pearson correlation coefficient ensures that the "positive correlation" judgment has clear statistical significance, rather than being a subjective assessment. The "dual threshold" design provided by this method, equivalent to a lock requiring two keys, effectively filters out false positives, ensuring that the identified FIREs are truly elite regions possessing both "high structural activity" and "strong functional associations." The sliding window mechanism guarantees an unbiased and systematic scan of the entire genome; the entire algorithm is clear and easily automated through computer programs.

[0126] like Figure 10As shown, the specific identification process for frequently interacting regions is illustrated, including data preprocessing, local interaction strength calculation, and integration analysis with the expressed data.

[0127] Furthermore, the Pearson correlation coefficient measures a linear relationship. If a non-linear relationship between structure and function is considered (e.g., expression levels tend to saturate after the interaction strength reaches a certain level), the Spearman rank correlation coefficient can be used instead or supplemented. Spearman's coefficient does not require the data to be linear; it only cares whether the order of the two variables is consistent, thus making it more robust.

[0128] Furthermore, a fixed "preset threshold" may not be suitable for all situations. More statistically significant adaptive thresholds can be used. For example, a permutation test can be used to determine the correlation threshold: the gene expression data are randomly shuffled multiple times, and the distribution of their correlation coefficients with the interaction strength is calculated to determine a statistically significant threshold (e.g., p-value < 0.05).

[0129] Furthermore, based on the "dual threshold," "triple" or even "multiple thresholds" can be introduced. For example, additional chromatin accessibility data (such as ATAC-seq) can be introduced, requiring that the final FIRE region not only meets the intensity and relevance thresholds, but also must be located in the "open" region of chromatin, further improving the accuracy of screening.

[0130] Example 5 This embodiment provides a method for screening highly expressed integration sites in Saccharomyces cerevisiae. In step S200, after identifying chromatin interaction domains on the genome, the method further includes: Step S200-6: Calculate the multidimensional structural features of each of the chromatin interaction domains.

[0131] The method provided in this embodiment is an optimization and refinement of the method described in Embodiment 1 above. It does not change the core logic of "boundary + FIRE", but adds a crucial "pre-screening" step, which first classifies all "chromatin interaction domains (CIDs)" by function, and then searches for the final integration site only in the "optimal" category.

[0132] This method introduces a two-stage filtering strategy of "classification first, then screening." Its core mechanism is that CIDs with similar three-dimensional structural features may also have similar intrinsic gene regulatory environments and functional states. By associating structure with function, we can identify which type of structure corresponds to the most ideal "high expression" environment.

[0133] For each identified CID, its complete "structural fingerprint" vector is calculated, thus obtaining a dataset in which each row represents a CID and each column represents a structural feature (such as average distance, insulation score, etc.).

[0134] Step S200-7: Based on the multi-dimensional structural features, perform cluster analysis on all the chromatin interaction domains to classify the chromatin interaction domains into different categories.

[0135] This step is an unsupervised learning step. It takes the dataset obtained in the previous step as input and allows the algorithm to automatically cluster all CIDs according to the similarity of their "structural fingerprints".

[0136] Specifically, the classic K-Means clustering algorithm can be used. The number of clusters to be divided is preset (e.g., K=3). The algorithm iteratively calculates and ultimately assigns each CID to one of the three clusters, making the structural features of CIDs within the same cluster as similar as possible, and maximizing the differences in structural features between CIDs in different clusters. This results in all CIDs being divided into several (e.g., three) different structural categories.

[0137] In addition to K-Means, other clustering algorithms can be tried, such as hierarchical clustering, which does not require a pre-defined number of clusters and can reveal the hierarchical relationship between clusters; or DBSCAN, which can identify clusters of arbitrary shapes and handle noise points well.

[0138] Step S200-8: Based on the gene expression data, assess the average transcriptional level of chromatin interaction domains in each category.

[0139] In this step, "functional annotations" are added to the "structural categories" identified in the previous step.

[0140] Specifically, for each category (e.g., category A): find all CIDs belonging to category A; identify all genes within the genomic regions covered by these CIDs; and use the obtained gene expression data to calculate the average expression level (average transcription level) of these genes. This assigns a "functional score" (i.e., average transcription level) to each structural category.

[0141] Step S200-9: The category with the highest average transcription level is identified as the target category.

[0142] In this step, the "functional scores" of all categories are compared, and the category with the highest score is identified as the "target category" most associated with high transcriptional activity.

[0143] Furthermore, step S400, which involves screening genomic sites located at the boundary of the chromatin interaction domain and overlapping with or adjacent to the frequently interacting regions, includes: Step S410: Select genomic sites that are located at the boundary of the chromatin interaction domain of the target category and overlap with or are adjacent to the frequently interacting regions.

[0144] This step is the final, refined screening step.

[0145] Specifically, we can consider only those CIDs that belong to the "target category" determined in the previous step, and then obtain the boundary locations of these "target CIDs". In these "elite boundaries", we can then apply the second criterion in the aforementioned implementation: filter out those sites that overlap with or are adjacent to the FIRE region.

[0146] The method presented in this embodiment no longer treats all CID boundaries "equally." Instead, it uses a biologically meaningful, data-driven filtering layer to first identify CIDs that are already in a "high expression-friendly" structural environment. This greatly eliminates interfering sites located at the boundaries of "silent" or "inactive" structural domains, thereby significantly improving the "hit rate" of the final screening.

[0147] This method not only identifies "where integration occurs" but also answers "why integration occurs here." By analyzing the structural characteristics of "target categories" (e.g., finding that they generally exhibit "high insulation scores" and "low average distances"), it reveals which specific three-dimensional structural conformations in Saccharomyces cerevisiae are associated with high gene expression, which is itself a valuable scientific discovery.

[0148] Furthermore, if experiments have verified the accumulation of a known set of "high-expression" and "low-expression" CID samples, the problem can be upgraded from unsupervised "clustering" to supervised "classification." A machine learning classifier (such as Support Vector Machine (SVM), Random Forest, etc.) can be trained to directly learn how to predict whether a CID belongs to the "high-expression" category based on structural features, and its prediction accuracy is usually higher.

[0149] In some implementations, the multi-dimensional structural features include at least one of the following: (1) Average domain distance: The interaction density within a domain is calculated by converting the contact matrix into a distance matrix. It is used to measure the compactness within a CID. This "distance" is not a physical distance, but a "genomic distance" calculated from the interaction frequency. The higher the interaction frequency, the closer the "distance".

[0150] Specifically, the contact matrix (C(i,j)) of Hi-C can be transformed into a distance matrix (e.g., D(i,j) = 1 / C(i,j)). The "average distance" of a CID is the average of the D(i,j) values ​​of all interval pairs within it. A low average distance indicates that the CID interacts very frequently and is structurally more compact.

[0151] (2) Domain Insulation Score: The absolute value of the chromatin interaction direction preference at the domain boundary. Used to measure the “isolation strength” or “clarity” of a CID boundary.

[0152] Specifically, this value is directly derived from the calculation results in the aforementioned implementation method. It is the absolute value of the "chromatin interaction direction preference score (CDS)" at the CID boundary location. A high absolute value means that at the boundary, the interaction direction has undergone a sharp and clear reversal, indicating that the boundary has a good "insulation" effect and can effectively isolate the internal and external environments.

[0153] (3) Average contact score of the structural domain: The total interaction strength within the structural domain. It is used to measure the total interaction activity within a CID.

[0154] Specifically, the interaction frequencies (C(i,j)) between all interval pairs within a CID are summed up to obtain a total. This total represents the overall interaction strength within that CID. A higher score means more frequent internal "communication".

[0155] (4) Domain contact variation coefficient: The degree of variation in the interaction strength within a domain. It is used to measure the uniformity of the interaction strength within a CID.

[0156] Specifically, this parameter can be a statistical indicator, typically calculated as "standard deviation / mean". It measures the dispersion of all interaction frequency values ​​within a CID. A low coefficient of variation means that the internal interactions are relatively evenly distributed; a high coefficient of variation means that there are a few "interaction hotspots" within the CID, while other areas have fewer interactions.

[0157] The above four features, from the four dimensions of compactness, boundary strength, overall activity, and activity uniformity, together construct a comprehensive and quantitative "structural profile" for each CID, providing rich and distinguishable input data for subsequent cluster analysis.

[0158] Furthermore, more bioinformatics data can be integrated as features. A highly promising extension is the incorporation of epigenetic data, such as: (1) Histone modification signal: Calculate the average signal intensity of activating modifications (such as H3K4me3, H3K27ac) and repressive modifications (such as H3K27me3) within each CID.

[0159] (2) Chromatin accessibility: The average ATAC-seq signal of each CID is calculated as a feature to measure its “openness”.

[0160] By incorporating these features into cluster analysis, a more comprehensive and accurate "structure-function" profile can be constructed.

[0161] Example 6 This embodiment aims to verify the robustness and universality of the screening method described in the foregoing embodiments.

[0162] 1. Experimental methods and grouping: To comprehensively evaluate the method's performance, this embodiment designed three levels of verification experiments: (1) Biological replication verification: The standard strain of Saccharomyces cerevisiae BY4741 was selected, and two biologically replicated Hi-C samples were prepared independently under the same conditions.

[0163] (2) Data downsampling verification: Based on the original BY4741 data, random sampling was used to simulate a dataset with sequencing depth gradually reduced from 90% to 30%.

[0164] (3) Verification of multiple genetic backgrounds: Five non-standard strains, including strains that have undergone random genome rearrangement (yKJP020, yKJP048, yKJP058), triploid strain (J005), and haploid strain (JCR27), were selected for testing.

[0165] 2. Experimental Results and Analysis: (1) Repeatability analysis: The directional preference score (CDS) for each region of the genome was calculated using the method described in this application. The results are as follows: Figure 11 As shown, the CDS values ​​calculated from two biological replicates exhibit a high degree of linear correlation (R=0.85), indicating that the method has excellent reproducibility.

[0166] (2) Deep dependency analysis: The results are as follows Figure 12 As shown, even when the amount of data is reduced to 30% of the original data, the correlation coefficient between the calculated results and the original full data is still greater than 0.84, proving that this method is not sensitive to sequencing depth and can effectively reduce experimental costs.

[0167] (3) Universality analysis: Analysis of strains with different genetic backgrounds revealed (data are shown in Table 1 and Table 2 below).

[0168] Table 1. Information on six strains of Saccharomyces cerevisiae

[0169] Table 2. Information on CID of Saccharomyces cerevisiae

[0170] Regardless of whether the genome has undergone rearrangement or ploidy changes, this method can identify structural domains with a median length consistently around 10 kb. This indicates that this method can capture biologically significant three-dimensional chromatin functional units that are ubiquitous in Saccharomyces cerevisiae.

[0171] Example 7 This embodiment aims to verify the quantitative correlation between the boundary and structural features of the “chromatin interaction domain (CID)” defined in this application and gene transcription activity.

[0172] 1. Experimental Method: (1) Enrichment analysis: Based on the "highly expressed genes" defined by RNA-seq data, the distribution frequency of these genes at different distances upstream and downstream of the CID boundary was statistically analyzed.

[0173] (2) Comparison of differences: The whole genome was divided into “CID boundary region” and “non-boundary region”, and the expression levels (TPM values) of genes in the two regions were compared.

[0174] (3) Feature correlation: The length of CID, average transcriptional activity and average chromatin accessibility (ATAC-seq data) were calculated, and the correlation among the three was analyzed.

[0175] 2. Experimental Results and Analysis: (1) Boundary enrichment effect: The results are as follows Figure 13 As shown, in all six strains tested, the highly expressed genes reached their peak at a distance of 0 from the CID boundary, indicating that the CID boundary is enriched with highly transcribed genes.

[0176] (2) Comparison of expression levels: Results are as follows Figure 14 As shown in the violin plot, genes located at the CID boundary have significantly higher average transcription levels than those in non-boundary regions (P<0.001), demonstrating that the boundary is a key functionally active region.

[0177] (3) Influence of structural features: The results are as follows Figure 15 As shown, the average transcriptional activity of CID was significantly positively correlated with its chromatin accessibility. Figure 15 (A), and is significantly negatively correlated with CID length ( Figure 15 (B). This reveals that a compact, open CID structure is more conducive to high-level expression, providing a theoretical basis for subsequent multi-dimensional screening.

[0178] Example 8 This embodiment is a specific verification of the clustering screening strategy described in Embodiment 5 above.

[0179] 1. Experimental Method: The average distance of CIDs, insulation score, average contact score, contact coefficient of variation, and chromatin accessibility were selected as feature dimensions to perform a consistent cluster analysis on the CIDs identified by the standard strain BY4741.

[0180] 2. Experimental Results and Analysis: (1) Clustering results: Analysis using the cumulative distribution function (e.g.) Figure 16 As shown in the figure, the optimal number of clusters is determined to be k=6, that is, CID can be stably divided into 6 structural categories.

[0181] (2) Category Feature Analysis: The features of these 6 CID categories are compared, and the results are as follows: Figure 17 to Figure 22 As shown. Figure 17 The study revealed significant differences in average transcriptional activity across various CID types, exhibiting a tiered distribution. Combined with... Figure 18 (Accessibility) Figure 19 (Insulation score) and Figure 20 (Internal distance) analysis shows that CID categories with characteristics such as "high chromatin accessibility", "low insulation score" and "specific internal distance" (e.g.) Figure 17 The fourth category (in the text) corresponds to the highest transcriptional activity.

[0182] This indicates that the cluster analysis in this application can accurately identify “high expression potential CID categories” with specific structural fingerprints, thereby further screening for integration sites within these categories and significantly improving the screening success rate.

[0183] Example 9 This embodiment verifies the functionality of Frequently Interacting Regions (FIRE) and demonstrates the final application effect of joint screening of CID and FIRE.

[0184] 1. Experimental Method: Calculate the local interaction strength across the entire genome and perform correlation analysis with gene expression data to identify FIRE regions. Compare gene expression levels between high-FIRE and low-FIRE regions.

[0185] 2. Experimental Results and Analysis: (1) Correlation between interaction and expression: Genome-wide frequent interaction score map as shown in Figure 23 As shown. Statistical analysis shows (e.g.) Figure 24 As shown in the figure, the intensity of local interactions on chromosomes is positively correlated with the overall gene expression level.

[0186] (2) High expression characteristics of the FIRE region: The results are as follows Figure 25 As shown, genes located in the High-FIRE region have significantly higher transcription levels than those in the Low-FIRE region, demonstrating that FIRE is another independent indicator of high expression.

[0187] (3) Joint screening strategy: Given that both the CID boundary and FIRE indicate high expression, the joint screening strategy proposed in this application is as follows: Figure 26 As shown. This strategy prioritizes genomic sites that are "located on the boundary of the target CID category and overlap with or are adjacent to the High-FIRE region." This dual-screening mechanism utilizes two independent structural features to mutually confirm the selected sites (i.e., Figure 26 The overlapping region (indicated by the dashed box) has the highest confidence level and is an ideal exogenous gene integration site for constructing efficient cell factories.

[0188] refer to Figure 27 This application also provides a screening device for highly expressed integration sites in Saccharomyces cerevisiae, comprising: Acquisition module 10 is used to acquire high-throughput chromatin conformation capture data and gene expression data of the Saccharomyces cerevisiae genome; The calculation module 20 is used to divide the genome into multiple continuous intervals, calculate the chromatin interaction direction preference of each interval based on the high-throughput chromatin conformation capture data, and determine the chromatin interaction domains on the genome according to the chromatin interaction direction preference. The identification module 30 is used to calculate the local interaction intensity of each interval based on the high-throughput chromatin conformation capture data, and in combination with the gene expression data, identify frequently interacting regions with high local interaction intensity and positive correlation with the gene expression level shown by the gene expression data. The screening module 40 is used to screen out genomic sites located at the boundary of the chromatin interaction domain and overlapping or adjacent to the frequently interacting region as the highly expressed integration sites.

[0189] It is understood that the device in this embodiment corresponds to the screening method for highly expressed integration sites of Saccharomyces cerevisiae in the above embodiments, and the options in the above embodiments are also applicable to this embodiment, so they will not be described again here.

[0190] This application also provides a computer device, which includes a processor and a memory. The memory stores a computer program, and the processor is used to execute the computer program to implement the screening method for highly expressed integration sites of Saccharomyces cerevisiae as described in any of the foregoing embodiments.

[0191] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0192] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving execution instructions.

[0193] This application embodiment also provides a computer storage medium storing a computer program, which, when executed on a processor, implements the screening method for highly expressed integration sites of Saccharomyces cerevisiae according to any one of the foregoing embodiments.

[0194] For example, the computer storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0195] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, in alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0196] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0197] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0198] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for screening highly expressed integration sites in Saccharomyces cerevisiae, characterized in that, include: To acquire high-throughput chromatin conformation capture data and gene expression data of the Saccharomyces cerevisiae genome; The genome is divided into multiple consecutive regions, and the chromatin interaction direction preference of each region is calculated based on the high-throughput chromatin conformation capture data. The chromatin interaction domains on the genome are then determined based on the chromatin interaction direction preference. Based on the high-throughput chromatin conformation capture data, the local interaction intensity of each interval is calculated, and combined with the gene expression data, frequent interaction regions with high local interaction intensity and positive correlation with the gene expression level shown by the gene expression data are identified. Genomic sites located at the boundary of the chromatin interaction domain and overlapping with or adjacent to the frequently interacting regions are selected as the highly expressed integration sites.

2. The method for screening highly expressed integration sites in *Saccharomyces cerevisiae* as described in claim 1, characterized in that, The expression for calculating the chromatin interaction direction preference for each of the aforementioned intervals is as follows: ; Among them, CDS (bin) i ) represents the chromatin interaction direction preference score of the i-th interval; The average of the scores between the i-th interval and the k upstream intervals is calculated as follows: C (i,j) This represents the interaction score between the i-th interval and the j-th upstream interval; The average of the interaction scores between the i-th interval and the k downstream intervals is calculated as follows: C (i,m) The value represents the interaction score between the i-th interval and the m-th downstream interval; k represents the number of intervals used to calculate the interaction range, determined according to a preset resolution.

3. The method for screening highly expressed integration sites in *Saccharomyces cerevisiae* as described in claim 2, characterized in that, The step of determining chromatin interaction domains on the genome based on the chromatin interaction direction preference includes: Identify locations on the genome where the directional preference scores of two adjacent regions have opposite signs; The opposite position of the symbol is defined as the boundary of the chromatin interaction domain; and / or, Before determining the chromatin interaction domains on the genome based on the chromatin interaction direction preference, the process also includes: Set the interference threshold; The intervals where the absolute value of the calculated directional preference score is lower than the interference threshold are excluded and are not used as the basis for identifying the boundaries of chromatin interaction structural domains.

4. The method for screening highly expressed integration sites in *Saccharomyces cerevisiae* as described in claim 1, characterized in that, Before determining the chromatin interaction domains on the genome based on the chromatin interaction direction preference, the process also includes: Obtain high-throughput chromatin conformation capture data from at least two biological replicates of the same strain of Saccharomyces cerevisiae; Different interval length parameters were set to identify the boundaries of the corresponding chromatin interaction domains. Compare the overlap of chromatin interaction domain boundaries among different biological replicates, and select the interval length parameter with the highest overlap as the final genome partitioning criterion; and / or, The calculation of the local interaction strength of each interval based on the high-throughput chromatin conformation capture data includes: for each interval, calculating the sum of the interaction frequencies of it and all other intervals within a preset genomic distance range to obtain the local interaction strength of that interval.

5. The method for screening highly expressed integration sites in *Saccharomyces cerevisiae* as described in claim 1, characterized in that, The identification of frequently interacting regions with high local interaction strength and positive correlation with the gene expression levels shown in the gene expression data includes: A sliding window was used to slide along the genome, and the Pearson correlation coefficient between the local interaction strength and the gene expression level within each window was calculated. The regions where the Pearson correlation coefficient is greater than a preset correlation threshold and the local interaction strength is higher than a preset strength threshold are selected as the frequent interaction regions.

6. The method for screening highly expressed integration sites in *Saccharomyces cerevisiae* as described in claim 1, characterized in that, After identifying the chromatin interaction domains on the genome, the following steps are also included: Calculate the multidimensional structural features of each of the chromatin interaction domains; Based on the multi-dimensional structural features, cluster analysis is performed on all the chromatin interaction domains to classify the chromatin interaction domains into different categories. Based on the gene expression data, the average transcriptional level of chromatin interaction domains in each category was assessed; The category with the highest average transcription level was identified as the target category. Furthermore, the screening of genomic sites located at the boundary of the chromatin interaction domain and overlapping with or adjacent to the frequently interacting regions includes: Genomic sites that are located at the boundary of the chromatin interaction domain of the target category and overlap with or are adjacent to the frequently interacting regions are selected.

7. The method for screening highly expressed integration sites in *Saccharomyces cerevisiae* as described in claim 6, characterized in that, The multi-dimensional structural features include at least one of the following: Average distance between domains: the interaction density within a domain calculated by converting the contact matrix into a distance matrix; Domain insulation score: the absolute value of the chromatin interaction direction preference at the domain boundary; Average contact score of structural domain: the total interaction strength within the structural domain; Domain contact variation coefficient: the degree of variation in the interaction strength within a domain.

8. A screening device for highly expressed integration sites in Saccharomyces cerevisiae, characterized in that, include: The acquisition module is used to acquire high-throughput chromatin conformation capture data and gene expression data of the Saccharomyces cerevisiae genome; The computation module is used to divide the genome into multiple continuous regions, calculate the chromatin interaction direction preference of each region based on the high-throughput chromatin conformation capture data, and determine the chromatin interaction domains on the genome according to the chromatin interaction direction preference. The identification module is used to calculate the local interaction intensity of each interval based on the high-throughput chromatin conformation capture data, and in combination with the gene expression data, identify frequently interacting regions with high local interaction intensity and positive correlation with the gene expression level shown by the gene expression data. The screening module is used to screen out genomic sites located at the boundary of the chromatin interaction domain and overlapping with or adjacent to the frequently interacting regions, as the highly expressed integration sites.

9. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the screening method for highly expressed integration sites of *Saccharomyces cerevisiae* according to any one of claims 1-7.

10. A computer storage medium, characterized in that, It stores a computer program that, when executed on a processor, implements the screening method for highly expressed integration sites in Saccharomyces cerevisiae according to any one of claims 1-7.

Citation Information

Patent Citations

  • Ssi cells with predictable and stable transgene expression and methods of formation

    CN113227388A

  • Method for high-yield biosynthetic gene or gene cluster product based on chromatin three-dimensional structure

    CN116153406A

  • Method for identifying chromosome translocation by integrating three-dimensional genome and three-generation genome data

    CN119541631A