Method and device for determining binding sites for binding proteins to nucleic acids, in particular for binding transcription factors to DNA
By using short laser pulses to convert hydrogen bonds into covalent bonds and employing nanopore sequencing, the method addresses inefficiencies and DNA damage in existing methods, enabling rapid and accurate determination of protein binding sites to nucleic acids.
Patent Information
- Application Number
- PCT/EP2025/054757
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-22
- Filing Date
- 2025-02-21
- Publication Date
- 2025-08-28
AI Technical Summary
Existing methods for determining protein binding sites to nucleic acids, such as FLIX-MS, suffer from complex procedures and undesired DNA damage due to femtosecond laser pulses, leading to inefficiencies and potential loss of information.
The method employs short laser pulses of less than 100 ns to convert hydrogen bonds into covalent bonds between proteins and nucleic acids, using nanopore sequencing to directly determine binding sites without extensive sample preparation, and employs nanopore technology to sequence nucleic acids with fixed bonds, allowing for direct determination of binding sites.
This approach enables rapid and efficient determination of protein binding sites with reduced DNA damage, facilitating automated and accurate identification of binding sites using nanopore sequencing technology.
Smart Images

Figure EP2025054757_28082025_PF_FP_ABST
Abstract
Description
[0001] Method and device for determining binding sites of proteins to nucleic acids, in particular transcription factors to DNA
[0002] TECHNICAL FIELD OF THE INVENTION
[0003] The invention relates to a method and a device for determining binding sites of proteins to nucleic acids. The focus is on interactions between transcription factors and DNA. Furthermore, the invention allows for the investigation of diverse interactions between DNA-binding proteins and DNA. More specifically, the invention relates to a method having the features of the preamble of independent patent claim 1 and to a device for carrying out such a method.
[0004] STATE OF THE ART
[0005] A method with the features of the preamble of independent patent claim 1 is known from Reim, A. et al., Atomic-resolution mapping of transcription factor-DNA interactions by femtosecond laser crosslinking and mass spectrometry, Nature Communications 11, 3019; 10.1038 / s41467-020-16837-x, 2020. In this method, also known as FLIX-MS, bindings of transcription factors to DNA are fixed by crosslinking using femtosecond laser pulses. Hydrogen bonds are converted into covalent bonds. The DNA is then degraded down to a few nucleotides, and the proteins are cleaved. After purification and enrichment of the peptide-nucleotide crosslinking products, they are analyzed by mass spectrometry to draw conclusions about the binding sites of DNA-protein interactions. The binding sites and attached peptides are calculated using computer algorithms. Overall, this procedure is comparatively complex.
[0006] In addition to the desired cross-linking of proteins initially covalently bound to the DNA, undesired changes in the DNA also occur due to the effect of femtosecond laser pulses with wavelengths in the UV range, see Russmann, C. et al. Crosslinking of progesterone receptor to DNA using tuneable nanosecond, picosecond and femtosecond UV laser pulses. Nucleic Acids Res. 1997 Jun 15; 25(12):2478-84, 10.1093 / nar / 25.12.2478. Here, a distinction must be made between mono- and biophotonic damage to the DNA. Due to the high intensities of the femtosecond laser pulses, biphotonic damage such as strand breaks, which arise from the simultaneous absorption of two photons, predominates. Monophotonic damage, such as pyrimidine dimers, occurs less frequently. However, these changes are not critical in the FLIX-MS method due to the subsequent degradation of the DNA.
[0007] Nanopore sequencing is a DNA sequencing method commercialized by Oxford Nanopore Technologies PLC. This nanopore technology generates electrical signals known as squiggle signals. These raw signals are assigned to the four canonical nucleotides using a bioinformatics method called basecalling. Modifications such as DNA methylation are also detected using appropriate basecalling models. The basecalling models used in basecalling were trained using known DNA sequences.
[0008] From Valenzuela-Gömez, F. et al., Nanopore sensing reveals a preferential pathway for the co-translocational unfolding of a conjugative relaxase-DNA complex, Nucleic Acids Research, Vol. 51, No. 13, 6357-6869, 10.1093 / nar / gkad492, 2023, it is known that a nanopore technology, which in this case does not originate from Oxford Nanopore Technologies PLC, can be used to analyze a protein-DNA complex that is many times larger than a single DNA strand.
[0009] From Georgieva, D. et al. Detection of base analogs incorporated during DNA replication by nanopore sequencing, Nucleic Acids Research, Vol. 48, Issue 15, September 4, 2020, page e88, 10.1093 / nar / gkaa517, the application of the MinlON tool from Oxford Nanopore Technologies PLC is known to detect various thymidine analogues, including CldU, BrdU, IdU, and EdU, alone or coupled to biotin and other larger adducts in synthetic DNA templates.
[0010] OBJECT OF THE INVENTION
[0011] The invention is based on the object of demonstrating a method and a device capable of determining protein binding sites to DNA at the molecular level. To this end, molecular binding sites of DNA-binding proteins should be able to be determined directly with less effort and more quickly than previously possible.
[0012] SOLUTION
[0013] The object of the invention is achieved by a method having the features of independent claim 1. Dependent claims 2 to 8 describe preferred embodiments of the method according to the invention. Claim 9 is directed to a device for carrying out the method according to the invention; and claim 10 relates to a preferred embodiment of the device according to the invention.
[0014] DESCRIPTION OF THE INVENTION
[0015] In the method according to the invention for determining binding sites of proteins to nucleic acids, bonds of the proteins to the nucleic acids are fixed by short laser pulses with a pulse duration of not more than 100 ns, and the nucleic acids with the fixed bonds are sequenced by passing them through nanopores.
[0016] Surprisingly, commercially available nanopore technology, as well as self-produced and / or application-specific nanopores, e.g., based on α-hemolysin, can be successfully used to sequence nucleic acids despite the bonds fixed by the laser pulses and the associated unintended changes to the nucleic acids, such as strand breaks and the formation of pyrimidine dimers, allowing the binding sites of proteins to the nucleic acids to be determined. The invention combines two easily automated process steps without complex sample preparation: cross-linking the proteins to the nucleic acids and sequencing the nucleic acids with the fixed bonds.
[0017] The short laser pulses are used to fix the bonds between proteins and nucleic acids in isolated substances (e.g., recombinant proteins and oligonucleotides or plasmid DNA) or directly in a biological culture (e.g., cell culture, tissue culture, organoids, tumor spheroids). In any case, however, the entire method according to the invention is carried out outside the human or animal body. The short laser pulses, with a pulse duration of no more than 100 ns, are designed in the method according to the invention to convert hydrogen bonds between proteins and nucleic acids into covalent bonds without the involvement of additional molecules in the bonding. The short laser pulses therefore cause cross-linking directly between the amino acid and its binding site on the nucleic acid, i.e., a so-called "zero-length crosslinking."Particularly suitable short laser pulses have pulse lengths in the femtosecond to picosecond range, with pulse durations of less than 1 ps, i.e., so-called femtosecond laser pulses, being preferred, and wavelengths in the UV range. Laser pulses in the UV range can be combined with further laser pulses, even of longer wavelengths, at a defined time interval; see Russmann, C. et al., Two wavelength femtosecond laser induced DNA-protein crosslinking. Nucleic Acid Research, 1998. A significant reduction in the DNA damage described above is advantageous here. The use of high-intensity visible pulses is also possible, e.g., at a multiple of the UV wavelength; see DE 10 2010 020 194 B4.
[0018] Further possible configurations of the short laser pulses can be found in the literature on so-called femtosecond laser-induced DNA-protein crosslinking (FLIX). The physicochemical principles of this technology are also described there.
[0019] The nucleic acids to which the method according to the invention relates are, in particular, DNA, with the proteins in question being, in particular, transcription factors. The DNA to which the proteins bind is usually double-stranded DNA. Alternatively, the nucleic acids can be single-stranded DNA or RNA, although appropriate adaptation of sample preparation prior to the application of nanopore technology may be necessary. When sequencing double-stranded DNA using nanopore technology, the double-stranded DNA is split into two single strands, one of which is then randomly selected and sequenced. It is generally advantageous to sequence both single strands of the double-stranded DNA. If only one of the two strands is sequenced, information about fixed bonds may be lost.Through so-called duplex basecalling, both strands can be sequenced directly one after the other and complementary information is available.
[0020] In the method according to the invention, the nucleic acids can be sequenced using nanopore technology directly after the binding of the proteins to the attached proteins has been established. Higher molecular weight proteins, such as transcription factors, bind to the nanopores, meaning that once the proteins reach the entrance of the nanopore, they block the respective nucleic acid in the respective nanopore. As a result, a squiggle signal generated during nanopore technology ends at a point representative of the binding site of the respective protein to the respective nucleic acid.
[0021] In the method according to the invention, the proteins are preferably degraded to such an extent before sequencing the nucleic acids that the nucleic acids can pass through the nanopores over the bonds. The resulting squiggle signals are then directly influenced by the binding to specific nucleic bases of the nucleic acids and the protein residues bound to them. Thus, the binding sites and the bound residues can be deduced from the squiggle signals.
[0022] When analyzing biological cultures, one approach—when specific binding sites of a known transcription factor are sought in the genome—is to perform chromatin immunoprecipitation (ChIP), see Orlando V., Mapping chromosomal proteins in vivo by formaldehyde-crosslinked-chromatin immunoprecipitation., Trends Biochem Sci. 2000, 25(3) pp. 99-104. In this case, this involves cross-linking using FLIX, lysis of the cell membranes, and fragmentation of the chromatin into pieces of several hundred base pairs using ultrasound and / or nucleases, immunoprecipitation, and subsequent sample preparation and analysis using nanopore sequencing.
[0023] Specifically, before sequencing the nucleic acids, the proteins can be degraded down to one to a maximum of ten amino acids, preferably to one to a maximum of three amino acids, and most preferably to one amino acid, namely the one with the fixed bond to the respective nucleic acid. Complete degradation of the protein in this respect is easily achieved with various proteases. It also ensures smooth passage of the nucleic acids with the attached protein residues through the nanopores. Proteinase K, for example, can be used for this purpose. This serine protease has a broad spectrum of activity. It hydrolyzes a large number of peptide bonds, which leads to the breakdown of proteins into smaller polypeptides and ultimately into individual amino acids. Proteinase K is typically used for the proteolysis step during DNA extraction from tissue or cell culture samples.
[0024] Protein degradation can be followed by purification of the nucleic acids with the attached protein residues from other degradation products. High-quality, ultrapure DNA is essential for many molecular biology projects. Established protocols such as phenol-chloroform extraction or commercially available DNA extraction kits for next-generation sequencing can be used.
[0025] An analysis step during or after DNA sequencing involves the assignment of specific squiggle signals to the binding sites of bound peptides or proteins – in this case, useful signals – and their differentiation from canonical nucleotide signals. This so-called basecalling, or a so-called basecalling model used in the process, can be trained using controlled samples of known nucleic acid-amino acid interactions and potential interference signals from DNA damage, such as pyrimidine dimers or hydrates.
[0026] Basecalling training involves machine learning or deep learning methods. The generated algorithm assigns the binding site (in the form of a nucleobase) and the bound amino acid to the useful signals, and the original nucleobases (e.g., the individual thymines in pyrimidine dimers) to the interfering signals.
[0027] The initial binding of the proteins to the nucleic acids can already occur before the method according to the invention is carried out. However, the nucleic acids can also be contacted with one or more proteins only at the beginning of the method according to the invention in order to bind these proteins to the nucleic acids.
[0028] A device according to the invention for the automated implementation of the method according to the invention comprises a fixation device that provides laser radiation in the form of a short laser pulse or pulse sequences with a single pulse duration of no more than 100 ns to fix protein bonds to nucleic acids in isolated substances (e.g., recombinant proteins and oligonucleotides or plasmid DNA) or directly in biological cultures (e.g., cell culture, tissue culture, organoids, tumor spheroids), and a nanopore sequencing device, wherein the nanopore sequencing device is connected to the fixation device to sequence the nucleic acids with the fixed bonds. The laser radiation can irradiate the entire sample at once, e.g., in an Eppendorf tube or a quartz cuvette, and scan, e.g.,in a cell culture dish or a microtiter plate, or in a microfluidic system, see for example Nebbioso, A. et al, Time-resolved analysis of DNA-protein interactions in living cells by UV laser pulses. Sci Rep. 2017 Sep 15;7(1): 11725. doi: 10.1038 / s41598-017-12010-5. A sample processing device can be connected between the fixing device and the nanopore sequencing device, which is configured to degrade the bound proteins down to one to a maximum of ten amino acids, preferably down to one to a maximum of three and most preferably down to a single amino acid bound to the respective nucleic acid. The degradation of the proteins can also be implemented in more complex samples such as cells or tissue. However, the nucleic acids are not disrupted, fragmented or otherwise degraded between the fixing device and the nanopore sequencing device.Rather, the nucleic acids, particularly in the form of double-stranded DNA, are purified in a manner typical for nanopore technology, but otherwise fed directly to the nanopore sequencing device. The nanopore sequencing device can, in particular, be a commercially available nanopore sequencing device designed for sequencing DNA without attached proteins or peptides.
[0029] Advantageous further developments of the invention emerge from the patent claims, the description and the drawings.
[0030] The advantages of features and combinations of several features mentioned in the description are merely exemplary and can be effective alternatively or cumulatively, without the advantages necessarily having to be achieved by embodiments according to the invention.
[0031] With regard to the disclosure content – not the scope of protection – of the original application documents and the patent, the following applies: Further features can be found in the drawings – in particular the illustrated geometries and the relative dimensions of several components to one another, as well as their relative arrangement and operative connection. The combination of features of different embodiments of the invention or features of different patent claims is also possible, deviating from the chosen references of the patent claims, and is hereby encouraged. This also applies to features that are illustrated in separate drawings or mentioned in their description. These features can also be combined with features of different patent claims.Likewise, features listed in the patent claims may be omitted for further embodiments of the invention, but this does not apply to the independent claims of the granted patent. The number of features listed in the patent claims and the description are to be understood as meaning that exactly this number or a greater number than the stated number is present, without the need for the explicit use of the adverb "at least." Thus, for example, if one protease is mentioned, this is to be understood as meaning that exactly one protease, two proteases, or more proteases are used. The features listed in the patent claims may be supplemented by further features or may be the only features present in the subject matter of the respective patent claim.
[0032] The reference signs contained in the patent claims do not represent a limitation of the scope of the subject-matter protected by the patent claims. They serve solely to make the patent claims easier to understand.
[0033] BRIEF DESCRIPTION OF THE CHARACTERS
[0034] In the following, the invention is further explained and described with reference to preferred embodiments shown in the figures.
[0035] Fig. 1 is a flowchart of steps of an embodiment of the method according to the invention and at the same time illustrates the structure of a device according to the invention.
[0036] Fig. 2 is a flowchart of analysis steps following the steps according to Fig. 1 of an embodiment of the method according to the invention.
[0037] Fig. 3A and B show quality scores of analyzed DNA libraries as they occur in the method according to the invention, Fig. 3B, right, and in comparison experiments, Fig. 3A and Fig. 3B, left.
[0038] Fig. 4A and B show quality control metrics of Nanopore sequencing data from additional analyzed DNA libraries consisting of two biological replicates, Buffer A and Buffer B, and three technical replicates each.
[0039] Figures 5A and B show sequence information from a proof-of-concept study of the method according to the invention. Figures 6A, B, C, and D show preliminary sequencing results of cross-linked interactions of TBP with double-stranded DNA and corresponding controls in the proof-of-concept study.
[0040] FIGURE DESCRIPTION
[0041] The method 1 according to the invention, shown in Fig. 1 in the form of a flow diagram, serves to determine binding sites of transcription factors 3 to DNA 4. The method 1 according to the invention can also be carried out to determine the binding sites 3 of other proteins 5 to other nucleic acids 6. In a binding reactor 7 of a device 22 for carrying out the method 1, the respective protein 5 is contacted with the nucleic acid 6 in order to enable a binding reaction 8. During this binding reaction, hydrogen bonds typically form between the protein 5 and the nucleic acid 6, which are not sufficiently stable for the determination of the binding sites 2. Therefore, the bonds are fixed in a fixing device 9. For this purpose, short laser pulses 11 with a pulse duration of less than 1 ps and a wavelength in the UV range are directed onto the nucleic acids 6 with the bound proteins 5 using a laser 10.Preferably, the laser pulses 11 have a pulse duration in the femtosecond range, so that the laser 10 is also referred to as a femtosecond laser. The laser pulses 11 cause cross-links 12 between the proteins 5 and the nucleic acids 6. The cross-links are based on the conversion of hydrogen bonds into covalent bonds and have a bond length of zero, meaning they bind the respective protein 5 directly to the respective nucleic acid 6.
[0042] Subsequently, in a sample processing device 13, the proteins 5 are degraded using a protease 14 down to the amino acid 21 directly bound to the respective nucleic acid 6. This is typically followed by purification of the nucleic acids 6 from protein degradation products, for example, using commercial kits from Oxford Nanopore PLC. This is followed by sequencing 16 of the nucleic acids 6 with the bound amino acid in a nanopore sequencing device 17. This produces raw data 23 in the form of a so-called squiggle signal 18.
[0043] According to Fig. 2, the raw data 23 from sequencing 16, i.e., the squiggle signal 18, which is typically in FAST5 or POD5 format, is subjected to analysis 20 in an analysis device 19. For this purpose, the squiggle signal 18 is subjected to so-called basecalling 25, for example, using the basecaller Dorado, an open source program from Oxford Nanopore Technologies. The accuracy of the basecalling 25 depends on the basecalling model used, which is typically trained on unmodified DNA but can be adapted to improve the detection of modifications. The results of the basecalling 25, referred to as reads 26, are subjected to quality control 27. Quality control 27 ensures that only reads 26 with high confidence, which are typically filtered based on Phred quality scores and read lengths, are retained for subsequent squiggle analysis 28.In squiggle analysis 28, for example, deviations that could correspond to altered bases or cross-linking events are identified with the help of Remora, another open source program from Oxford Nanopore Technologies. Basecalling training 29 involves an iterative refinement of the basecalling model used in basecalling 25, incorporating known modifications or experimental artifacts to improve the detection of DNA changes caused by cross-linking. The identified cross-linked DNA bases are then used in downstream analyses not shown separately here to infer specific peptide residues that are bound to these sites. The results of analysis 20 are the binding site of interest, 2 of protein 5 to nucleic acid 6, and the peptide residue 21 bound there.
[0044] In a variant of method 1, the degradation 15 of protein 5 is omitted. Sequencing 16 in the nanopore sequencing device 17 then ends whenever the bound protein 5 hits the entrance of the respective nanopore and prevents further passage of the nucleic acid 6 or the respective strand of nucleic acid 6 through the respective nanopore. If the sequence of nucleic acid 6 is known, the resulting end of the respective squiggle signal 18 provides an indication of the binding site 2 of interest. However, the bound peptide residue 21 cannot be determined in this way.
[0045] The cross-linking 12 in the fixation device 9 with the aid of the short laser pulses 11 is carried out according to the known FLIX technology, for example, in the same way as in the so-called FLIX-MS technology, in which the cross-linking is followed not only by a degradation 15 of the protein 5, but also by a degradation of the nucleic acid 6 and then a purification of the protein residues bound to the nucleic acid residues, as well as a mass spectrometry of these binding products. The sequencing 16 of the nucleic acids 6 in the nanopore sequencing device 17 can be carried out without complex prior preparation using commercially available products, in-house or customer-specific nanopores. However, the analysis 20 of the squiggle signal 18 must be modified compared to these commercial products so that the binding sites 2 of interest of the bound amino acids 21 are determined. For this purpose, the analysis 20 orThe basecalling model used is learned using known combinations of binding sites and bound amino acids as well as potential other DNA changes as a result of the laser pulses 11.
[0046] The laser pulses 11 can cause damage to the nucleic acids 5 in the fixation device 9, such as strand breaks or the formation of pyrimidine dimers. However, this does not impede the sequencing 16 in the nanopore sequencing device 17. Radiation-induced molecular modifications of the DNA, such as pyrimidine dimers and hydrates, can be detected by base calling after the signal has been learned and assigned to the original bases.
[0047] Fig. 3 shows quality scores of analyzed DNA libraries output by a commercial nanopore technology product. Fig. 3A shows the quality score for 200 bp double-stranded DNA fragments without UV irradiation (“(-)UVR”) as a comparison value on the left, and the corresponding quality score after UV irradiation (“(+)UVR”) on the right. Fig. 2B shows the quality scores of 200 bp double-stranded DNA fragments and transcription factors without cross-linking (“(-)DPC”) on the left, and with cross-linking (“(+)DPC”) on the right. The protein hydrolysis and purification steps of DNA fragments following UV irradiation and cross-linking, respectively, were performed in parallel for both sample types. The quality score of Fig. 3B, right, with an average of 9.8, is not significantly worse than the comparison value of Fig. 2A, left, with an average of 11.1.
[0048] Fig. 4 shows quality control metrics for Nanopore sequencing data from additional analyzed DNA libraries consisting of two biological replicates, Buffer A and Buffer B, as well as three technical replicates each. Fig. 4A shows the distribution of read lengths of the sequenced DNA libraries from control samples (dsDNA only) and two independently cross-linked samples (dsDNA+TBP+FLIX in Buffer A and dsDNA+TBP+FLIX in Buffer B). The raw sequencing data were converted to sequence information using a canonical basecalling model, and only reads with a quality score (q-score) of 10 or higher were selected for analysis. Fig. 4B shows the Phred quality score distribution of the analyzed DNA libraries. The table below provides a statistical summary of the number of generated reads according to Phred quality assessments and the average quality scores.
[0049] Number of reads after quality filtering
[0050] — Control 9,301 — Control 404,249 — Control 839,238
[0051] — BufferA 25,722 — BufferA 304,667 — BufferA 467,703
[0052] — - BufferB 13,073 — - BufferB 110,819 — - BufferB 442,953
[0053] Average Quality Scores
[0054] — Control 17.5 — Control 20.5 — ■ Control 20.6
[0055] — BufferA 12.4 — BufferA 12.7 — BufferA 12.8
[0056] ■— BufferB 12.2 — - BufferB 12.7 ■— BufferB 12.9
[0057] While the average quality scores of the control datasets are higher than those of the cross-linked datasets, all datasets meet the standard threshold for good-quality Nanopore sequencing data (q-score >10). The lower quality scores in the cross-linked samples may be due to cross-linking-induced DNA modifications that are not detected by canonical basecalling.
[0058] Fig. 5 shows sequence information for a proof-of-concept study of the method according to the invention, in which the binding behavior of the TATA-box binding protein (TBP) to corresponding double-stranded target sequences was investigated. Fig. 5A shows i) the consensus sequence of the TATA-box motif (Seq_1) in reverse complementary orientation. Listed below are ii) the reference information for mapped cross-linking points (five-pointed empty stars) of TBP on short, double-stranded oligonucleotides (Seq_2) using MS (Reim et al. 2020) and iii) a section of the EF1A promoter region (Seq_3) with the TATA-box motif used in the proof-of-concept study and the expected (five-pointed stars with question marks) cross-linking points of TBP. Fig. 5B shows the 200 bp long dsDNA sequence (Seq_4) used for the in vitro binding studies in combination with TBP.Two reference regions, Region #1 and Region #2, are marked, which were used to analyze the squiggle signals outside the TATA box. Fig. 6 shows preliminary sequencing results of cross-linked interactions of TBP with double-stranded DNA and corresponding controls in the proof-of-concept study. The sequencing data were generated and analyzed using the methods and equipment described in the previous figures. For this analysis, the raw data 23 from control and cross-linked samples were subjected to canonical basecalling 25 using a model trained on unmodified DNA. Fig. 6 compares representative Nanopore signal trajectories of 50 subsampled reads (n=3) between control (dsDNA, no TBP, no cross-linking; “Ctrl”, black) and cross-linked samples relative to the TATA box. Fig. 6A: dsDNA, TBP, FLIX (“FLIX - Buffer A”, gray); Fig. 6C: dsDNS, TBP, FLIX (“FLIX - Buffer B”, gray). Fig. 6B and Fig.6D show randomly selected, size-matched sequence regions of the same 200 bp fragments outside the TBP binding motif. The upper half of each overview plot shows strand-specific signals for reads in the 5'-3' direction, while the lower half represents the signals in the 3'-5' direction. For each strand, the upper plot shows the squiggle signals, while the lower plot shows the corresponding standard deviations to illustrate the variability of the signals. The sequence regions are aligned to the region of interest (Fig. 6A and Fig. 6C: TATABOX; Fig. 6B and Fig. 6D: comparable regions outside the binding motif). Cross-linking analysis at DNA bases of interest is illustrated. The expected cross-linking points (see Fig.3) are marked by diagonal lines; detected changes in the squiggle signal are marked by horizontal lines; where the expected and observed locations overlap, the location is filled with vertical lines.
[0059] Initial analyses using canonical basecalling and squiggle analysis of the TATA box motif show signal deviations in cross-linked samples compared to controls, suggesting altered bases and potential binding sites. While the signal patterns outside the TBP binding motif (Fig. 6B, Fig. 6D) appear comparable between conditions on both strands, clear deviations are observed within the TBP binding motif (Fig. 6A, Fig. 6C) – particularly at cytosine and adenine – as well as at an adenine located upstream (5') of the TATA box on the 5'-3' strand. Outside the TBP binding motif, the signal patterns remain largely homogeneous. However, isolated deviations occur in both the control and cross-linking datasets. In contrast to deviations at the expected cross-linking sites (Fig. 6A, Fig.6C), these do not form plateaus, which would indicate base modifications, and may be due to factors such as repetitive sequences. LIST OF REFERENCE SYMBOLS.
[0060] Proceedings
[0061] Binding place
[0062] Transcription factor
[0063] DNS
[0064] protein
[0065] nucleic acid
[0066] Binding reactor
[0067] Binding reaction
[0068] Fixing device
[0069] Laser
[0070] laser pulse
[0071] Cross-linking
[0072] Sample processing facility
[0073] Protease
[0074] Degradation of the protein
[0075] Sequencing
[0076] Nanopore sequencing device
[0077] Squiggle signal
[0078] Analysis device
[0079] analysis
[0080] Bound peptide
[0081] device
[0082] Raw data
[0083] Computer files
[0084] Basecalling
[0085] Read
[0086] Quality control
[0087] Squiggle analysis
[0088] Basecalling training
Claims
PATENT CLAIMS 1. Method (1) for determining binding sites (2) of proteins (5) to nucleic acids (6), wherein bonds of the proteins (5) to the nucleic acids (6) are fixed by short laser pulses (11) with a pulse duration of not more than 100 ns, characterized in that the nucleic acids (6) with the fixed bonds are sequenced by passing through nanopores.
2. Method (1) according to claim 1, characterized in that the short laser pulses are formed in such a way that hydrogen bonds of the proteins (5) to the nucleic acids (6) are converted into covalent bonds.
3. Method (1) according to claim 1 or 2, characterized in that the proteins (5) are degraded before the sequencing of the nucleic acids (6) to such an extent that the nucleic acids (6) pass through the nanopores via the bonds.
4. Method (1) according to claim 3, characterized in that the proteins (5) are each degraded down to one to a maximum of ten amino acids, preferably down to one to a maximum of three and most preferably down to a single amino acid.
5. Method (1) according to claim 3 or 4, characterized in that the proteins (5) are degraded with the aid of a protease (14).
6. Method (1) according to one of claims 3 to 5, characterized in that squiggle signals (18) arising when the nucleic acids (6) are passed through the nanopores are analyzed in order to determine the amino acid and a nucleobase to which the amino acid is covalently bound and, optionally, DNA damage.
7. Method (1) according to claim 6, characterized in that a basecalling model is trained to recognize the residue and the nucleic base.
8. Method (1) according to claim 1 or 2, characterized in that the nucleic acids (6) with the bound proteins (5) are sequenced until the proteins (6) bind to the nanopores.
9. Device (22) for the automated implementation of the method (1) according to one of the preceding claims, with a fixing device (9) which has a laser (10) emitting short laser pulses (11) with a pulse duration of not more than 100 ns in order to fix bonds of proteins (5) to nucleic acids (6), and with a nanopore sequencing device, wherein the nanopore sequencing device (17) is connected to the fixing device (9) in order to sequence the nucleic acids (6) with the fixed bonds.
10. Device (22) according to claim 9, characterized in that a sample processing device (13) is connected between the fixing device (9) and the nanopore sequencing device (17), which is configured to degrade the bound proteins (5) down to one to a maximum of ten amino acids, preferably down to one to a maximum of three and most preferably down to a single amino acid bound to the respective nucleic acid (6).
Citation Information
Patent Citations
Device for stabilizing the cornea
DE102010020194B4
Method for detecting protein-DNA interaction
EP3404113A1
Method and system for analysis of protein and other modifications on DNA and RNA
WO2014052433A2
Detection and quantification of methylation in DNA
WO2015138405A2
Biomolecule measurement apparatus
WO2017104398A1