DNA motifs for identifying, isolating and assembling the nucleotide sequence of a chromosome centromere
The use of specific DNA motifs with defined patterns and distances addresses the challenge of centromere sequence assembly, enabling precise characterization and disease diagnosis by distinguishing and validating centromere sequences across human chromosomes.
Patent Information
- Application Number
- PCT/IT2024/050269
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-24
- Publication Date
- 2025-07-03
AI Technical Summary
Current methods struggle to accurately identify, isolate, and assemble the nucleotide sequence of chromosome centromeres due to their repetitive and highly polymorphic nature, which complicates genome assembly and hinders the detection of associated diseases.
A method utilizing specific DNA motifs, such as CENP-B box and pJa protein binding motifs, with defined patterns and distances, to distinguish and validate centromere sequences across human chromosomes, enabling precise characterization and assembly of centromeres.
Enables rapid and reliable identification of centromere sequences, facilitating genome assembly, detecting aberrations, and diagnosing diseases related to centromere DNA changes, providing a chromosome-specific roadmap for centromere analysis and diagnosis.
Smart Images

Figure IT2024050269_03072025_PF_FP_ABST
Abstract
Description
[0001] DNA MOTIFS FOR IDENTIFYING, ISOLATING AND ASSEMBLING THE NUCLEOTIDE SEQUENCE OF A CHROMOSOME CENTROMERE
[0002] The present invention concerns DNA motifs for identifying, isolating and assembling the nucleotide sequence of a chromosome centromere. In particular, these motifs are designed for the identification, isolation, and assembly of the nucleotide sequence of a chromosome's centromere and the entire human genome. Additionally, the invention encompasses methods for validating these sequences. The primary application of these DNA motifs and methods lies in their ability to facilitate the detailed characterization and analysis of centromeres and whole chromosomes. This innovation stands to make significant contributions to the fields of genomic research and diagnostic genetics, offering a novel approach to centromere study and genome assembly.
[0003] Human centromeres are regions of repetitive alpha-satellite DNA present in each chromosome (Wevrick and Willard, 1989) to support faithful segregation of the genetic information. Advances in sequencing enabled the first complete linear annotation of human centromeres for all chromosomes for the pseudo-haploid human cell line (CHM-13 hTERT) derived from embryonic tissue of an inviable molar pregnancy (Altemose et al., 2022; Bzikadze and Pevzner, 2020; Gershman et al., 2022; Nurk et al., 2022). More recently, complete human diploid genomes have been assembled, including the HG002 fully phased maternal and paternal haplotypes using trio-binning (Jarvis et al., 2022) and the reference genome for the RPE1 cell line with phased haplotypes (Volpe et al., 2023). The latest development has been to build a pangenome map of near complete human diploid assemblies (Liao et al., 2023). These studies highlight the genetic divergence between individuals. It is emblematic that centromeres and rDNA remain the hardest regions to assemble and linearly annotate of the human genome. In fact, centromere sequences are absent from the current human pangenome draft (Liao et al., 2023). This is due to their genetic structure made of long repetitive tandem arrays repeated near identical in sequence. Human pericentromeric and centromeric satellite DNA occupies an estimate of 5-8% of the human genome. Precisely, for T2T-CHM13 genome centromeres represent 6.2% of the whole genome, equal to 189.9 Mb as reported (Altemose et al., 2022). The centromere contains tandemly repeated alphasatellite (85.2 Mb total genome-wide in CHM13) that individually occupy ~170 nucleotides (referred to as ‘nt’ or ‘base pairs’ or ‘bp’, all to mean single DNA bases) of repeat units called “monomer” (Altemose et al., 2022). Human monomers are further organized in chromosome specific higher-order repeats (HORs) which are multi-monomeric repeat units that share high sequence homology and therefore a very low divergence in sequence (average 1 .5%). HOR-array can be classified into active, inactive and divergent. The active centromere region is involved in ensuring a faithful chromosome segregation due to its association with the kinetochore. This region can vary in size spanning between 340 kb on chromosome 21 to 4.8 Mb on chromosome 18, depending on the chromosome (Altemose et al., 2022). The flanking inactive region includes alpha-satellite monomers that lack HORs-unit composition and become progressively smaller and more diverged arrays. It can also contain transposable elements (TEs), segmental duplications, and non-alpha- satellite repeat families. Moreover, there are divergent HOR-arrays adjacent to the more homogeneous HOR arrays in which HOR periodicity is almost completely eroded as well as the presence of highly divergent alpha-satellite monomeric layers and other repetitive elements. The active HORs host a region where the canonical nucleosome is interspersed with nucleosomes containing a centromere-specific histone H3 variant called CENP-A (Blower et al., 2002; Bodor et al., 2014; Gershman et al., 2022; Glynn et al., 2010). Recently, the T2T Consortium has highlighted the centromere-dip-regions (CDRs) within the human centromere which are hypomethylated DNA arrays to which the CENP-A protein is closely associated (Altemose et al., 2022; Gershman et al., 2022). Within the CDRs, CENP-A serves as the site for kinetochore attachment for chromosomes to divide. CENP-B is another centromeric protein, which maintains centromere organization and structure binding to a 17 base pairs (referred to as ‘bp’) motif within the alpha-satellite DNA (Masumoto et al., 1989; Miga, 2015). Recognition of the binding motif is essential for the positioning of the protein within alpha-satellite repeats. Moreover, in vitro experiments showed that only 9 bp of the CENP-B box are essential for CENP-B binding to DNA (Masumoto et al., 1993).
[0004] Due to the challenge of describing and correctly assembling a centromere of a chromosome, experts in human centromere genomics have harnessed the identification of SLINKs (single unique nucleotide kmers), unique DNA motif kmers within the genome of interest that can help to guide the assembly, structure and order to correctly place repetitive elements within the centromere of a specific chromosome (Logsdon et al., 2021 ).
[0005] However, because centromere DNA sequence changes between individuals, it’s a painstaking exercise for each human genome to identify and reliably validate the distribution, amount and location for each sequencing read to correctly build a complete linear assembly based on SLINKs. Because centromere DNA is some of the most rapidly evolving across species, the essential function carried out by this locus in enabling chromosome segregation is epigenetically determined by CENP- A. Very recently, evidence of conservation amongst human and primate species of chromatin organization has reinforced the long-standing notion that chromatin, but not DNA sequence, determines functional conservation of centromeres (Dubocanin et al., 2023). Importantly, however, while CENP-A positioning does not seem to be determined by sequences, CENP-B is the only sequence-specific centromere protein localized to centromeres of all chromosomes except chromosome Y. CENP- B binds to a specific DNA motif and it’s thought to localize mainly within the active centromere. Recent evidence suggest that CENP-B positioning is functionally relevant to build the correct chromatin structure (Chardon et al., 2022), implying that the positioning of CENP-B boxes may be the only feature within an extremely polymorphic centromere domain to be conserved between individuals in human and potentially other closely related species given its functionality.
[0006] Notwithstanding the knowledge concerning centromeres, they are some of the most repetitive and hard to resolve regions of any genome. In spite of their essential role in enabling chromosome segregation, and recent association of centromere DNA instability with human diseases (Black and Giunta, 2018; Giunta et al., 2021 ), rapid and reliable approaches to interrogate centromeric loci are lacking. The resolution of human centromere DNA sequences has always been a major challenge for sequencing technologies and assembly algorithms. Also, methods that can identify changes within centromere DNA, including those associated with pathological states, are currently lacking.
[0007] In the light of the above it is evident the need for new methods and tools for analyzing, identifying, isolating, assembling, validating the nucleotide sequence of a centromere of a chromosome and diagnosing diseases related to aberrant centromere DNA that overcome the disadvantages of known methods and tools.
[0008] With this background, the present invention introduces a method whereby the identification of a centromere directly leads to the determination of its corresponding chromosome. This method employs specific motifs which have significant utility in various diagnostic and genetic methodologies. These motifs are particularly effective for analyzing the nucleotide sequence of a chromosome's centromere. Such analysis is crucial for identifying aberrations, variants, mutations, or changes within the centromere, including those associated with human diseases and pathological mutagenesis. Furthermore, the invention provides a means to assess the correct order, organization, and size of a centromere's nucleotide sequence for each human chromosome. This aspect is essential for tasks such as karyotyping, detecting the absence or presence of individual chromosomes, and aiding or validating the assembly of the human genome and its completeness. One of the novel features of this invention is the ability to determine the size of the functional centromere nucleotide sequence. This determination is based on the maximum number of repetitions of the identified motifs, offering a precise and effective approach to centromere analysis. Overall, the present invention represents a significant advancement in the field of genomic research, providing a reliable and efficient tool for detailed centromere study, crucial for both diagnostic and research applications.
[0009] The present invention reveals for the first time an ancestral footprint associated with the positioning of functional centromeric DNA motifs along the chromosome present in conserved positions at the same distances across the human population. The present invention identifies the chromosome-specific distribution of distances between a specific centromeric DNA sequence motif. The present invention applies to detect repetitive regions of the genome, which traditionally are the most complex and hard to resolve regions of any genome and which are often associated with human diseases.
[0010] To reliably identify changes with centromeric DNA, the Applicant established a genome-wide mapping using a set of computational methods collectively referred to as the Genomic-Centromere Finder (GCF) and determined the conserved DNA fingerprinting characteristic of each human chromosome. The present invention therefore identifies for the first time a chromosome-specific centromere pattern derived by the above-mentioned sequences distances through computational methods. Founded on evidence, the DNA motif-based centromere fingerprinting is conserved across human genomes and extends to non-human primate species, in spite of the extensive underlying polymorphism in the DNA sequence implying its functional relevance in centromere function and genome architecture. According to the present invention, evidence is provided for the use of the centromere fingerprinting based on specific motifs and their distance, genomic coordinates and orientation as an effective genetic handle to (1 ) distinguish and isolate chromosomespecific centromeric DNAfrom sequencing reads, (2) validate and analyze complete centromeres during genome reconstruction and (3) predicts the size of the functional centromere chromatin for kinetochore assembly.
[0011] Altogether, according to the present invention, a much-needed rapid and reliable tool is established for the interrogation of centromere DNA sequences, including detection of pathological changes associated with human diseases. According to the present invention, the set of computational tools, collectively referred as Genomic-Centromere Finder (GCF), is provided to build a genome-wide map of centromere motifs, extrapolate the position thereof and calculate the distance between the motifs, in particular between the DNA motif of CENP-B boxes. The distance between each CENP-B box constructs a specific barcode for each chromosome that allows to distinguish them from each other based on the CENP-B box distribution. The precise distribution of these sequences defines the organization of the centromere array of each chromosome, allowing targeted analysis of the centromeric regions of the target chromosomes and improving the ability to isolate the centromeres of specific chromosomes by exploiting the distribution of this sequence identified by the ‘centromere fingerprint’. With the advantage of long reads sequencing, the analysis according to the present invention can drastically improve and facilitate the assembly, validation, study and characterization of centromere DNA in research, clinical settings and beyond.
[0012] According to the present invention, it is now possible to correctly assess, validate and analyze a complete haploid and diploid human reference genome. In the analysis, the Applicant mined the first fully assembled human centromere sequence to reveal pattern and distribution of the CENP-B box sequence. Based on previous data on distances between the CENP-B boxes, a periodicity of approximately one CENP-B box every other monomer was expected following the scheme described by Rosandic et al. in 2006 (Rosandic et al., 2006). Moreover, Henikoff et al. in 2015 investigated the cadence of the CENP-B box in the human alpha-satellite HOR arrays describing a consensus distribution every ~340 base pairs (distance plus CENP-B box length) (Henikoff et al., 2015). According to the present invention, the analysis highlighted the centromere identity of each chromosome based on the consensus distribution of the CENP-B box sequence identifying many additional distances specific to the centromere of a chromosome and conserved across individuals. The main pattern genome-wide remains represented by the cadency of this sequence every other monomer that is enriched in the first and second cluster. Conversely, according to the present invention, it was found that the third cluster started to lose the main pattern, achieving a new consensus distribution of one CENP-B box every monomer or one CENP-B box every three monomers. In addition, the fourth cluster represents a special case in which chromosome 17 and chromosome X construct unique pattern consisting of several different distance values and almost lose the pattern every other monomer (e.g. distance values between 320-325). Altogether, the present invention provides a chromosome specific roadmap of CENP-B box organization at base pair resolution and informs on the sequence, organization, structure, chromatin landscape, stiffness, and propensity to mis-segregate of specific chromosomes. The present invention is extremely important in diagnostic fields for laboratories where repetitive regions of specific chromosomes are highly studied. On the other hand, using pJa binding motif as a second control sequence, the Applicant also discovered a different cadence compared to the CENP-B box but also that this sequence is lost in almost all the active HORs, except those in chromosome 2, 4, 8, 13, 15, 20, 21. This specific positioning of pJa binding motif in the active HORs has been harnessed by the Inventors to further advance centromere characterization of a genome. The lateral position of the pJa binding motif may reflect the theory of layered expansion playing at the centromere in which has been demonstrated that new alpha-satellite repeats periodically emerge and expand within an active array, displacing the older repeats laterally, and becoming the site of kinetochore assembly (Altemose et al., 2022). Taking advantage of long sequences, any laboratories can sequence any DNA specimen, cell line or patient sample and extrapolate information on the centromere by simply visualizing whether the pattern of the two motifs CENP-B boxes and pJa binding motif, alone or in combination, is maintained and whether this region is physiologically present or lost or aberrantly altered under different conditions. Additionally, this precise organization allows discriminating sequences belonging to specific chromosomes from a pool of raw sequences from any DNA sample of any subject or cell line. Furthermore, the sequence analysis according to the present invention demonstrates its use to further attempt to call variants in the human centromeres, assemble the diploid genome of different individuals or cell lines and help the resolution of highly repetitive DNA sequences using these ‘anchor’ sequences whom the Inventors have established positioned with chromosome-specific patterns. Additionally, the Inventors identified that not just the specific centromere DNA motif sequences and their specific distances, but also the orientation (forward or reverse complement) of the motifs, are used to identify the centromere and chromosome of a human genome.
[0013] Based on the above, the present invention demonstrates for the first time that the centromere barcode for each individual chromosome is advantageously exploitable to detect, analyze, assemble, validate, diagnose DNA changes, and predict diseases related to the human centromere and whole chromosomes arms. Thanks to the finding that specific centromere DNA motifs tend to occur at certain distances, frequencies and orientation that are specific to each human chromosome, the invention can be used to distinguish centromeres for each human chromosome solely looking at DNA sequencing reads, without the need to fully reconstruct the completely assembled sequence, and thus to improve and facilitate the interrogation of changes in centromere DNA region and in the overall genome. Therefore, the present invention provides a new, fast and simpler method for identifying and characterizing the repetitive sequences of human centromeres. In addition, the centromere barcode according to the present invention can be used as biomarker to assess genomic changes in a variety of diagnostic, clinical and therapeutic research interventions, but can also be used to validate and facilitate human genome assemblies providing the possibility of validating the fidelity of the centromere sequences and isolating chromosome-specific centromeric sequences from raw sequences (fastq) of any human subject or cell line for the study of the chromosome of interest. Altogether, the present invention drastically improves all- around analyses of centromere and human DNA in clinical settings and genetic diagnostics, as well as other fields of research.
[0014] Therefore, the present patent application seeks protection for a novel method and associated computational tools designed for analyzing chromosome-specific distances and orientation of a specific DNA motif. The method has several applications:
[0015] - Genetic and Genomic Codification of Human Centromeres: Using the discovery of the most frequently occurring distance of a DNA motif to define the structure of human centromeres.
[0016] - Extraction of Long DNA Sequencing Reads from Centromeres: Facilitating the retrieval of extended DNA sequences from centromeric regions.
[0017] - Chromosomal Origin Determination of DNA Reads: Identifying the specific chromosome from which DNA reads originate, applicable across various human specimens, cell lines, and individual DNA samples.
[0018] - Validation of Centromere Sequences and Human Genome Assemblies: Offering a method to direct and confirm the accuracy of centromere and chromosome sequences within the context of complete human genome assemblies.
[0019] - Guidance for Correct Assembly of Human Centromeres and Genomes: Providing a directional approach to accurately assemble human centromeres and genomes.
[0020] - Identification of Aberrations and Structural Variations in Centromere DNA: Detecting structural variations, insertions / deletions (INDELS), and other alterations in centromere DNA, including those of pathological origin that are not discernible using current sequencing and diagnostic methods.
[0021] - Prediction of Active Centromere Size: Estimating the active and functional size of centromeres for each human chromosome, utilizing an automated process based on the modality of distances identified by our method.
[0022] - Rapid Identification of Genome Assembly Errors: Swiftly flagging inaccurately assembled regions of the human genome, including pathologically mutated areas and other aberrations at the DNA sequence level.
[0023] Therefore, it is a specific object of the present invention the use of a DNA motif as a marker for characterizing, identifying, isolating, assembling or validating the nucleotide sequence of the centromere of a human chromosome, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or its reverse complement sequence *CCCG**T****CGAA*; or a pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or its reverse complement sequence CTA[C or T]GGTG[A or G]AAAAGGAA (SEQ ID NO:2); said motif being repeated several times along the nucleotide sequence of the centromere according to a pattern of distances between subsequent motifs, said pattern being repeated several times along the nucleotide sequence of the centromere and being specific for each chromosome, wherein the distance is the number of nucleotides between the last nucleotide of a motif and the first nucleotide of the subsequent motif.
[0024] According to the present invention, the motif *TTCG****A**CGGG* can be the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4); whereas the reverse motif *CCCG**T****CGAA* can be the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO:6).
[0025] Therefore, the pattern of distances comprises one or more motifs that are repeated with specific distances, said distances being one after the other consecutively. Every chromosome presents a specific pattern of distances in the centromere nucleotide sequence, so that the pattern of distances of the centromere nucleotide sequence is a distinctive barcode of a chromosome. As mentioned above, in turn the pattern of distances is repeated along the nucleotide sequence of the centromere. The pattern of distances can be repeated (several times) along the nucleotide sequence of the centromere one after the other consecutively, i.e. one attached to the next, in particular in the active (or functional) nucleotide sequence of the centromere. The pattern of distances in the centromere nucleotide sequence of each chromosome is described below. In particular, the patterns described below refer to the patterns of CENP-B box protein binding motif, that can comprise *TTCG****A**CGGG* or *CCCG**T****CGAA* or both *TTCG****A**CGGG* and *CCCG**T****CGAA* in the same pattern, and to the patterns of pJa protein binding motif, that can comprise SEQ ID NO: 1 or 2 or both SEQ ID NO: 1 and 2 in the same pattern. The distances represent the most represented distances between subsequent motifs, however each distance value could vary by + / - 1 nucleotide. When the pattern of distances comprises more than one distance among the motifs, the pattern of distances is described below with the distance values separated by a comma or by a hyphen, each distance being the number of nucleotides between the last nucleotide of a motif and the first nucleotide of the subsequent motif. It is clear to a person skilled in the art that, when the pattern of distances comprises different distances, since the patterns can be repeated consecutively, the same pattern can be represented by starting the series of consecutive distance values of the pattern from any of the distance values indicated below for each pattern of distances; in this case the distance values of the pattern that are upstream of the chosen initial distance value have to be considered downstream of the last distance value of the pattern as described below. When the pattern comprises only one distance value, a minimum number of consecutive repetitions of this distance value in the centromere nucleotide sequence is reported in order to provide a distinctive pattern of distance for the specific chromosome.
[0026] According to the use of the present invention the pattern of distances between subsequent CENP-B box protein binding motifs within the centromere can be:
[0027] - chosen among (322 ± 1 )nnucleotides, (323 ± 1 )nnucleotides or (324 ± 1 )nnucleotides, preferably (323 ± 1 )nnucleotides, more preferably (323)nnucleotides, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 5, for example at least 6, at least 7, at least 8, at least 9 or at least 10, in the nucleotide sequence of the centromere of chromosome 1 (i.e. the three distance values of 322 nucleotides, 323 nucleotides or 324 nucleotides are the three most represented distances in chromosome 1 . In particular, the most represented distance value is 323 and is repeated in tandem so that the unit of distancepattern is 323n, wherein n is greater than or equal to 5. Each distance value can vary by + / - 1 nucleotide; for example the distance value 322 ± 1 nucleotides means 321 or 322 or 323. The term “in tandem” means that the distance value in the pattern is repeated one after the other consecutively);
[0028] - 321 ± 1 , 325 ± 1 nucleotides in the nucleotide sequence of the centromere of chromosome 2 (i.e. the pattern consists of CENP-B box protein binding motif repeated with these main different alternated distances, 321 and 325 nucleotides. In other words, the pattern is 321-325, that is repeated along the nucleotide sequence of the centromere for example as follows: ...321 -325- 321 -325-321 -325... , wherein each distance can vary by + / - 1 nucleotide. It is clear to a person skilled in the art that the pattern can be read as 321 ± 1 - 325 ± 1 or 325 ± 1 - 321 ± 1 , i.e. the pattern can start on any value mentioned in the pattern string. Said pattern can be described also as (321 ± 1 , 325 ± 1 )n, wherein n is the number of repetition of said distance values in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively, or more times);
[0029] - 497 ± 1 , 321 ± 1 , 324 ± 1 , 324 ± 1 , 322 ± 1 , 321 ± 1 and 663 ± 1 in the nucleotide sequence of the centromere of chromosome 3 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motif is repeated with these main different nucleotide distances organized in tandem as follows: ...497 — 321 - 324-324-322-321 -663... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 321 ± 1 , 324 ± 1 , 324 ± 1 , 322 ± 1 , 321 ± 1 , 663 ± 1 and 497 ± 1 ; 324 ± 1 , 324 ± 1 , 322 ± 1 , 321 ± 1 , 663 ± 1 , 497 ± 1 and 321 ± 1 ; 324 ± 1 , 322 ± 1 , 321 ± 1 , 663 ± 1 , 497 ± 1 , 321 ± 1 and 324 ± 1 ; 322 ± 1 , 321 ± 1 ,
[0030] 663 ± 1 , 497 ± 1 , 321 ± 1 , 324 ± 1 and 324 ± 1 ; 321 ± 1 , 663 ± 1 , 497 ± 1 ,
[0031] 321 ± 1 , 324 ± 1 , 324 ± 1 and 322 ± 1 ; 663 ± 1 , 497 ± 1 , 321 ± 1 , 324 ± 1 ,
[0032] 324 ± 1 , 322 ± 1 and 321 ± 1 ; 497 ± 1 , 321 ± 1 , 324 ± 1 , 324 ± 1 , 322 ± 1 ,
[0033] 321 ± 1 and 663 ± 1 );
[0034] - 664 ± 1 , 153 ± 1 , 323 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 491 ± 1 and 324 ± 1 in the nucleotide sequence of the centromere of chromosome 4 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...664- 153-323-321 -325-495-491 -324... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 153 ± 1 , 323 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 491 ± 1 , 324 ± 1 and 664 ± 1 ; 323 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 491 ± 1 , 324 ± 1 , 664 ± 1 and 153 ± 1 ; 321 ± 1 , 325 ± 1 , 495 ± 1 , 491 ± 1 , 324 ± 1 , 664 ± 1 , 153 ± 1 and 323 ± 1 ; 325 ± 1 , 495 ± 1 , 491 ± 1 , 324 ± 1 , 664 ± 1 , 153 ± 1 , 323 ± 1 and 321 ± 1 ; 495 ± 1 , 491 ± 1 , 324 ± 1 , 664 ± 1 , 153 ± 1 , 323 ± 1 ,
[0035] 321 ± 1 and 325 ± 1 ; 491 ± 1 , 324 ± 1 , 664 ± 1 , 153 ± 1 , 323 ± 1 , 321 ± 1 ,
[0036] 325 ± 1 and 495 ± 1 ; 324 ± 1 , 664 ± 1 , 153 ± 1 , 323 ± 1 , 321 ± 1 , 325 ± 1 ,
[0037] 495 ± 1 and 491 ± 1 );
[0038] - 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 in the nucleotide sequence of the centromere of chromosome 5 (i.e. this is the unit of distancepattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...323-323-323-323- 323-322... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 322 ± 1 and 323 ± 1 ; 323 ± 1 , 323 ±
[0039] 1 , 323 ± 1 , 322 ± 1 , 323 ± 1 and 323 ± 1 ; 323 ± 1 , 323 ± 1 , 322 ± 1 , 323 ± 1 ,
[0040] 323 ± 1 and 323 ± 1 ; 323 ± 1 , 322 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 and 323 ±
[0041] 1 ; 322 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 and 323 ± 1 );
[0042] - 665 ± 1 , 323 ± 1 , 323 ± 1 , 320 ± 1 , 153 ± 1 , 323 ± 1 and 832 ± 1 in the nucleotide sequence of the centromere of chromosome 6 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...665-323- 323-320-152-323-832... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 323 ± 1 , 320 ± 1 , 153 ± 1 , 323 ± 1 , 832 ± 1 and 665 ± 1 ; 323 ± 1 , 320 ± 1 , 153 ± 1 , 323 ± 1 , 832 ± 1 , 665 ± 1 and 323 ± 1 ; 320 ± 1 , 153 ± 1 , 323 ± 1 , 832 ± 1 , 665 ± 1 , 323 ± 1 and 323 ± 1 ; 153 ± 1 , 323 ± 1 , 832 ± 1 , 665 ± 1 , 323 ± 1 , 323 ± 1 and 320 ± 1 ; 323 ± 1 , 832 ± 1 , 665 ± 1 , 323 ± 1 , 323 ± 1 , 320 ± 1 and 153 ± 1 ; 832 ± 1 , 665 ± 1 , 323 ± 1 , 323 ± 1 , 320 ± 1 , 153 ± 1 and 323 ± 1 );
[0043] - 325 ± 1 , 323 ± 1 , 323 ± 1 and 325 ± 1 in the nucleotide sequence of the centromere of chromosome 7 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...325-323-323-325... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 323 ± 1 , 325 ± 1 and 325 ± 1 ; 323 ± 1 , 325 ± 1 , 325 ± 1 and 323 ± 1 ; 325 ± 1 , 325 ± 1 , 323 ± 1 and 323 ± 1 );
[0044] - 321 ± 1 , 154 ± 1 and 667 ± 1 in the nucleotide sequence of the centromere of chromosome 8 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -154-667... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 154 ± 1 , 667 ± 1 and 321 ± 1 ; 667 ± 1 , 321 ± 1 and 321 ± 1 );
[0045] - 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 and 495 ± 1 in the nucleotide sequence of the centromere of chromosome 9 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -325- 321 -325-321 -325-495... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 and 321 ± 1 ; 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 321 ± 1 and 325 ± 1 ; 325 ±
[0046] 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 321 ± 1 , 325 ± 1 and 321 ± 1 ; 321 ± 1 , 325 ± 1 ,
[0047] 495 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 and 325 ± 1 ; 325 ± 1 , 495 ± 1 , 321 ± 1 ,
[0048] 325 ± 1 , 321 ± 1 , 325 ± 1 and 321 ± 1 ; 495 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 ,
[0049] 325 ± 1 , 321 ± 1 and 325 ± 1 ); - 323 ± 1 , 323 ± 1 , 322 ± 1 and 321 ± 1 in the nucleotide sequence of the centromere of chromosome 10 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with 3 different nucleotide distances organized in tandem as follows: ...323-323-322-323-323-322-321 ... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 322 ± 1 , 321 ± 1 and 323 ± 1 ; 322 ± 1 , 321 ± 1 , 323 ± 1 and 323 ± 1 ; 321 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 );
[0050] - 321 ± 1 , 495 ± 1 in the nucleotide sequence of the centromere of chromosome 11 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -495-321 -495... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read also as: 495 ± 1 , 321 ± 1. Said pattern can be also described as ( 321 ± 1 , 495 ± 1 )n, wherein n is the number of repetition of said distance values in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively);
[0051] - 323 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 in the nucleotide sequence of the centromere of chromosome 12 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...323-323-323-322... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 323 ± 1 , 322 ± 1 and 323 ± 1 ; 323 ± 1 , 322 ± 1 , 323 and 323 ± 1 ; 322 ± 1 , 323, 323 ± 1 and 323 ± 1 );
[0052] - chosen between a first pattern 323 ± 1 , 324 ± 1 and 492 ± 1 and a second pattern 325 ± 1 , 321 ± 1 , 323 ± 1 , 324 ± 1 and 492 ± 1 in the nucleotide sequence of the centromere of chromosome 13 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances that can be organized in tandem as follows with 2 options: ...323-324-492... and ...325-321-323-324-492... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 324 ± 1 , 492 ± 1 and 323 ± 1 ; 492 ± 1 , 323 ± 1 and 324 ± 1 the first pattern, and 321 ± 1 , 323 ± 1 , 324 ± 1 , 492 ± 1 and 325 ± 1 ; 323 ± 1 , 324 ± 1 , 492 ± 1 , 325 ± 1 and 321 ± 1 ; 324 ± 1 , 492 ± 1 , 325 ± 1 , 321 ± 1 and 323 ± 1 ; 492 ± 1 , 325 ± 1 , 321 ± 1 , 323 ± 1 and 324 ± 1 the second pattern);
[0053] - 321 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 and 321 ± 1 n in the nucleotide sequence of the centromere of chromosome 14 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -325-325-325-321 ... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 325 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 and 321 ± 1 ; 325 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 and 325 ± 1 ; 325 ± 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 and 325 ± 1 ; 321 ± 1 , 321 ± 1 , 325 ± 1 , 325 ± 1 and 325 ± 1 );
[0054] - 154 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 and 325 ± 1 in the nucleotide sequence of the centromere of chromosome 15 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...154- 325-325-325-321 -325-321 -325... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 325 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 and 154 ± 1 ; 325 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 154 ± 1 and 325 ± 1 ; 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 154 ± 1 , 325 ± 1 and 325 ± 1 ; 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 154 ± 1 , 325 ± 1 ,
[0055] 325 ± 1 and 325 ± 1 ; 325 ± 1 , 321 ± 1 , 325 ± 1 , 154 ± 1 , 325 ± 1 , 325 ± 1 ,
[0056] 325 ± 1 and 321 ± 1 ; 321 ± 1 , 325 ± 1 , 154 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 ,
[0057] 321 ± 1 and 325 ± 1 ; 325 ± 1 , 154 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 ,
[0058] 325 ± 1 and 321 ± 1 );
[0059] - 322 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 in the nucleotide sequence of the centromere of chromosome 16 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...322-323-323-323-323... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 ; 323 ± 1 , 323 ± 1 , 323 ± 1 , 322 ± 1 and 323 ± 1 ; 323 ± 1 , 323 ± 1 , 322 ± 1 , 323 ± 1 and 323 ± 1 ; 323 ± 1 , 322 ± 1 , 323 ± 1 , 323 ± 1 and 323 ± 1 );
[0060] - 497 ± 1 , 496 ± 1 , 150 ± 1 , 151 ± 1 , 320 ± 1 , 149 ± 1 , 667 ± 1 and 149 ± 1 in the nucleotide sequence of the centromere of chromosome 17 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...497- 496-150-151 -320-149-667-149... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 496 ± 1 , 150 ± 1 , 151 ± 1 , 320 ± 1 , 149 ± 1 , 667 ± 1 , 149 ± 1 and 497 ± 1 ; 150 ± 1 , 151 ± 1 , 320 ± 1 , 149 ± 1 , 667 ± 1 , 149 ± 1 , 497 ± 1 and 496 ± 1 ; 151 ± 1 , 320 ± 1 , 149 ± 1 , 667 ± 1 , 149 ± 1 , 497 ± 1 , 496 ± 1 and 150 ± 1 ; 320 ± 1 , 149 ± 1 , 667 ± 1 , 149 ± 1 , 497 ± 1 , 496 ± 1 ,
[0061] 150 ± 1 and 151 ± 1 ; 149 ± 1 , 667 ± 1 , 149 ± 1 , 497 ± 1 , 496 ± 1 , 150 ± 1 ,
[0062] 151 ± 1 and 320 ± 1 ; 667 ± 1 , 149 ± 1 , 497 ± 1 , 496 ± 1 , 150 ± 1 , 151 ± 1 ,
[0063] 320 ± 1 and 149 ± 1 ; 149 ± 1 , 497 ± 1 , 496 ± 1 , 150 ± 1 , 151 ± 1 , 320 ± 1 ,
[0064] 149 ± 1 and 667 ± 1 );
[0065] - 321 ± 1 , 320 ± 1 , 321 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 in the nucleotide sequence of the centromere of chromosome 18 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -320-321 -325-325-321 ... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start considering a sliding window of 1 over each distance, specifically the pattern can be read as: 320 ± 1 , 321 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 and 321 ± 1 ; 321 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 and 320 ± 1 ; 325 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 , 320 ± 1 and 321 ± 1 ; 325 ± 1 , 321 ± 1 , 321 ± 1 , 320 ± 1 , 321 ± 1 and 325 ± 1 ; 321 ± 1 , 321 ± 1 , 320 ± 1 , 321 ± 1 , 325 ± 1 and 325 ± 1 );
[0066] - chosen among (323 ± 1 )nnucleotides, (322 ± 1 )nnucleotides or (663 ± 1 )nnucleotides, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 5, for example at least 6, at least 7, at least 8, at least 9 or at least 10, in the nucleotide sequence of the centromere of chromosome 19 (i.e. the three distance values of 323 ± 1 nucleotides, 322 ± 1 nucleotides or 663 ± 1 nucleotides are the three most represented distances in chromosome 19). For example CENP-B box protein binding motifs are repeated in tandem as follows: ...323-323-323-323-323- 323-323-323-323-323-323-323-323-323-323... wherein each distance can vary by + / - 1 nucleotide);
[0067] - 663 ± 1 , 324 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 in the nucleotide sequence of the centromere of chromosome 20 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence). In other words, CENP-B box protein binding motifs are repeated with these-main different nucleotide distances organized in tandem as follows: ...663-324- 321 -325-321 -321 -325... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 324 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 and 663 ± 1 ; 321 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 , 663 ± 1 and 324 ± 1 ; 325 ±
[0068] 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 , 663 ± 1 , 324 ± 1 and 321 ± 1 ; 321 ± 1 , 321 ± 1 ,
[0069] 325 ± 1 , 663 ± 1 , 324 ± 1 , 321 ± 1 and 325 ± 1 ; 321 ± 1 , 325 ± 1 , 663 ± 1 ,
[0070] 324 ± 1 , 321 ± 1 , 325 ± 1 and 321 ± 1 ; 325 ± 1 , 663 ± 1 , 324 ± 1 , 321 ± 1 ,
[0071] 325 ± 1 , 321 ± 1 and 321 ± 1 );
[0072] - 492 ± 1 , 325 ± 1 , 321 ± 1 , 323 ± 1 , 324 ± 1 in the nucleotide sequence of the centromere of chromosome 21 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence). In other words, CENP-B box protein binding motifs are repeated with different nucleotide distances organized in tandem as follows: ...492-325-321 -323-324... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 325 ± 1 , 321 ± 1 , 323 ± 1 , 324 ± 1 and 492 ± 1 ; 321 ± 1 , 323 ± 1 , 324 ± 1 , 492 ± 1 and 325 ± 1 ; 323 ± 1 , 324 ± 1 , 492 ± 1 , 325 ± 1 and 321 ± 1 ; 324 ± 1 , 492 ± 1 , 325 ± 1 , 321 ± 1 and 323 ± 1 );
[0073] - 321 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 in the nucleotide sequence of the centromere of chromosome 22 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -325-325-325... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 325 ± 1 , 325 ± 1 , 325 ± 1 and 321 ± 1 ; 325 ± 1 , 325 ± 1 , 321 ± 1 and 325 ± 1 ; 325 ± 1 , 321 ± 1 , 325 ± 1 and 325 ± 1 ); - chosen between 834 ± 1 , 494 ± 1 , 511 ± 1 , 150 ± 1 or 834 ± 1 , 150 ± 1 , 511 ± 1 , 494 ± 1 in the nucleotide sequence of the centromere of chromosome X (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated according to these main different distance pattern in which the nucleotide distances are organized in tandem as follows: ...834-494-511 -150... and ...834-150-511 -494... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 494 ± 1 , 511 ± 1 , 150 ± 1 and 834 ± 1 ; 511 ± 1 , 150 ± 1 , 834 ± 1 and 494 ± 1 ; 150 ± 1 , 834 ± 1 , 494 ± 1 and 511 ± 1 the first pattern and 150 ± 1 , 511 ± 1 , 494 ± 1 and 834 ± 1 ; 511 ± 1 , 494 ± 1 , 834 ± 1 and 150 ± 1 ; 494 ± 1 , 834 ± 1 , 150 ± 1 and 511 ± 1 the second pattern).
[0074] According to the use of the present invention, the pattern of distances between subsequent pJ-alpha (pJa) DNA motifs can be:
[0075] - 321 ± 1 and 325 ± 1 nucleotides in the nucleotide sequence of the centromere of chromosome 2 (i.e. the pattern consists of pJa motif repeated with these main different alternated distances, 321 and 325 nucleotides. In other words, the pattern is 321 -325 (as described even for the CENP-B box in this chromosome), that is repeated along the nucleotide sequence of the peri / centromere of chromosome 2: ...321 -325-321 -325-321 -325... , wherein each distance can vary by + / - 1 nucleotide. It is clear to a person skilled in the art that the pattern can be read as 321 ± 1 - 325 ± 1 or 325 ± 1 - 321 ± 1 , i.e. the pattern can start on any distance value mentioned in the pattern string according to a sliding window of 1 value over each distance. Said pattern can be described also as (321 ± 1 , 325 ± 1 )n, wherein n is the number of repetition of said distance values in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively);
[0076] - 832 ± 1 , 664 ± 1 , 1173 ± 1 , 495 ± 1 in the nucleotide sequence of the centromere of chromosome 4 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated with these main different nucleotide distances organized in tandem as follows: ...832-664-1173-495... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 664 ± 1 , 1173 ± 1 , 495 ± 1 and 832 ± 1 ; 1173 ± 1 , 495 ± 1 , 832 ± 1 and 664 ± 1 ; 495 ± 1 , 832 ± 1 , 664 ± 1 and 1173 ± 1 );
[0077] - chosen between a first pattern 320 ± 1 , 1002 ± 1 or a second pattern 320 ± 1 , 2195 ± 1 in the nucleotide sequence of the centromere of chromosome 8 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated along the nucleotide sequence of the peri / centromere of chromosome 8 with these main nucleotide distances according to two patterns: ...320-1002... and ...320- 2195... , wherein each distance can vary by + / - 1 nucleotide. It is clear to a person skilled in the art that the pattern 320 ± 1 - 1002 ± 1 can be read also as 1002 ± 1 - 320 ± 1 , and the pattern 320 ± 1 , 2195 ± 1 can be read also as 2195 ± 1 - 320 ± 1 , i.e. the pattern can start on any distance value mentioned in the pattern string according to a sliding window of 1 value over each distance. Said patterns can be described also as (320 ± 1 , 1002 ± 1 )n or (320 ± 1 , 2195 ± 1 )n, wherein n is the number of repetition of said distance values in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively);
[0078] - chosen between a first pattern (1173 ± 1 , 1853 ± 1 )n or a second pattern (1173 ± 1 )n, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively, in the nucleotide sequence of the centromere of chromosome 13 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated along the nucleotide sequence of the peri / centromere of chromosome 13 with these main different nucleotide distances organized according two patterns: ... 1173-1853-... or ... 1173-1173-1173... , wherein each distance can vary by + / - 1 nucleotide It is clear to a person skilled in the art that the pattern (1173 ± 1 , 1853 ± 1 )n can be read also as (1853 ± 1 , 1173 ± 1 )n, i.e. the pattern can start on any distance value mentioned in the pattern string according to a sliding window of 1 value over each distance);
[0079] - 838 ± 1 , 1005 ± 1 , 663 ± 1 in the nucleotide sequence of the centromere of chromosome 15 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated with these maindifferent nucleotide distances organized in tandem as follows: ...838- 1005-663... wherein each distance can vary by + / - 1 nucleotide. The pattern ...838-1005-663... is repeated along the nucleotide sequence of the peri / centromere. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 1005 ± 1 , 663 ± 1 and 838 ± 1 ; 663 ± 1 , 838 ± 1 and 1005 ± 1 );
[0080] - chosen between a first pattern (2702 ± 1 )n or a second pattern (1343 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 20 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated along the nucleotide sequence of the peri / centromere of chromosome 20 with these main different nucleotide distances organized according to two patterns: ...2702-2702- 2702... and ...1343-1343-1343... wherein each distance can vary by + / - 1 nucleotide);
[0081] - (1853 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 21 , wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence). In other words, pJa motif is repeated along the nucleotide sequence of the peri / centromere of chromosome 21 with a nucleotide distance organized in tandem as follows: ...1853-1853-1853- 1853-1853... wherein each distance can vary by + / - 1 nucleotide); - chosen between a first pattern (498 ± 1 , 153 ± 1 )n or a second pattern (668 ± 1 , 2725 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 22, wherein n is the number of repetition of said distance values in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively, (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated along the nucleotide sequence of the peri / centromere of chromosome 22 with these main different nucleotide distances organized in tandem according two patterns: ...498-153... and ...668-2725... , wherein each distance can vary by + / - 1 nucleotide. It is clear to a person skilled in the art that the pattern (498 ± 1 , 153 ± 1 )n can be read also as (153 ± 1 , 498 ± 1 )n, and the pattern (668 ± 1 , 2725 ± 1 )n can be read also as (2725 ± 1 , 668 ± 1 )n, i.e. the pattern can start on any distance value mentioned in the pattern string according to a sliding window of 1 value over each distance).
[0082] Therefore, according to the present invention, by the characterization of each chromosome-specific centromere patterns, one can consequently identify the chromosome to which the centromere belongs from sequences and raw-sequencing data (fastq).
[0083] The motifs of present invention can be advantageously used in in vitro diagnostic and genetic methods. For example, the motifs of the present invention can be advantageously used for studying the nucleotide sequence of a centromere of a chromosome in order to identify: chromosome specific reads from rawsequencing data; centromere variants, mutations or changes associated with aberrant DNA, human diseases and pathological mutagenesis; to assess the correct order, organization and size of the centromere nucleotide sequence; to karyotype and detect absence / presence of individual chromosomes; to facilitate and / or validate human genome assembly. According to the present invention, the size of the active-centromere nucleotide sequence can be inferred based on the number and mode of motifs-specific patterns occurring in an assembled genome or from raw-sequencing data (after normalizing and considering the coverage, misjoined or collapsed reads.
[0084] The present invention concerns also a method for characterizing, identifying, isolating, assembling or validating the assembling of the nucleotide sequence of a centromere of a human chromosome, said method comprising the following steps: a) identifying in a fully or partially assembled genome or in raw sequencing long-reads (obtained by third-generation sequencing) or in contigs a motif, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or *CCCG**T****CGAA*; or pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2); b) identifying and isolating a centromere of a chromosome by selecting in the partially or fully assembled genome or in the raw sequencing long reads or in the contigs, a nucleotide sequence wherein said motif is repeated several times along the nucleotide sequence according to a defined pattern of distances, (for example that the inventors have discovered and defined), between subsequent motifs, said pattern being specific for each chromosome, wherein the distance is the number of nucleotides between the last nucleotide of a motif and the first nucleotide of the subsequent motif; and optionally c) assembling the nucleotide sequences selected from the raw sequencing long reads or the contigs (considering the motif’s distance specific for each centromere of a chromosome) in step b) or, alternatively, validating the nucleotide sequences selected from the assembled genome in step b), so that the pattern are repeated as defined in step b) for each chromosome (i.e. according to the aforementioned periodicity and the correct orientation, positioning and distance to build the correct sequence of a centromere and of a chromosome for example as defined by the inventors).
[0085] According to the invention, said motifs are identified in the same DNA strand. Specifically, CENP box motif *TTCG****A**CGGG* or its reverse complement sequence *CCCG**T****CGAA* are both present on the same DNA strand; analogously, pJa protein binding motif SEQ ID NO:1 or its reverse complement sequence SEQ ID NO:2 are both present on the same DNA strand. Therefore, the above-mentioned CENP box and PJa sequences, *TTCG****A**CGGG* and SEQ ID NO:1 , are present on the same DNA strand with different orientations.
[0086] As stated above, the motif *TTCG****A**CGGG* can be the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence I I I CGTTGGAAACGGGA (SEQ ID NO:4); whereas the reverse motif *CCCG**T****CGAA* can be the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO:6).
[0087] According to the above-mentioned method of the present invention, the pattern of distances of step b) between subsequent CENP-B box protein binding motifs within the centromere can be:
[0088] - chosen among (322 ± 1 )nnucleotides, (323 ± 1 )nnucleotides or (324 ± 1 )nnucleotides, preferably (323 ± 1 )nnucleotides, more preferably (323)nnucleotides, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 5, for example at least 6, at least 7, at least 8, at least 9 or at least 10, in the nucleotide sequence of the centromere of chromosome 1 (i.e. the three distance values of 322 nucleotides, 323 nucleotides or 324 nucleotides are the three most represented distances in chromosome 1 . In particular, the most represented distance value is 323 and is repeated in tandem so that the unit of distancepattern is 323n, wherein n is greater than or equal to 5. Each distance value can vary by + / - 1 nucleotide; for example the distance value 322 ± 1 nucleotides means 321 or 322 or 323. The term “in tandem” means that the distance value in the pattern is repeated one after the other consecutively);
[0089] - 321 ± 1 , 325 ± 1 nucleotides in the nucleotide sequence of the centromere of chromosome 2 (i.e. the pattern consists of CENP-B box protein binding motif repeated with these main different alternated distances, 321 and 325 nucleotides. In other words the pattern is 321 -325, that is repeated along the nucleotide sequence of the centromere for example as follows: ...321 -325- 321 -325-321 -325... , wherein each distance can vary by + / - 1 nucleotide. It is clear to a person skilled in the art that the pattern can be read as 321 ± 1 - 325 ± 1 or 325 ± 1 - 321 ± 1 , i.e. the pattern can start on any value mentioned in the pattern string. Said pattern can be described also as (321 ± 1 , 325 ± 1 )n, wherein n is the number of repetition of said distance values in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively, or more times); - 497 ± 1 , 321 ± 1 , 324 ± 1 , 324 ± 1 , 322 ± 1 , 321 ± 1 and 663 ± 1 in the nucleotide sequence of the centromere of chromosome 3 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motif is repeated with these main different nucleotide distances organized in tandem as follows: ...497 — 321- 324-324-322-321-663... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 321 ± 1 , 324 ± 1 , 324 ± 1 , 322 ± 1 , 321 ± 1 , 663 ± 1 and 497 ± 1 ; 324 ± 1 , 324 ± 1 , 322 ± 1 , 321 ± 1 , 663 ± 1 , 497 ± 1 and 321 ± 1 ; 324 ± 1 , 322 ± 1 , 321 ± 1 , 663 ± 1 , 497 ± 1 , 321 ± 1 and 324 ± 1 ; 322 ± 1 , 321 ± 1 ,
[0090] 663 ± 1 , 497 ± 1 , 321 ± 1 , 324 ± 1 and 324 ± 1 ; 321 ± 1 , 663 ± 1 , 497 ± 1 ,
[0091] 321 ± 1 , 324 ± 1 , 324 ± 1 and 322 ± 1 ; 663 ± 1 , 497 ± 1 , 321 ± 1 , 324 ± 1 ,
[0092] 324 ± 1 , 322 ± 1 and 321 ± 1 ; 497 ± 1 , 321 ± 1 , 324 ± 1 , 324 ± 1 , 322 ± 1 ,
[0093] 321 ± 1 and 663 ± 1 );
[0094] - 664 ± 1 , 153 ± 1 , 323 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 491 ± 1 and 324 ± 1 in the nucleotide sequence of the centromere of chromosome 4 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...664- 153-323-321-325-495-491-324... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 153 ± 1 , 323 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 491 ± 1 , 324 ± 1 and 664 ± 1 ; 323 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 491 ± 1 , 324 ± 1 , 664 ± 1 and 153 ± 1 ; 321 ± 1 , 325 ± 1 , 495 ± 1 , 491 ± 1 , 324 ± 1 , 664 ± 1 , 153 ± 1 and 323 ± 1 ; 325 ± 1 , 495 ± 1 , 491 ± 1 , 324 ± 1 , 664 ± 1 , 153 ± 1 ,
[0095] 323 ± 1 and 321 ± 1 ; 495 ± 1 , 491 ± 1 , 324 ± 1 , 664 ± 1 , 153 ± 1 , 323 ± 1 ,
[0096] 321 ± 1 and 325 ± 1 ; 491 ± 1 , 324 ± 1 , 664 ± 1 , 153 ± 1 , 323 ± 1 , 321 ± 1 ,
[0097] 325 ± 1 and 495 ± 1 ; 324 ± 1 , 664 ± 1 , 153 ± 1 , 323 ± 1 , 321 ± 1 , 325 ± 1 ,
[0098] 495 ± 1 and 491 ± 1 );
[0099] - 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 in the nucleotide sequence of the centromere of chromosome 5 (i.e. this is the unit of distancepattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...323-323-323-323- 323-322... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 322 ± 1 and 323 ± 1 ; 323 ± 1 , 323 ±
[0100] 1 , 323 ± 1 , 322 ± 1 , 323 ± 1 and 323 ± 1 ; 323 ± 1 , 323 ± 1 , 322 ± 1 , 323 ± 1 ,
[0101] 323 ± 1 and 323 ± 1 ; 323 ± 1 , 322 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 and 323 ±
[0102] 1 ; 322 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 and 323 ± 1 );
[0103] - 665 ± 1 , 323 ± 1 , 323 ± 1 , 320 ± 1 , 153 ± 1 , 323 ± 1 and 832 ± 1 in the nucleotide sequence of the centromere of chromosome 6 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...665-323- 323-320-152-323-832... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 323 ± 1 , 320 ± 1 , 153 ± 1 , 323 ± 1 , 832 ± 1 and 665 ± 1 ; 323 ± 1 , 320 ± 1 , 153 ± 1 , 323 ± 1 , 832 ± 1 , 665 ± 1 and 323 ± 1 ; 320 ±
[0104] 1 , 153 ± 1 , 323 ± 1 , 832 ± 1 , 665 ± 1 , 323 ± 1 and 323 ± 1 ; 153 ± 1 , 323 ± 1 ,
[0105] 832 ± 1 , 665 ± 1 , 323 ± 1 , 323 ± 1 and 320 ± 1 ; 323 ± 1 , 832 ± 1 , 665 ± 1 ,
[0106] 323 ± 1 , 323 ± 1 , 320 ± 1 and 153 ± 1 ; 832 ± 1 , 665 ± 1 , 323 ± 1 , 323 ± 1 ,
[0107] 320 ± 1 , 153 ± 1 and 323 ± 1 );
[0108] - 325 ± 1 , 323 ± 1 , 323 ± 1 and 325 ± 1 in the nucleotide sequence of the centromere of chromosome 7 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...325-323-323-325... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 323 ± 1 , 325 ± 1 and 325 ± 1 ; 323 ± 1 , 325 ± 1 , 325 ± 1 and 323 ± 1 ; 325 ± 1 , 325 ± 1 , 323 ± 1 and 323 ± 1 );
[0109] - 321 ± 1 , 154 ± 1 and 667 ± 1 in the nucleotide sequence of the centromere of chromosome 8 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -154-667... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 154 ± 1 , 667 ± 1 and 321 ± 1 ; 667 ± 1 , 321 ± 1 and 321 ± 1 );
[0110] - 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 and 495 ± 1 in the nucleotide sequence of the centromere of chromosome 9 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -325- 321 -325-321 -325-495... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 and 321 ± 1 ; 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 321 ± 1 and 325 ± 1 ; 325 ±
[0111] 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 321 ± 1 , 325 ± 1 and 321 ± 1 ; 321 ± 1 , 325 ± 1 ,
[0112] 495 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 and 325 ± 1 ; 325 ± 1 , 495 ± 1 , 321 ± 1 ,
[0113] 325 ± 1 , 321 ± 1 , 325 ± 1 and 321 ± 1 ; 495 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 ,
[0114] 325 ± 1 , 321 ± 1 and 325 ± 1 );
[0115] - 323 ± 1 , 323 ± 1 , 322 ± 1 and 321 ± 1 in the nucleotide sequence of the centromere of chromosome 10 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with 3 different nucleotide distances organized in tandem as follows: ...323-323-322-323-323-322-321 ... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 322 ± 1 , 321 ± 1 and 323 ± 1 ; 322 ± 1 , 321 ± 1 , 323 ± 1 and 323 ± 1 ; 321 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 );
[0116] - 321 ± 1 , 495 ± 1 in the nucleotide sequence of the centromere of chromosome 11 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -495-321 -495... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read also as: 495 ± 1 , 321 ± 1. Said pattern can be also described as ( 321 ± 1 , 495 ± 1 )n, wherein n is the number of repetition of said distance values in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively);
[0117] - 323 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 in the nucleotide sequence of the centromere of chromosome 12 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...323-323-323-322... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 323 ± 1 , 322 ± 1 and 323 ± 1 ; 323 ± 1 , 322 ± 1 , 323 and 323 ± 1 ; 322 ± 1 , 323, 323 ± 1 and 323 ± 1 );
[0118] - chosen between a first pattern 323 ± 1 , 324 ± 1 and 492 ± 1 and a second pattern 325 ± 1 , 321 ± 1 , 323 ± 1 , 324 ± 1 and 492 ± 1 in the nucleotide sequence of the centromere of chromosome 13 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances that can be organized in tandem as follows with 2 options: ...323-324-492... and ...325-321-323-324-492... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 324 ± 1 , 492 ± 1 and 323 ± 1 ; 492 ± 1 , 323 ± 1 and 324 ± 1 the first pattern, and 321 ± 1 , 323 ± 1 , 324 ± 1 , 492 ± 1 and 325 ± 1 ; 323 ± 1 , 324 ± 1 , 492 ± 1 , 325 ± 1 and 321 ± 1 ; 324 ± 1 , 492 ± 1 , 325 ± 1 , 321 ± 1 and 323 ± 1 ; 492 ± 1 , 325 ± 1 , 321 ± 1 , 323 ± 1 and 324 ± 1 the second pattern);
[0119] - 321 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 and 321 ± 1 n in the nucleotide sequence of the centromere of chromosome 14 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -325-325-325-321 ... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 325 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 and 321 ± 1 ; 325 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 and 325 ± 1 ; 325 ± 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 and 325 ± 1 ; 321 ± 1 , 321 ± 1 , 325 ± 1 , 325 ± 1 and 325 ± 1 );
[0120] - 154 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 and 325 ± 1 in the nucleotide sequence of the centromere of chromosome 15 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...154- 325-325-325-321 -325-321 -325... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 325 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 and 154 ± 1 ; 325 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1, 154 ±1 and 325 ± 1 ; 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 154 ± 1 , 325 ± 1 and 325 ± 1; 321 ± 1, 325 ± 1, 321 ± 1, 325 ± 1, 154 ± 1, 325 ± 1,
[0121] 325 ± 1 and 325 ± 1; 325 ± 1, 321 ± 1, 325 ± 1, 154 ± 1, 325 ± 1, 325 ± 1,
[0122] 325 ± 1 and 321 ± 1; 321 ± 1, 325 ± 1, 154 ± 1, 325 ± 1, 325 ± 1, 325 ± 1,
[0123] 321 ± 1 and 325 ± 1; 325 ± 1, 154 ± 1, 325 ± 1, 325 ± 1, 325 ± 1, 321 ± 1,
[0124] 325 ±1 and 321 ±1);
[0125] - 322 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 in the nucleotide sequence of the centromere of chromosome 16 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...322-323-323-323-323... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 ; 323 ± 1 , 323 ± 1 , 323 ± 1 , 322 ± 1 and 323 ± 1 ; 323 ± 1 , 323 ± 1 , 322 ± 1 , 323 ± 1 and 323 ± 1 ; 323 ± 1 , 322 ± 1, 323 ± 1, 323 ± 1 and 323 ± 1);
[0126] - 497 ± 1, 496 ±1, 150 ± 1, 151 ± 1 , 320 ± 1 , 149 ± 1, 667 ±1 and 149 ±1 in the nucleotide sequence of the centromere of chromosome 17 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...497- 496-150-151-320-149-667-149... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 496 ± 1 , 150 ± 1 , 151 ± 1 , 320 ± 1 , 149 ± 1 , 667 ±1, 149 ±1 and 497 ±1; 150 ± 1 , 151 ± 1 , 320 ± 1 , 149 ± 1 , 667 ± 1 , 149 ± 1, 497 ± 1 and 496 ± 1 ; 151 ± 1 , 320 ± 1 , 149 ± 1 , 667 ± 1 , 149 ± 1 , 497 ± 1 ,
[0127] 496 ± 1 and 150 ± 1; 320 ± 1, 149 ± 1, 667 ± 1, 149 ± 1, 497 ± 1, 496 ± 1,
[0128] 150 ± 1 and 151 ± 1; 149 ± 1, 667 ± 1, 149 ± 1, 497 ± 1, 496 ± 1, 150 ± 1,
[0129] 151 ± 1 and 320 ± 1 ; 667 ± 1 , 149 ± 1 , 497 ± 1 , 496 ± 1 , 150 ± 1 , 151 ± 1 , 320 ± 1 and 149 ± 1 ; 149 ± 1 , 497 ± 1 , 496 ± 1 , 150 ± 1 , 151 ± 1 , 320 ± 1 , 149 ± 1 and 667 ± 1 );
[0130] - 321 ± 1 , 320 ± 1 , 321 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 in the nucleotide sequence of the centromere of chromosome 18 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -320-321 -325-325-321 ... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start considering a sliding window of 1 over each distance, specifically the pattern can be read as: 320 ± 1 , 321 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 and 321 ± 1 ; 321 ± 1 , 325 ± 1 , 325 ±
[0131] 1 , 321 ± 1 , 321 ± 1 and 320 ± 1 ; 325 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 , 320 ± 1 and 321 ± 1 ; 325 ± 1 , 321 ± 1 , 321 ± 1 , 320 ± 1 , 321 ± 1 and 325 ± 1 ; 321 ±
[0132] 1 , 321 ± 1 , 320 ± 1 , 321 ± 1 , 325 ± 1 and 325 ± 1 );
[0133] - chosen among (323 ± 1 )nnucleotides, (322 ± 1 )nnucleotides or (663 ± 1 )nnucleotides, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 5, for example at least 6, at least 7, at least 8, at least 9 or at least 10, in the nucleotide sequence of the centromere of chromosome 19 (i.e. the three distance values of 323 ± 1 nucleotides, 322 ± 1 nucleotides or 663 ± 1 nucleotides are the three most represented distances in chromosome 19). For example CENP-B box protein binding motifs are repeated in tandem as follows: ...323-323-323-323-323- 323-323-323-323-323-323-323-323-323-323... wherein each distance can vary by + / - 1 nucleotide);
[0134] - 663 ± 1 , 324 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 in the nucleotide sequence of the centromere of chromosome 20 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence). In other words, CENP-B box protein binding motifs are repeated with these-main different nucleotide distances organized in tandem as follows: ...663-324- 321 -325-321 -321 -325... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 324 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 and 663 ± 1 ; 321 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 , 663 ± 1 and 324 ± 1 ; 325 ±
[0135] 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 , 663 ± 1 , 324 ± 1 and 321 ± 1 ; 321 ± 1 , 321 ± 1 ,
[0136] 325 ± 1 , 663 ± 1 , 324 ± 1 , 321 ± 1 and 325 ± 1 ; 321 ± 1 , 325 ± 1 , 663 ± 1 ,
[0137] 324 ± 1 , 321 ± 1 , 325 ± 1 and 321 ± 1 ; 325 ± 1 , 663 ± 1 , 324 ± 1 , 321 ± 1 ,
[0138] 325 ± 1 , 321 ± 1 and 321 ± 1 );
[0139] - 492 ± 1 , 325 ± 1 , 321 ± 1 , 323 ± 1 , 324 ± 1 in the nucleotide sequence of the centromere of chromosome 21 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence). In other words, CENP-B box protein binding motifs are repeated with different nucleotide distances organized in tandem as follows: ...492-325-321 -323-324... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 325 ± 1 , 321 ± 1 , 323 ± 1 , 324 ± 1 and 492 ± 1 ; 321 ± 1 , 323 ± 1 , 324 ± 1 , 492 ± 1 and 325 ± 1 ; 323 ± 1 , 324 ± 1 , 492 ± 1 , 325 ± 1 and 321 ± 1 ; 324 ± 1 , 492 ± 1 , 325 ± 1 , 321 ± 1 and 323 ± 1 );
[0140] - 321 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 in the nucleotide sequence of the centromere of chromosome 22 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated with these main different nucleotide distances organized in tandem as follows: ...321 -325-325-325... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 325 ± 1 , 325 ± 1 , 325 ± 1 and 321 ± 1 ; 325 ± 1 , 325 ± 1 , 321 ± 1 and 325 ± 1 ; 325 ± 1 , 321 ± 1 , 325 ± 1 and 325 ± 1 );
[0141] - chosen between 834 ± 1 , 494 ± 1 , 511 ± 1 , 150 ± 1 or 834 ± 1 , 150 ± 1 , 511 ± 1 , 494 ± 1 in the nucleotide sequence of the centromere of chromosome X (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, CENP-B box protein binding motifs are repeated according to these main different distance pattern in which the nucleotide distances are organized in tandem as follows: ...834-494-511 -150... and ...834-150-511 -494... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 494 ± 1 , 511 ± 1 , 150 ± 1 and 834 ± 1 ; 511 ± 1 , 150 ± 1 , 834 ± 1 and 494 ± 1 ; 150 ± 1 , 834 ± 1 , 494 ± 1 and 511 ± 1 the first pattern and 150 ± 1 , 511 ± 1 , 494 ± 1 and 834 ± 1 ; 511 ± 1 , 494 ± 1 , 834 ± 1 and 150 ± 1 ; 494 ± 1 , 834 ± 1 , 150 ± 1 and 511 ± 1 the second pattern).
[0142] In addition, according to the above-mentioned method of the present invention, the pattern of distances of step b) between subsequent pJ -alpha (pJa) DNA motifs can be:
[0143] - 321 ± 1 and 325 ± 1 nucleotides in the nucleotide sequence of the centromere of chromosome 2 (i.e. the pattern consists of pJa motif repeated with these main different alternated distances, 321 and 325 nucleotides. In other words, the pattern is 321 -325 (as described even for the CENP-B box in this chromosome), that is repeated along the nucleotide sequence of the peri / centromere of chromosome 2: ...321 -325-321 -325-321 -325... , wherein each distance can vary by + / - 1 nucleotide. It is clear to a person skilled in the art that the pattern can be read as 321 ± 1 - 325 ± 1 or 325 ± 1 - 321 ± 1 , i.e. the pattern can start on any distance value mentioned in the pattern string according to a sliding window of 1 value over each distance. Said pattern can be described also as (321 ± 1 , 325 ± 1 )n, wherein n is the number of repetition of said distance values in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively);
[0144] - 832 ± 1 , 664 ± 1 , 1173 ± 1 , 495 ± 1 in the nucleotide sequence of the centromere of chromosome 4 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated with these main different nucleotide distances organized in tandem as follows: ...832-664-1173-495... wherein each distance can vary by + / - 1 nucleotide. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 664 ± 1 , 1173 ± 1 , 495 ± 1 and 832 ± 1 ; 1173 ± 1 , 495 ± 1 , 832 ± 1 and 664 ± 1 ; 495 ± 1 , 832 ± 1 , 664 ± 1 and 1173 ± 1 );
[0145] - chosen between a first pattern 320 ± 1 , 1002 ± 1 or a second pattern 320 ± 1 , 2195 ± 1 in the nucleotide sequence of the centromere of chromosome 8 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated along the nucleotide sequence of the peri / centromere of chromosome 8 with these main nucleotide distances according to two patterns: ...320-1002... and ...320- 2195... , wherein each distance can vary by + / - 1 nucleotide. It is clear to a person skilled in the art that the pattern 320 ± 1 - 1002 ± 1 can be read also as 1002 ± 1 - 320 ± 1 , and the pattern 320 ± 1 , 2195 ± 1 can be read also as 2195 ± 1 - 320 ± 1 , i.e. the pattern can start on any distance value mentioned in the pattern string according to a sliding window of 1 value over each distance. Said patterns can be described also as (320 ± 1 , 1002 ± 1 )n or (320 ± 1 , 2195 ± 1 )n, wherein n is the number of repetition of said distance values in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively);
[0146] - chosen between a first pattern (1173 ± 1 , 1853 ± 1 )n or a second pattern (1173 ± 1 )n, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively, in the nucleotide sequence of the centromere of chromosome 13 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated along the nucleotide sequence of the peri / centromere of chromosome 13 with these main different nucleotide distances organized according two patterns: ... 1173-1853-... or ... 1173-1173-1173... , wherein each distance can vary by + / - 1 nucleotide It is clear to a person skilled in the art that the pattern (1173 ± 1 , 1853 ± 1 )n can be read also as (1853 ± 1 , 1173 ± 1 )n, i.e. the pattern can start on any distance value mentioned in the pattern string according to a sliding window of 1 value over each distance);
[0147] - 838 ± 1 , 1005 ± 1 , 663 ± 1 in the nucleotide sequence of the centromere of chromosome 15 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated with these main different nucleotide distances organized in tandem as follows: ...838- 1005-663... wherein each distance can vary by + / - 1 nucleotide. The pattern ...838-1005-663... is repeated along the nucleotide sequence of the peri / centromere. The pattern can be repeated consecutively in tandem along the centromere sequence. It is clear to a person skilled in the art that the pattern can start at any value within the defined sequence of distances, specifically the pattern can be read as: 1005 ± 1 , 663 ± 1 and 838 ± 1 ; 663 ± 1 , 838 ± 1 and 1005 ± 1 );
[0148] - chosen between a first pattern (2702 ± 1 )n or a second pattern (1343 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 20 (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated along the nucleotide sequence of the peri / centromere of chromosome 20 with these main different nucleotide distances organized according to two patterns: ...2702-2702- 2702... and ...1343-1343-1343... wherein each distance can vary by + / - 1 nucleotide);
[0149] - (1853 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 21 , wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence). In other words, pJa motif is repeated along the nucleotide sequence of the peri / centromere of chromosome 21 with a nucleotide distance organized in tandem as follows: ...1853-1853-1853- 1853-1853... wherein each distance can vary by + / - 1 nucleotide);
[0150] - chosen between a first pattern (498 ± 1 , 153 ± 1 )n or a second pattern (668 ± 1 , 2725 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 22, wherein n is the number of repetition of said distance values in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively, (i.e. this is the unit of distance-pattern that is repeated along the centromere sequence. In other words, pJa motif is repeated along the nucleotide sequence of the peri / centromere of chromosome 22 with these main different nucleotide distances organized in tandem according two patterns: ...498-153... and ...668-2725... , wherein each distance can vary by + / - 1 nucleotide. It is clear to a person skilled in the art that the pattern (498 ± 1 , 153 ± 1 )n can be read also as (153 ± 1 , 498 ± 1 )n, and the pattern (668 ± 1 , 2725 ± 1 )n can be read also as (2725 ± 1 , 668 ± 1 )n, i.e. the pattern can start on any distance value mentioned in the pattern string according to a sliding window of 1 value over each distance).
[0151] According to the present invention, the above-mentioned method of the present invention can be a computer-implemented method and said steps a), b) and c) can be carried out by a processor according to code written by the inventors.
[0152] It is a further object of the present invention also a method for identifying possible pathologic changes (such as structural alternation, inversions, insertions, deletions or any other sequence change that alters the distances, orientation or motif) in the centromere nucleotide sequence of a human chromosome of a subject, said method comprising the following steps: a) identifying in the chromosome nucleotide sequence (for example in the genome or DNA or sequencing reads or contigs) of said subject a motif, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or *CCCG**T****CGAA*; or pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2); b) identifying the position of said motif in the chromosome nucleotide sequence; c) comparing the motifs, orientation and the positions thereof in the chromosome nucleotide sequence of the subject with the motifs, orientation and positions thereof in a reference nucleotide sequence wherein when a position of the motif, the orientation of the motif and / or the motif in the chromosome nucleotide sequence of the subject is different from the position and / or the orientation of the same motif and / or from the motif itself in the reference chromosome nucleotide sequence a possible pathologic change of the centromere nucleotide sequence of the chromosome is present. For example Figure 14 and Figure 15 show a pathological chromosomal aberration. The reference or benchmark chromosome nucleotide sequence is a correctly assembled chromosome nucleotide sequence of a reference / benchmark subject, for example a healthy subject. For example, in a healthy subject the positions of the motifs are organized in the centromere according to the patterns described above.
[0153] As stated above, according to the invention, said motifs are identified in the same DNA strand. Specifically, CENP box motif *TTCG****A**CGGG* or its reverse complement sequence *CCCG**T****CGAA* are both present on the same DNA strand; analogously, pJa protein binding motif SEQ ID NO:1 or its reverse complement sequence SEQ ID NO:2 are both present on the same DNA strand. Therefore, the above-mentioned CENP box and PJa sequences, *TTCG****A**CGGG* and SEQ ID NO:1 , are present on the same DNA strand with different orientations.
[0154] According to an embodiment of this method of the invention, the motif *TTCG****A**CGGG* can be the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4); whereas the reverse motif *CCCG**T****CGAA* can be the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO.6).
[0155] According to an embodiment of this method of the present invention, when the change in the centromere nucleotide sequence of a chromosome of a subject is a structural inversion, (as shown in Figure 14 and Figure 15), said structural inversion is present when in the same position in the centromere nucleotide sequence of the subject the motif is *TTCG****A**CGGG*, whereas in the reference centromere nucleotide sequence the motif is *CCCG**T****CGAA* or vice versa, and / or when in the same position in the centromere nucleotide sequence of the subject the motif is SEQ ID NO:1 , whereas in the reference centromere nucleotide sequence the motif is SEQ ID NO:2 or vice versa.
[0156] According to the present invention, also this method can be a computer- implemented method and said steps a), b) and c) can be carried out by a processor.
[0157] The present invention concerns also a method for calculating the size of the centromere nucleotide sequence of a human chromosome, preferably of the active (or functional) centromere nucleotide sequence, said method comprising the step of: a) identifying a motif in the centromere nucleotide sequence, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or *CCCG**T****CGAA*; or pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2), said motif being repeated several times along the nucleotide sequence of the centromere according to a pattern of distances between subsequent motifs; b) identifying the pattern of distances of said motif, wherein said pattern comprises a number K of distance values, said pattern being repeated several times along the nucleotide sequence of the centromere and being specific for each chromosome, wherein the distance is the number of nucleotides between the last nucleotide of a motif and the first nucleotide of the subsequent motif; c) calculating the frequency (i.e. the number of times) with which the pattern is repeated in the centromere nucleotide sequence and d) calculating the size of the centromere nucleotide sequence by the following formula: centromere size (bp) = {(sum of distance values of the pattern (bp)) + [(17bp x k) + 17bp]} x Frequency.
[0158] As stated above, according to the invention, said motifs are identified in the same DNA strand. Specifically, CENP box motif *TTCG****A**CGGG* or its reverse complement sequence *CCCG**T****CGAA* are both present on the same DNA strand; analogously, pJa protein binding motif SEQ ID NO:1 or its reverse complement sequence SEQ ID NO:2 are both present on the same DNA strand. Therefore, the above-mentioned CENP box and PJa sequences, *TTCG****A**CGGG* and SEQ ID NO:1 , are present on the same DNA strand with different orientations.
[0159] According to the present invention, the motif *TTCG****A**CGGG* can be the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4); whereas the reverse motif *CCCG**T****CGAA* can be the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence
[0160] ACCCGTTTCCAACGAAA (SEQ ID NO.6).
[0161] According to the present invention, the method for calculating the size of the centromere nucleotide sequence of a chromosome as defined above can be a computer-implemented method and said steps a), b), c) and d) are carried out by a processor.
[0162] It is a further object of the present invention the use of a DNA motif as a marker for characterizing, identifying, isolating, assembling or validating the nucleotide sequence of a human chromosome, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or *CCCG**T****CGAA*; or a pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2); said motif being repeated several times along the nucleotide sequence of the chromosome in conserved positions forming a specific pattern (for instance in the form of a bar code) for each chromosome.
[0163] According to the use of the invention, the motif *TTCG****A**CGGG* can be the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4); whereas the reverse motif *CCCG**T****CGAA* can be the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO.6).
[0164] The present invention also concerns a method for characterizing, identifying, isolating, assembling or validating the nucleotide sequence and structural organization of a chromosome, said method comprising the following steps: a) identifying, in an assembled genome or in raw sequencing long-reads or in contigs, one or more motifs, wherein said one or more motifs are: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or its reverse complement *CCCG**T****CGAA*; and / or
[0165] > pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2), wherein said one or more motifs are repeated along the same strand of the nucleotide sequence of a chromosome according to a pattern of positions (for instance in the form of a barcode), said pattern being specific for each human chromosome; b) visually comparing the pattern of positions of said one or more motifs in said assembled genome or said raw sequencing long-reads or said contigs with the pattern of positions of the same motifs in a correctly assembled and validated reference chromosome (for instance in the form of a barcode), wherein said assembled genome or said raw sequencing long-reads or said contigs belong to the chromosome having a pattern which is visually superimposable with their pattern; and optionally c) assembling the nucleotide the raw sequencing long reads or contigs or, alternatively, validating the nucleotide sequence of the assembled genome on the basis of the visual comparison of step b).
[0166] As stated above, according to the invention, said motifs are identified in the same DNA strand. Specifically, CENP box motif *TTCG****A**CGGG* or its reverse complement sequence *CCCG**T****CGAA* are both present on the same DNA strand; analogously, pJa protein binding motif SEQ ID NO:1 or its reverse complement sequence SEQ ID NO:2 are both present on the same DNA strand. Therefore, the above-mentioned CENP box and PJa sequences, *TTCG****A**CGGG* and SEQ ID NO:1 , are present on the same DNA strand with different orientations.
[0167] According to an embodiment of the method of the invention, CENP box motif *TTCG****A**CGGG* and its reverse complement sequence *CCCG**T****CGAA* are identified in the same DNA strand in order to obtain the pattern of positions of these two motifs in the form of a barcode. Preferably, the two motifs are visually represented in the barcode with bars of two different colors. In other words, CENP box motif *TTCG****A**CGGG* and its reverse complement sequence *CCCG**T****CGAA*, that are the same sequence with opposite orientation, identified in the same strand, can be represented with two different colors.
[0168] As stated above, the motif *TTCG****A**CGGG* can be the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4); whereas the reverse motif *CCCG**T****CGAA* can be the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO.6). According to the invention, step a) of said method can be carried out by identifying said motifs outside the centromere region, i.e. by identifying said motifs in the chromosome arms. In addition, step c) of aligning and visually comparing of said method can be carried out for example by using artificial intelligence or machine learning.
[0169] In particular, a comparison between the positions of the above-described motifs in a chromosome nucleotide sequence of a subject with the positions of the same motifs can be used for: a) assigning reads or contigs to a specific chromosome; b) validating if the assembly of the genomes and chromosomes / scaffolds is correct; c) validating if the analyzed piece of DNA belongs to a specific chromosome; d) evaluating if the chromosome is normal or altered, for example in pathological states.
[0170] More in particular, said comparison can be carried out by considering the positions of the motifs outside the centromere region.
[0171] Moreover, said comparison can be a visual comparison of the positions of the motifs in a chromosome nucleotide sequence of a subject with the positions of the motifs in a reference chromosome nucleotide sequence, which can be carried out for example by using artificial intelligence or machine learning.
[0172] According to the invention, said method can be a computer-implemented method and said steps a), b), c) and d) are carried out by a processor.
[0173] In addition, the present invention concerns a method for identifying an alteration of human chromosome nucleotide sequence of a subject, said method comprising the following steps: a) identifying in the chromosome nucleotide sequence (for example in the genome or DNA or sequencing reads) of said subject a motif, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or *CCCG**T****CGAA*; or pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2); b) identifying the positions of the motifs identified in step a) in the chromosome nucleotide sequence; and c) comparing the positions of motifs identified in step b) with the positions of the same motifs in a reference chromosome nucleotide sequence, wherein when the positions of the motifs in the chromosome nucleotide sequence of the subject are different from the positions of the same motifs in the reference chromosome nucleotide sequence (with the periodicity hereby described and defined by the Inventors) an alteration of chromosome nucleotide sequence is present. The reference or benchmark chromosome nucleotide sequence is a correctly assembled chromosome nucleotide sequence of a reference / benchmark subject, for example a healthy subject. For example, in a healthy subject the positions of the motifs are organized in the centromere according to the patterns described above. According to the invention, step a) of said method can be carried out by identifying said motifs outside the centromere region, i.e. by identifying said motifs in the chromosome arms.
[0174] According to the invention, the motif *TTCG****A**CGGG* can be the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4); whereas the reverse motif *CCCG**T****CGAA* can be the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO.6) According to an embodiment of the present invention, step c) of the method of the present invention can be carried out by visually comparing the positions of the motifs in a chromosome nucleotide sequence of a subject with the positions of the motifs in a reference chromosome nucleotide sequence, for example by using artificial intelligence or machine learning.
[0175] Also this method can be a computer-implemented method and said steps a), b) and c) can be carried out by a processor. According to the present invention, the methods described above can be used for the in vitro diagnosis of diseases characterized by chromosomal aberrations in the centromere and / or outside the centromere region (i.e. in the chromosome’s arms), also as a prenatal screening. Moreover, the methods described above according to the invention can be used for the in vitro diagnosis of tumors characterized by chromosomal aberrations in the centromere and / or outside the centromere region.
[0176] In addition, the present invention also concerns a method for diagnosing and treating a disease characterized by chromosomal aberrations in the centromere and / or outside the centromere region (i.e. in the chromosome’s arms), said method comprising a) carrying out the steps of any of the methods according to the invention as described above in order to diagnose said disease and b) treating a subject in need thereof by administering to the subject a therapy which is suitable for said disease.
[0177] The present invention also concerns a system for characterizing, identifying, isolating, assembling or validating the nucleotide sequence of a centromere of a chromosome, said system comprising:
[0178] (A) storage media in which the following data are stored: an assembled genome or raw sequencing long-reads or contigs, such as in fasta or fastq format; and (B) control logic unit connected to said storage media and configured to perform steps a) to c) of the method as defined in any of claims 5-9.
[0179] In addition, the present invention concerns a system for identifying possible pathologic changes in the centromere nucleotide sequence of a chromosome of a subject, said system comprising:
[0180] (A) storage media in which the following data are stored: a chromosome nucleotide sequence of said subject, such as in fasta or fastq format; and
[0181] (B) control logic unit connected to said storage media and configured to perform steps a) to c) of the method as defined in any of claims 10-13.
[0182] The present invention concerns also a system for calculating the size of the centromere nucleotide sequence of a chromosome, said system comprising:
[0183] (A) storage media in which the following data are stored: a centromere nucleotide sequence of said subject, such as in fasta or fastq format; and
[0184] (B) control logic unit connected to said storage media and configured to perform steps a) to d) of the method as defined in any of claims 14-16.
[0185] The invention also concerns a system for characterizing, identifying, isolating, assembling or validating the nucleotide sequence of a chromosome, said system comprising:
[0186] (A) storage media in which the following data are stored: assembled genome or raw sequencing long-reads or contigs; and
[0187] (B) control logic unit connected to said storage media and configured to perform steps a) to c) of the method as defined in any of claims 19-21 .
[0188] Moreover, the present invention concerns a system for identifying an alteration of chromosome nucleotide sequence of a subject, said system comprising: (A) storage media in which the following data are stored: a chromosome nucleotide sequence of said subject, such as in fasta or fastq format; and
[0189] (B) control logic unit connected to said storage media and configured to perform steps a) to c) of the method as defined in any of claims 22-25.
[0190] According to the present invention, the systems as defined above can further comprise display means for displaying a report of the results of the performed steps, such as the motifs patterns.
[0191] The methods according to the present invention can be computational methods, such as a method for the characterization of Centromere Nucleotide Sequences, or other computer-generated methods, artificial intelligence methods and non computer-implemented methods.
[0192] In conclusion, this invention presents DNA motifs for identifying, evaluating and assembling chromosome centromeres and also chromosome zones outside the centromere, essential in genomic engineering, personalized medicine, and chromosomal therapy. It addresses the challenge of identifying and validating centromeric sequences amid genetic variability, with applications in pharmaceutical research, genetic diagnostics, and biotechnology. The invention includes the Giunta-Corda Genomic-Centromere Finder (GCF) for enhanced genome assembly, particularly in repetitive regions. It also offers a method for analyzing chromosomespecific CENP-B box organization, aiding in the study of centromeres and chromosomal abnormalities. Additionally, the invention's 'centromere barcode' facilitates the detection of centromeric DNA changes and assists in diagnosing related diseases.
[0193] In consideration of the described advancements, the present invention has significant applicability in the industrial sector, particularly in genomic engineering, personalized medicine, and chromosomal assembly and studies. The invention encompasses DNA motifs and methodologies that facilitate the manipulation of centromeric regions, which is vital for the development of artificial chromosomes, genetic interventions and therapies. The capability of this invention to identify and validate centromeric sequences amid variability is a notable development in genomic research.
[0194] The invention's ability to recognize CENP-B motifs as fixed spacers within polymorphic centromere domains addresses a key challenge in genomic research and offers new possibilities in therapeutic and diagnostic applications. This technology is particularly relevant in the diagnosis of genetic disorders involving chromosomal segregation and others involving changes in centromere DNA or chromosomal DNA that alter the established distances of the centromere footprint hereby defined by the Inventors. The conservation of the centromeric features identified by the Inventors across species contributes to comparative genomics, aiding in genomic, evolutionary and medical research.
[0195] The invention introduces the Giunta-Corda Genomic-Centromere Finder (GCF), a computational toolset that enhances genome assembly and validation, especially in complex repetitive regions. This development introduces new methods for centromeres and genome assemblies and validation useful in personalized medicine and genetic research and also improves the functionality of existing sequencing technologies to help with analyses and assembly.
[0196] Additionally, the invention provides a detailed analysis of chromosomespecific CENP-B box organization, which is essential for advancing diagnostic capabilities in the study of chromosomal abnormalities. This analysis enables targeted and efficient research, increasing the accuracy of chromosome-specific studies pertaining to centromeres. The discovery of variant patterns in different chromosome clusters has implications for understanding chromosomal dynamics, evolution and centromere biology.
[0197] The invention's establishment of a 'centromere barcode' for each chromosome is a significant contribution to genomic analysis. This tool aids in the detection and analysis of centromeric DNA changes, the diagnosis of diseases related to these regions, and the enhancement and validation of genome assembly. The invention thus represents a notable development in genomic research, with applications in diagnostics, clinical settings, and broader scientific inquiry.
[0198] The present invention will now be described, by way of not limitative illustration, with particular reference to some examples and figures, wherein:
[0199] Referring to Figure 1 , panels A-B illustrate the Genomic Centromere Footprint method, the core workflow of the present invention (panel A: distance distribution for centromere DNA motifs is specific within human chromosomes; legend of panel A: 1 : Starting from a complete genome or from long-reads sequencing, conserved chromosome-specific distances identified with sequence motif *TTCG****A**CGGG* and *CCCG**T****CGAA*; 2: Giunta-Corda Genomic Centromere distance finder to retrieve the distance of each motif; 3: FastQ reads extraction based on chromosome-specific pattern; 4: Reads filtering and variant calling at the human centromeres. Panel B: Computational characterization of centromere nucleotide sequences based on distance distribution of the CENP-B box motif (*TTCG****A**CGGG* & *CCCG**T****CGAA*) specific to chromosome and conserved in human genomes). This method provides a comprehensive analysis of human centromeres, facilitating the identification of the conserved, correct and chromosome-specific centromere organization and also flags dysfunctional organization that differs from the pattern identified in this invention. The invention is applied to DNA sequences. Obtaining DNA sequences encompasses both laboratory (wet lab) and computational (dry lab) procedures, beginning with DNA extraction, followed by library preparation and DNA long-read sequencing. Subsequent analyses focus on the repetitive pattern of a specific centromeric sequence, pivotal for assembling, validating and characterizing each human chromosome. A novel aspect of this invention is its ability to identify and isolate these centromeres from sequencing data, offering a large-scale organizational overview of human centromeric sequences and immediate chromosome-specific assignment. This facilitates the characterization of centromeres across different genomes. The method employs mainly one conserved short DNA sequence, present in all chromosomes except the Y chromosome, that the Inventors have discovered having a specific pattern, distance, orientation different for each human chromosome to interrogate the structure of human centromeres at base pair resolution. Utilizing computational tools using custom R code written and assembled by the Inventors, this invention analyzes the distribution of this sequence throughout the entire genome, particularly in active higher order repeats (HORs) and centromere dip regions (CDRs) of each chromosome. Further detailed in Figure 1, panels C-F (heatmaps showing the normal distance distribution of the CENP-B box motif specific for each chromosome), are heatmaps that represent the 'reference genome fingerprint' of distances for the DNA motif specific for each chromosome and common to humans. These heatmaps display the most frequent distances of the CENP-B box within various human reference genomes examples, including CHM13, HG002, RPE1 , and HG38. For these heatmaps, only distances that account for more than 1 % in at least one chromosome have been included in the analysis. This approach underscores the invention's utility in providing a rapid, detailed and quantifiable understanding of centromeric sequence organization for each chromosome, a critical factor in genomic research and diagnostic applications. Regarding Figure 1 D, variation in the normal DNA Fingerprint identified in Fig. 1 C as shown in the dotted box indicate chromosomes with centromeres that are aberrant or not assembled correctly. Indeed, in HG38 some centromeres are incorrectly assembled.
[0200] The present invention, as depicted in Figure 2, encompasses a comprehensive analysis of CENP-B box enrichment and orientation within various human reference genomes, namely CHM13, RPE1 , HG002, and HG38. Specifically, Figure 2A (six most represented distances of the CENP-B box within each active HOR of the human genome - from CHM13) illustrates the count and orientation of CENP-B boxes in these genomes. Figure 2B (six most represented distances of the pJalpha motif within each active HOR of the human genome - from CHM13) further extends this analysis to T2T-CHM13v2.0, highlighting the Inventors discovery that the most prevalent distances of the CENP-B box sequence and pJa motif within all active higher-order-repeats (HORs) has a spacing specific to each chromosome.
[0201] In Figure 3 (CENP-B box distribution in each centromere CDR with specific position for chromosome 17 and X - derived from CHM13 human reference genome), the Inventors show the distribution of the CENP-B box within the CDRs of chromosomes 17 and X as an example for the T2T-CHM13 reference genome. This figure details the distance distribution at each position within the CDRs, offering insights into the structural organization of these specific chromosomal regions determined by the motif distances.
[0202] Figure 4, encompassing panels A-F, showcases the method of isolating sequencing reads that contain centromere DNA using the centromere barcode defined by the Inventors (panels 4A-4B: Isolating sequencing reads containing centromere DNA using the distances between motifs; legend of panel 4B 1 : Common pattern of chromosome 2 (chr2), 2: Chr4 specific pattern, 3: Chr8 specific pattern, 4: Chr11 specific pattern, 5: Chr15 specific pattern, 6: Chr17 specific pattern, 7: ChrX specific pattern; panel 4C: Identify and isolate sequencing reads containing centromere DNA of chromosome 3; legend of panel 4C: 1 and 2: Common pattern, 3: Chr3 specific pattern exist and was identified to isolate exclusively chr3 centromere sequencing reads, 4: Others chr specific pattern, 5: Chr4 specific pattern does not work to isolate chr3 reads given high specificity, 6: Chr17 specific pattern does not work to isolate chr3 reads given high specificity, 7: ChrX specific pattern does not work to isolate chr3 reads given high specificity; panel 4D: Identify and isolate sequencing reads containing centromere DNA of chromosome 4; legend of panel 4D 1 : Common pattern, 2: Common pattern, 3: Common pattern, 4: Common pattern, 5: Chr4 specific pattern, 6: Chr17 specific pattern, 7: ChrX specific pattern; panel 4E: Identify and isolate sequencing reads containing centromere DNA of chromosome 17; panel 4F: Identify and isolate sequencing reads containing centromere DNA of chromosome X). This process involves the alignment of reads post-isolation, which are catalogued based on their chromosome-specific centromere patterns determined by motifs distances.
[0203] Figures 5 and 6 (Changes in centromere distances identify aberrant DNA in a specific chromosome), across panels A-F and A-G respectively, demonstrate the application of the invention in variant calling for centromere DNA. Specifically, Figure 5 addresses chromosome 17, and Figure 6 focuses on chromosome X. In particular, panels 5A and 5B show deletions at the centromere of chromosome 17, panels 5C, 5D, 5E and 5F show insertions at the centromere of chromosome 17, panels 6A, 6B, 6C and 6D show deletions at the centromere of chromosome X, panels 6E, 6F and 6G show insertions at the centromere of chromosome X. Both figures 5 and 6 utilize only the raw reads extracted based on the unique pattern of each chromosome to identify variants such as abnormal deletions or insertions or other changes that cause a variation from the expected normal distribution of motifs that is essential for centromeres to work at least for part of the centromere. While complete deletion or other major changes of the centromere are likely incompatible with chromosome retention, smaller changes that weaken the centromere or cause chromosome dysfunction and related diseases can be detected with the invention and method as shown in Figures 5-6.
[0204] Figure 7, in panels A-B, introduces a method for assessing centromere size based on this invention. In particular, Figure 7 shows that pattern of distances (k- pattern) using Sequence *TTCG****A**CGGG*can be used to predict centromere size. More in detail, Figure 7A shows the assessment of centromere size of chromosome 6, whereas Figure 7B shows the assessment of centromere size of chromosome 10. By analyzing the frequency and repetition of each centromerespecific pattern, the invention allows for the inference of active centromere size, correlated with CENP-A peaks against each reference genome. The distribution of motifs identified in this invention is common to all humans; conversely, centromere size and in part sequence change from person to person, and even within the maternal and paternal haplotypes of the same genome, so using the pattern to establish the centromere size for each genomes haplotype is very useful and clinically relevant in physiology and pathology. Using the script pattern_finder.R, the Inventors were able to search for the shortest and longest unique patterns of distances for each chromosomes. The term k-pattern was devised, all the subpatterns for a specific motif that contains k number of distances present within the centromeric sequences. The k-pattern profile (k-pattern spectrum) of each chromosome in a genome assembly may represent an indicator of the complexity of the centromere under examination. The Figure 7 A-B represent the k-pattern spectrum for the chromosome 6 and 10, respectively. In both Figures 7A and 7B, y- axis is the frequency of centromeric sequence with a specific pattern of the CENP- B box, whereas x-axis comprises the specific patterns. The longest k-pattern is used to assess and predict the active centromere size for these chromosomes. In both figures 7A and 7B the symbol * indicates the longest k-pattern for the specific chromosome with the highest frequency, where k = n. of distances. The longest k- pattern (highlighted with the symbol * in Figure 7 A) with highest frequency for chromosome 6 is:
[0205] - CHM13 (15-patterns, k = 15): box-322 -box-831 -box-664-box-322-box-322- box-319-box-151 -box-322-box-831 -box-664-box-322-box-322-box-319-box- 151 -box-322-box;RPE1 (15-patterns, k = 15): box-152-box-323-box-832- box-665-box-323-box-323-box-320-box-152-box-323-box-832-box-665-box- 323-box-323-box-320-box-152-box;
[0206] - HG002 maternal (15-patterns, k = 15): box-320-box-323-box-323-box-665- box-832-box-323-box-152-box-320-box-323-box-323-box-665-box-832-box- 323-box-152-box-320-box;
[0207] - HG002 paternal (15-patterns, k = 15): box-320-box-152-box-323-box-832- box-665-box-323-box-323-box-320-box-152-box-323-box-832-box-665-box- 323-box-323-box-320-box;
[0208] - HG38 (12-patterns, k = 12): box-323-box-323-box-320-box-152-box-323- box-832-box-665-box-323-box-323-box-320-box-152-box-323-box;
[0209] The longest k-pattern (highlighted with the symbol * in Figure 7 B) with highest frequency for chromosome 10 is:
[0210] - CHM13 (14-patterns, k = 14): box-322 -box-322-box-321 -box-320-box-322- box-322-box-321 -box-322-box-322 -box-321 -box-320-box-322-box-322-box- 321 -box;
[0211] - RPE1 (15-patterns, k = 15): box-323-box-322-box-323-box-323-box-322- box-321 -box-323-box-323-box-322-box-323-box-323-box-322-box-321 -box- 323-box-323-box;
[0212] - HG002 maternal (11 -patterns, k = 11 ): box-323-box-323-box-320-box-323- box-323-box-322-box-323-box-323-box-320-box-323-box-323-box;
[0213] - HG002 paternal (9-patterns, k = 9): box-323-box-323-box-322-box-323-box- 323-box-320-box-323-box-323-box-322-box;
[0214] - HG38 (8-patterns, k = 8): box-321 -box-323-box-323-box-323-box-321 -box- 323-box-323-box-323-box;
[0215] The formula to calculate the active functional centromere size is the following: Centromere size = {sum of distances of the longest k-patterns + [(17 bp motif x k) + 17bp]} x Frequency where is used the longest k-pattern for this specific chromosome with the highest frequency, where k = n. of distances, where x is a symbol for multiplication.
[0216] For example, for the following longest possible string of values to make the longest 15-pattern (k-pattern where k = 15) for chromosome 6 in the CHM13 reference genome, the size calculation of the active centromere array is:
[0217] Centromere size =
[0218] {322+831 +664+322+322+319+151 +322+831 +664+322+322+319+151 +322+[(17 x 15) + 17]} x 27
[0219] Figures 7A and 7B also show that the size of centromeres changes between genomes, while pattern is conserved.
[0220] In Figure 8, the invention provides a schematic representation of the most represented distances within the CHM13 and HG38 reference genomes. In particular, Fig. 8 shows how the method according to the invention can be a powerful tool for testing correct centromere assembly, genetic screenings and diagnostic purpose. This schematic shows the invention as instrumental in identifying discrepancies in the assembly of specific chromosomal regions, particularly highlighting potential assembly inaccuracies in chromosomes 13 through 22 in HG38 flagged using the invention hereby described. Such insights are valuable for identifying centromere-associated diseases in pathological conditions and as a rapid assessment of human genome assembly correctness. Figure 9 emphasizes the invention’s capability for fast and accurate identification of human centromeres through chromosome-specific reads-mapping post-isolation from raw sequencing data. Legend of Figure 9 1 : Common pattern, 2: Common pattern, 3: Common pattern, 4: Common pattern, 5: Common pattern, 6: Chr17 specific pattern, 7: ChrX specific pattern.
[0221] Figure 10 further elucidates the potential for identifying mutations and changes in centromere DNA. This figure showcases variant calling post-mapping of catalogued reads against the CHM13 reference genome.
[0222] Figure 11 demonstrates the isolation of sequencing reads containing centromere DNA, providing examples of chromosome-specific reads mapping. This includes the identification of common patterns shared among chromosomes 2, 9, 15, 20, and 21 , as well as unique patterns in chromosomes 4, 8, 11 , 15, 17, and X. Legend of Figure 11: 1 : Chr2 pattern, 2: Chr4 specific pattern, 3: Chr8 specific pattern, 4: Chr11 specific pattern, 5: Chr15 specific pattern, 6: Chr17 specific pattern, 7: ChrX specific pattern.
[0223] Figure 12 shows how the present invention enables the validation of de novo human genome assemblies for assignment of centromere reads in the correct order and to the correct chromosome based on the barcode. It presents the centromere organization in the RPE1 genome, illustrating consistency with the centromere organization in the CHM13 benchmark genome but also differences where the RPE1 genome was not assembled correctly by the latest automatic assembler due to current scarcity in validation methods for centromere assembly.
[0224] Figure 13 shows the conserved pattern in primates using centromere DNA motif sequence *TTCG****A**CGGG*& *CCCG**T****CGAA*. In particular, this figure expands the scope of the invention to primates, showing the CENP-B box distances within the reference genomes of species such as gorillas, chimpanzees, bonobos, orangutans, gibbons, baboons, and green monkeys. This comparative analysis underscores the broader applicability of the invention across primate species where the distances and motif are conserved.
[0225] Figure 14 shows the scope of this invention even outside of centromeres to chromosome arms. In particular, Figure 14 (with or without Mb values) shows the visual identification of the correct structure of assembled chromosomes using CENP-B box barcode at chromosome arms, in order to test correct assembly and rapidly identifying structural inversion or misassembled chromosome-arms including mapping the chromosome due to a visual representation of conserved barcode specific to each chromosome arm in the CHM13 reference genome.
[0226] Figure 15 shows the visual identification of correct structure or aberrant variants using CENP-B box barcode at chromosome arms, by showing a comparison of the CENP-B box organization in the chromosome arms within multiple reference genomes. Having the barcode of the CENP-B positioning along the entire sequence of each chromosome arm allows to pinpoint the correct assembled-chromosome structure or aberrant variants. Examples show the DNA motif CENP-B box positioning along chromosome arms identifying a barcode common to all humans that is used as an indicator of aberrant or misassembled chromosome (e.g. chromosome 9 q-arm) and presence of large structural variation (e.g. region of chromosome 10 q-arm translocated in chromosome X).
[0227] Lastly, Figure 16 shows two contigs assigned to the correct chromosome based on pattern, distance and orientation of CENP-B boxes. Data from HGSVC (Logsdon et al., 2024 - Complex genetic variation in nearly complete human genomes). Contigs are publicly available here: https: / / ftp.1000genomes.ebi.ac.uk / vol1 / ftp / data_collections / HGSVC3 / working....
[0228] EXAMPLE 1 : Genome-wide analysis of centromeric functional DNA motifs
[0229] Materials
[0230] No declaration according to article 170 bis CPI is necessary, since in this study the data concerning DNA sequences are available open access online without the need to obtain biological material:
[0231] - CHM13 (T2T-CHM13v2.0) reference genome (Nurk et al., 2022);
[0232] HG38 (GRCh38.p14) reference genome (“https: / / www.ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.40 / ,” n.d.);
[0233] - HG002 (Jarvis et al., 2022);
[0234] - RPE1 reference genome (Volpe et al., 2023);
[0235] - Primates reference genomes (Kuderna et al., 2023).
[0236] Methods
[0237] To interrogate the structure of the human centromeres at base pair resolution, two conserved short sequences and their reverse complement sequences were used which are present in the centromere of each chromosome and refer to the CENP-B box motif (*TTCG****A**CGGG* or *CCCG**T****CGAA*) and pJa motif (SEQ ID NO: 1 TTCCTTTT[C or T]CACC[A or G]TAG or SEQ ID NO:2: CTA[C or T]GGTG[A or G]AAAAGGAA). For the CENP-B box, the whole 17 base pair sequence was not used, allowing mismatches outside the nine biologically assigned functional bases essential for the binding of the CENP-B protein. The goal of the analysis was to investigate the distances between functional motifs within repetitive regions of the human centromeres and in the chromosome arms. fuzznuc.sh (bash)
[0238] # This script uses a tool from EMBOSS that allows users to scan each file in fasta format derived from the fastq to find the query of interest (as for NTTCGNNNNANNCGGGN). Using the parameter -complement you can get the position and the orientation of the query, (+) stands for forward (*TTCG****A**CGGG*) and (-) for complement reverse orientation (*CCCG**T****CGAA*). As described above for the CENP-B box DNA motif is applied using pJa motif as query in forward (+) (SEQ ID NO:1 TTCCTTTT[C or T]CACC[A or G]TAG) and reverse (-) (SEQ ID NO:2: CTA[C or T]GGTG[A or G]AAAAGGAA). convert_table.R ®
[0239] # This R function assign the color blue (#0000CC) to forward orientations and red (#CC0000) to complement reverse directions. Default to black if orientation is neither + nor - (#000000). distance_finder®(R)
[0240] # This R operation uses the package dplyr and takes the table from fuzznuc and calculates the distance between each query in each chromosome. Fuzznuc outputs a table with multiple columns that describes the position of the query within the genomes: chromosome (1 ), start coordinate of the query (2), end coordinate of the query (3) and the orientation (4). The R script considers each row of the table as a query-hit and calculates the distance between each query. The function subtracts the end coordinate of query(n-l ) from the start coordinate of query(n) to calculate the nucleotide distance (fragment length) between each query. Each distance will be reported in a new column called ‘distance’. From the column ‘distance’ it is possible to get several information both for the genome-wide organization of the query and for the chromosome specific organization. You can get this information by using the aggregate function in R choosing the parameter of interest (e.g., retrieve how many queries for each chromosome with a specific orientation and distance and / or the most represented distance genome-wide, without specifying for which chromosome and orientation).
[0241] After the use of distance_finder.R, the Inventors obtain a new table with the following columns: chromosome (1 ), start coordinate of the query (2), end coordinate of the query (3), the orientation (4) and the distance of each query from the previous one (5). pattern_finder.R
[0242] # Having this list of chromosome-specific distances it is possible to apply the concept of K-merizing distances to obtain which is the most represented and specific pattern of each chromosome. The k-patterns are all the subpatterns of length k (k is the number of distances that defines a specific pattern) for a specific query present within a chromosome. Here, the Inventors proceed in k-merizing the distance between our chosen motifs.
[0243] If there is this consecutive distance values between each CENP-B box (referred to as ‘box’):
[0244] - box-497-box-496-box-150-box-151 -box-320-box-149-box-667-box-149-box —
[0245] K-pattern of length: k = 1 would be all the possible occurrences of a single distance value.
[0246] K = 2 will be all pairs formed by two distance values: 497-496, 496-150, 150-151 , 151 -320, 320-149, 149-667, 667-149.
[0247] The same for k = 3, 497-496-150, 496-150-151 , 150-151 -320, 151 -320-149, 320- 149-667, 149-667-149.
[0248] Then the Inventors are able to count the occurrences of each k-pattern to get the most represented and conserved sequence of distances for each chromosome. This operation can be performed in R by using the Inventors own algorithm. Here, the parameter max_n allows users to search the linear combination of 1 to 15 distances (and can be changed according to demand). This outputs a .txt file containing all the possible pattern combinations that occur at least 10 times (which can be changed according to demand) with length 1 to 15.
[0249] By counting the number of patterns and nucleotides you can predict the length of the array occupied by a specific pattern, answering also to the claim of predicting centromere size and locate the functional region within the centromere where the pattern is most stringently repeated. This is important to be calculated rapidly because size changes between individuals and in pathological conditions. reads_finder.R
[0250] # Once each chromosome-specific pattern is characterized, the Inventors use this information to extrapolate chromosome-specific-centromeric reads after long-reads sequencing (current technologies to obtain long-reads are Oxford Nanopore Technologies or PacBio) directly from the fastq file (the raw file containing all the reads sequenced during the experiment). Again, fuzznuc.sh and distance_fmder.R are used to catalogue all the reads that contain the query of interest and the distance values in each read. Then, the Inventors’ own function in R is able to isolate only the reads with a specific pattern (for example: if get only the centromeric reads of chromosome 20 using the information obtained with pattern_finder.R with the specific pattern of distances for chromosome 20). This can be applied to all autosomal chromosomes 1 to 22 and the X.
[0251] # Loop through the ‘distance’ column in the table calculated within the raw reads to find only reads with a specific pattern and that belong to a specific chromosome.
[0252] With this function, you will get a .txt file with all the read IDs with the specific pattern. Then, the Inventors create a catalog with only the reads of interest by filtering the fastq file by read IDs using the tool seqtk. Then, users can use winnowmap to map the reads against the reference genome to validate that the reads belong to the specific centromere of the chromosome of interest, therefore, leading to determine which chromosome the reads belong to, across different human specimens, cell lines and individuals from which DNA can be isolated.
[0253] Having a reliable mapping of these sequences against its own centromere and avoiding an incorrect placement of the reads, centromere assembly and variant calling can be performed with confidence, improving the reliability of the results.
[0254] # Last but not least, using fuzznuc.sh and convert_table.R the position and orientation of each query is mapped genome-wide. This sequence (the CENP-B box motif *TTCG****A**CGGG* or *CCCG**T****CGAA*) appears to establish a distinct pattern of positioning and orientation, emphasizing its significance in occupying the correct chromosome-specific location. Looking at the organization of this sequence genome-wide, we observed a specific barcode (similar to what others have shown using many different SUNKs-like sequences) of our query even in the p-arms and q-arms of the reference genomes available at the time. The considerable advantage of this invention is the use of only two *TTCG****A**CGGG*and *CCCG**T****CGAA* which are conserved across all human individuals and do not need to be identified and re-validated for each genome like for SUNKs. The inventors show that the CENP-B box motif (*TTCG****A**CGGG* or *CCCG**T****CGAA*) builds a chromosome and centromere specific pattern of distances and orientations that reflects the status and the quality of each assembled chromosome, and therefore can be used as an indicator of the genome-wide quality of the assembled genomes or to identify pathological changes in the DNA or chromosomes as shown in Fig. 14).
[0255] Source code
[0256] Code for performing the described analyses is available from Giunta repository on Github (fuzznuc.sh, convert_table.R (R), distance_finder.R (R), pattern_finder.R, reads_finder.R, pattern_search.R).
[0257] Tool
[0258] - http: / / emboss.toulouse.inra.fr / cgi-bin / emboss / fuzznuc
[0259] - https: / / github.com / lh3 / seqtk
[0260] - https: / / github.com / marbl / Winnowmap
[0261] Results
[0262] The Applicant delved into gaining an understanding of the organization of alpha-satellite monomers using consensus motifs known to be specific for the centromere locus. The two motifs that were chosen as landmarks are the CENP-B box motif and the pJa motif which are highly enriched in centromeric and pericentromeric regions. Even if the function is still unknown, pJa protein was discovered in HeLa nuclear extract and binds a 17 base pair (referred to as ‘bp’) motif within centromeric and pericentromeric regions (Gaff et al., 1994). For CENP- B instead the function is known for structural organization of the centromere (Chardon et al., 2022) and binds DNA with the following sequence motif: *TTCG****A**CGGG*. Using fuzznuc, the CHM13 reference genome was scanned for these CENP-B box and pJa protein binding motifs, using the sequence *TTCG****A**CGGG* or *CCCG**T****CGAA* and SEQ ID NO:1 TTCCTTTT[C or T]CACC[A or G]TAG or SEQ ID NO:2: CTA[C or T]GGTG[A or G]AAAAGGAA). Custom R codes written by the Inventors were used to calculate the distance between each motif and search for a consensus distribution of these sequences in each chromosome along the entire reference genome. Here, the newly available genomic loci of the core centromere DNA in the latest human assemblies are probed to obtain a refined and updated view of human centromere sequences at base pair resolution. A previous accepted model called ‘every other monomer scheme’ described CENP-B boxes (Rosandic et al., 2006) regularly distributed in every other monomer unit in dimers which belong to suprachromosomal family 1 (SF1 ) and 2 (SF2) whereas irregularly distributed in dimers belong to suprachromosomal family 3 (SF3). Suprachromosomal families consist in centromere arrays characterized by a specific monomeric organization (dimers, pentamers, etc.) (Alexandrov et al., 1993).
[0263] Profiling and characterization of chromosome-specific organization of the CENP-B box and pJa motif
[0264] To define the exact placement of CENP-B boxes within each centromere at base pair resolution, the annotated position of the CENP-B box (in forward and reverse complement orientation) sequence was extracted within the T2T-CHM13 haploid genome, and fuzznuc was used also to find the motifs positioning in the HG002 diploid genome, GRCh38.p14 and in the latest complete diploid reference genome of the RPE1 cell line (Volpe et al., 2023). To extrapolate information on the position and spacing of all CENP-B boxes regardless of their orientation, the Applicant wrote and devised the custom R code, which allows the calculation of the nucleotide distance genome-wide. This set of computational approaches is collectively named herewith as Genomic Centromere Finder (GCF) that can be used either on complete genome references or within raw reads produced obtained by long-reads sequencing experiments. Firstly, the previous assumption was tested that CENP-B boxes are positioned within the centromere with a frequency of “every other monomer” (Rosandic et al., 2006). This means that there is a near regular alternation between a ~170 bp monomer of alpha-satellite that contains a CENP-B box followed by a monomer that does not have a CENP-B box. Genome-wide mapping of CENP-B boxes on the CHM13 reference genome yielded the distribution of each motif across the whole genome, including HOR and CDR calculated as the fragment length by subtracting the end coordinate from the start coordinate of the previous motif sequence (Fig. 1a-b). To simplify, the resulting value is referred to as ‘distance’ observed between each CENP-B box in the human genomes. Next, the distance values found within the CHM13, HG38, HG002 and RPE1 and present at least >1 % of the time along each chromosome were plotted. In line with previous observations (Rosandic et al., 2006), it was found that not all CENP-B boxes are positioned with an “every other monomer” frequency. Surprisingly, however, CENP-B box distribution was more varied than previously anticipated, with many recurrent distances found across the genome. Importantly, the CENP-B box distances build a pattern that is chromosome-specific. Specifically, four chromosome clusters can be readily identified based on the genetic distribution of the CENP-B box: a recurrent pattern common to chromosome 6, 10, 1 , 12, 5, 16, 7, 19; a pattern that encompasses chromosome 15, 14, 22, 2, 20, 9, 18, 11 ; chromosome s, 8, 4, 13, 21 can be grouped automatically through hierarchical clustering based on their common pattern. Intriguingly, chromosomes 17 and X display a completely unique pattern of distance values (Fig. 1b) (Fig. 1c, e, f) leading to a new cluster. The inventors however discovered that even first and second clusters group the chromosomes that are most representative of the pattern of a CENP-B box every- other-monomer can be discriminated by specific distances enough for each individual chromosome. The most representative distances of these two clusters are precisely 320, 322, 324 nt, which when added to the length of the CENP-B box (17 base pairs) generates fragments of ~337-339-341 nucleotides (referred to as ‘nt’) as previously described by Rosandic et al. in 2006 (Rosandic et al., 2006) and Henikoff et al. in 2015 (Henikoff et al., 2015). The third cluster presents chromosomes with a mix of these distances and therefore clusters alone due to new and highly represented distance values such as 152-153-494-496-662-663 and 666 nt. This cluster lost the main pattern of a box every-other-monomer, showing that the functional CENP-B box can also be present monomer after monomer, once every four monomers or more rarely once every five monomers. Two chromosomes behave differently, allowing the separation of a new small subcluster consisting of chromosome 17 and X. These two chromosomes lost the consensus distribution of a box every-other-monomer, going on to construct a new specific chromosome pattern that arranges the box monomer after monomer (every 148-149-150 nt), rarely every other monomer (every 319-320 nt), a box and then two monomers without it (495-496 nt) and a box and then three or more monomers without it (666 and 833 nt). Given the potential widespread applications of the centromere fingerprinting approach according to the present invention, whether the chromosome-specific consensus distribution of the CENP-B box is maintained across different human cell lines and individuals was tested. Three additional assembled human reference genomes were used, the centromere-incomplete haploid human GRCh38 reference genome (Schneider et al., 2017), the diploid human reference genome of the HG002 cell line (Jarvis et al., 2022) and the diploid RPE1 reference genome (Volpe et al., 2023). The GRCh38 reference genome has an incomplete centromere organization causing the impossibility to distinguish between centromere and pencentromere regions and therefore may highlight variations in this consensus distribution of the functional motif (Miga, 2015). Indeed, the invention allows to see that the GRCh38 reference genome lost the correct CENP-B box distribution (Fig.ld and Fig. 8) especially in the third cluster, whereas the barcode is maintained in the diploid reference genome of the HG002 cell line in both the paternal (Fig.le) and the maternal lineage (Fig.lf), because the HG002 genome is assembled correctly at the centromere. This was possible thanks to a consortium of people that over many years worked to assemble the centromeres of HG002. The same concept to evaluate the CENP-B box distribution was used within the new assembled diploid reference genome of the RPE1 cell line, the first human diploid genome assembled by a single laboratory in less than one year in the Giunta Laboratory (Sapienza, Rome, Italy). The method of the present invention was used to establish the correct assembly of the centromere regions in RPE1 , which in fact reflects the same organization of the CHM13-benchmark reference genome (Fig. 12). Moreover, by plotting the CENP-B box distance found in the GRCh38 reference genome, a different chromosome clustering was observed due to the poor centromere linear sequence annotation of this reference genome, confirming that the distribution of this sequence can be used as a proof of concept to validate a correct assembly of human centromeres in a genome. Conversely, the four clusters are maintained in the maternal and paternal lineage of the HG002 and in the RPE1 genome supporting and emphasizing that the arrangement of the CENP-B box is maintained across different cell lines and individuals, because it likely confers structural fidelity to the centromere through subsequent binding of the CENP-B protein. In the light of this chromosome-specific consensus distribution of the CENP- B box, whether it is maintained even across different primate species was investigated. To evaluate whether the distribution of the CENP-B box is conserved in other human related species, the distribution in different primates was analyzed and its distance was scanned along the entire reference genome of gorilla, chimpanzee, bonobo, orangutan, gibbon, baboon and green monkey considering that these genomes are not fully assembled, the amount of CENP-B box is consistent only in humans, gorilla and chimpanzee. The most represented distances in the human genome are partially shared in the genomes of the primates under consideration, underlining a main close relationship with gorillas and chimpanzees (Fig.13). This reinforces the importance of the correct sequence distribution in maintaining the structural integrity and fidelity of human and other primate centromeres by subsequent protein binding leading to centromere stability and underscore the validity of the invention due to the evolutionary conservation of the sequence motif chosen and its position along the chromosomes.
[0265] Taking advantage of the meticulous characterization of each centromeric region of the CHM13 reference genome, the most represented distances in the active HORs of each chromosome were also investigated to evaluate which pattern remains conserved in the active centromere. The nucleotide distance between each CENP-B box was calculated and the six most represented distances for each chromosome were plotted according to the sequence orientation and normalized for the total number of CENP-B boxes in each array. As shown in Altemose et al. in 2022 (Altemose et al., 2022), the method confirmed the presence of a 1.7 Mb inversion inside active alpha-satellite HOR array D1Z7 of chromosome 1 as shown by the enrichment of both forward and reverse orientation for the CENP-B box sequence (Fig. 2a and Fig. 2b). Thus, the invention allows the most rapid and visually striking representation of changes, aberration and structural variants in the centromere of a chromosome. According to the heatmap clusters, CENP-B box has a predominantly periodicity of 322 nucleotides (nt) even in the active HOR of chromosome 6, 10, 1 , 12, 5, 16, 7, 19 whereas 320, 322, 324 nt is the most represented in the active HOR of chromosome 15, 14, 22, 2, 20, 9, 18, 11 (Fig. 2b). The Inventors identified and leveraged distances other than 320 and 324 nt with unexpected variation in the periodicity between a CENP-B box and the next one in the active array of all chromosomes. Chromosome 17 and X show unique periodicity identified by the Inventors, chromosome 17 displaying enriched values equal to 148 nt, 150 nt and 495 nt, and the latter chromosome X with pattern at 149 nt, 493 nt, 510 nt, and 833 nt, showing that the intact CENP-B box can also be present once every four monomers and more rarely once every five monomers even in the active HOR of the functional centromere. No other enriched periodicity was observed indicating that CENP-B box are precisely spaced within the alpha-satellite active HORs, in line with recent evidence of a structural role for CENP-B binding to centromere in organizing the repetitive DNA and supporting centromeric chromatin replication and function (Chardon et al., 2022). Conversely, pJa behaves differently but add an additional layer of value to the invention and the method because: 1 ) the pJa motif is only interspersed in the active HORs of chromosome 2, 4, 8, 13, 15, 20, 21 and 2) it constructs a different pattern compared to CENP-B box (Fig.2b). In the invention, the pJa motif is used as an additional identifier for specific chromosomes and to confirm centromere pattern in those chromosomes.
[0266] To deeper visualize the distribution of each CENP-B box within the active regions of each centromere, the consensus distance distribution of each CENP-B box in the CDR for each chromosome was constructed, except in chromosome 21 , which has the methylation dip in the centromere transition region that resides outside the active HOR and is devoid of CENP-B boxes. Each line was constructed by plotting the relative distance (y-axis) of the CENP-B box in a specific coordinate of this region (x-axis) (Fig. 3). The trace shows how the distribution of the CENP-B box forms a pattern which is variable along the array but specific to the chromosome alternating CENP-B with different distance values and only partially respecting the pattern of a box every-other-monomer that was previously suggested. This alternation of distance values is lost in the flat spots of the graph, where it follows the 'same distance' scheme, and the CENP-B boxes are distributed approximately at similar distance. Here, the Inventors harnessed the variation between the distances to build a chromosome-specific pattern of distances even when the line looks flatters due to more similar distances. Therefore, the flat line represents regions in which the CENP-B box is distributed with similar but crucially different distance values that coincide with the most represented range of distance 320 nt - 324 nt and with the every-other-monomers scheme. This scheme is highly represented in the CDRs of chromosome 1 , 5, 10, 14, 16, 18, 19 and 22 but still with subtle differences that the Inventors have identified, classified and harnessed pertaining to individual chromosomes (Fig.3). The traces of chromosome 17 and X were also analyzed in which the most represented distance values are almost lost, constructing a completely different and unique pattern for these chromosomes. Zooming in the individual traces of both chromosomes, the unique distribution of the CENP-B box was observed (Fig.3). Along the entire CDR array of chromosome 17, the CENP-B box is arranged with the well-defined block of distances 496-495-149-150-319-148- 666-148 base pair repeated ~61 times, describing a pattern which has lost the expected and previously described periodicity of one CENP-B box every-other- monomer but displays other distances which potentially lead to place CENP-B protein every monomer, every two monomers or every three monomers as uncovered by the Inventors (Fig. 3). Importantly, chromosome X has completely lost the pattern of the CENP-B box every-other-monomer and gained new placements of this sequence. The distance values building the block in chromosome X are 149- 833-493-510 nt and are repeated ~47 times in the CDR array (Fig. 3). While the distances are conserved, the number of possible repetitions varies between genomes so this value is not a suitable reference as not universal. The distances, on the other hand, can be rapidly used to determine the minimum and maximum number of repetitions for each pattern.
[0267] Knowledge on the distribution of functional motifs in the centromeric array led to the first complete and detailed characterization of the distances in the human genome by the Inventors that can be used for more accurate centromere analyses, allowing the fishing of reads that belong to a specific chromosome, avoiding the problem of multimapping and incorrect assignment of reads in these highly repeated regions. With the method of the present invention, it is now possible to interrogate every read produced by long-read sequencing and investigate the presence of a chromosome-specific pattern therein. Since the pattern is maintained across chromosomes, the Inventors show that the read containing a specific pattern maps exclusively to the chromosome to which it belongs (Fig.4). Taking advantage of the long read nanopore sequencing performed on the RPE1 cell line, achieving 80X coverage, the reads were interrogated by using the combination of fuzznuc and the custom R codes as described in the methods of the invention to extrapolate CENP- B box coordinates and calculate the genomic distances. Each raw read with a specific chromosome pattern was cataloged into a new pool of reads that represent a specific chromosome pattern. These reads were mapped against the entire reference genome (CHM13v2.0) expecting the mapping only in the chromosome with that pattern to confirm the centromere barcode of distances to be a suitable tool for isolation of chromosome-specific reads. This allows a collection of centromeric reads for each chromosome, and therefore when they were mapped, they were expected to map only to the chromosome containing that functional motif distribution pattern. This filtering method greatly increases the reliability of the reads that are mapped to the centromere, in fact, while the reads were given the possibility of mapping in each chromosome, the Inventors demonstrate that they continue to map in the chromosome where the characteristic pattern is present and in no other chromosome (Fig. 4a-f, Fig.9 and Fig. 11), indicating the specificity and reliability of the methods hereby presented. This process can also be applied to drastically improve the performance in the variant calling procedure, leading to a more accurate calling of genomic variant such as insertions and deletions in the centromere arrays as the Inventors have demonstrated (Fig. 5a-f, Fig. 6a-g and Fig. 10).
[0268] The characterization of each chromosome-pattern and its content allowed to compare the amount of each pattern against the spread of the CENP-A protein within the centromere arrays. Taking the example of chromosome 6 and 10, it was observed that the amount of the patterns is proportional to the spread of the protein, which is strictly associated with the active HORs. The HG38 reference genome, which has an incomplete organization of the centromere arrays, showed that the amount of the pattern is always less compared to the other genomes, and that is proportionally correlated with the spread of the CENP-A protein. Hence, the method can be used as a proxy to infer and calculate the size of the active centromere regions that is functionally relevant for kinetochore attachment and chromosome segregation even without epigenetic, protein information or the need for experiments using CENP-A or other proteins but simply estimating the mode of the most repeated pattern of distances.
[0269] With all these methods described in the Genomic Centromere Finder (GCF) collection, it is possible to analyze, study, evaluate and characterize the human centromeres from different points of view. Taking advantage of ultra-long sequences any laboratories can sequence any cell lines, biological specimen or patient sample and extrapolate information on the centromere by simply visualizing whether the status of the chromosome-specific patterns using our identified sequence motifs, distances, orientation and distribution in centromeres and along chromosomes is maintained as hereby described and whether this region is affected under different aberrant conditions, hence differing from the normal organization discovered by the Inventors and described in all details in this invention. Additionally, this precise organization allows to discriminate sequences belonging to both chromosomes from an assembled genome as well as from a pool of raw sequences from any human. Furthermore, the sequence analysis can be used to call variants in the human centromeres, assemble the complete human genome of different cell lines or individuals, and help the resolution of highly repetitive DNA sequences using these "anchor" sequences whose we established chromosome-specific patterns.
[0270] On the basis of the above, the information of the chromosome specific distances of DNA motifs and computational tools according to the present invention can be advantageously used for the following applications:
[0271] (1 ) Codifying the genetic and genomic structure of human centromeres based on our novel finding the most represented distance of a specific DNA motif;
[0272] (2) Extracting centromere long DNA sequencing reads across different human specimens, cell lines and individuals from which DNA can be isolated;
[0273] (3) Determining which chromosome the reads belong to;
[0274] (4) Validation method for centromere sequences and complete human genome assemblies;
[0275] (5) Guiding correct assembly of human centromeres and human genomes;
[0276] (6) Determining structural variations (SVs), insertion / deletions (INDELS), DNA aberrations and other changes in centromere and chromosomal DNA including of pathological origins, undetectable with current sequencing and diagnostic or other methods;
[0277] (7) Predicting active and functional centromere size based on mode of distances for each human chromosome for each individual automatically extracted with our method;
[0278] (8) Rapid flagging of incorrectly assembled regions of the human genome, of pathologically mutated and other human centromere aberrations at the DNA sequence level.
[0279] References
[0280] Alexandrov IA, Medvedev LI, Mashkova TD, Kisselev LL, Romanova LY, Yurov YB. 1993. Definition of a new alpha satellite suprachromosomal family characterized by monomeric organization. Nucleic Acids Res 21:2209- 2215.
[0281] Altemose N, Logsdon GA, Bzikadze AV, Sidhwani P, Langley SA, Caldas GV, Hoyt SJ, Uralsky L, Ryabov FD, Shew CJ, Sauria MEG, Borchers M, Gershman A, Mikheenko A, Shepelev VA, Dvorkina T, Kunyavskaya O, Vollger MR, Rhie A, McCartney AM, Ash M, Lorig-Roach R, Shafin K, Lucas JK, Aganezov S, Olson D, de Lima LG, Potapova T, Hartley GA, Haukness M, Kerpedjiev P, Gusev F, Tigyi K, Brooks S, Young A, Nurk S, Koren S, Salama SR, Paten B, Rogaev El, Streets A, Karpen GH, Dernburg AF, Sullivan BA, Straight AF, Wheeler TJ, Gerton JL, Eichler EE, Phillippy AM, Timp W, Dennis MY, O’Neill RJ, Zook JM, Schatz MC, Pevzner PA, Diekhans M, Langley CH, Alexandrov IA, Miga KH. 2022. Complete genomic and epigenetic maps of human centromeres. Science 376:eabl4178. doi: 10.1126 / science.abl4178
[0282] Black EM, Giunta S. 2018. Repetitive Fragile Sites: Centromere Satellite DNA As a Source of Genome Instability in Human Diseases. Genes 9:615. doi: 10.3390 / genes9120615
[0283] Blower MD, Sullivan BA, Karpen GH. 2002. Conserved organization of centromeric chromatin in flies and humans. Dev Cell 2:319-330. doi: 10.1016 / s1534-5807(02)00135-1
[0284] Bodor DL, Mata JF, Sergeev M, David AF, Salimian KJ, Panchenko T, Cleveland DW, Black BE, Shah JV, Jansen LE. 2014. The quantitative architecture of centromeric chromatin. eLife 3:e02137. doi:10.7554 / eLife.02137
[0285] Bzikadze AV, Pevzner PA. 2020. Automated assembly of centromeres from ultralong error-prone reads. Nat Biotechnol 38:1309-1316. doi: 10.1038 / s41587- 020-0582-4
[0286] Chardon F, Japaridze A, Witt H, Velikovsky L, Chakraborty C, Wilhelm T, Dumont M, Yang W, Kikuti C, Gangnard S, Mace A-S, Wuite G, Dekker C, Fachinetti D. 2022. CENP-B-mediated DNA loops regulate activity and stability of human centromeres. Mol Cell 82: 1751 -1767. e8. doi:10.1016 / j.molcel.2022.02.032
[0287] Dubocanin D, Cortes AES, Hartley GA, Ranchalis J, Agarwal A, Logsdon GA, Munson KM, Real T, Mallory BJ, Eichler EE, O’Neill RJ, Stergachis AB. 2023. Conservation of chromatin organization within human and primate centromeres, doi: 10.1101 / 2023.04.20.537689
[0288] Gaff C, du Sart D, Kalitsis P, lannello R, Nagy A, Choo KH. 1994. A novel nuclear protein binds centromeric alpha satellite DNA. Hum Mol Genet 3:711-716. doi:10.1093 / hmg / 3.5.711
[0289] Gershman A, Sauria MEG, Guitart X, Vollger MR, Hook PW, Hoyt SJ, Jain M, Shumate A, Razaghi R, Koren S, Altemose N, Caldas GV, Logsdon GA, Rhie A, Eichler EE, Schatz MC, O’Neill RJ, Phillippy AM, Miga KH, Timp W. 2022. Epigenetic patterns in a complete human genome. Science 376:eabj5089. doi: 10.1126 / science.abj5089
[0290] Giunta S, Herve S, White RR, Wilhelm T, Dumont M, Scelfo A, Gamba R, Wong CK, Rancati G, Smogorzewska A, Funabiki H, Fachinetti D. 2021. CENP-A chromatin prevents replication stress at centromeres to avoid structural aneuploidy. Proc Natl Acad Sci U S A 118:e2015634118. doi: 10.1073 / pnas.2015634118
[0291] Glynn M, Kaczmarczyk A, Prendergast L, Quinn N, Sullivan KF. 2010. Centromeres: assembling and propagating epigenetic function. Subcell Biochem 50:223-249. doi: 10.1007 / 978-90-481 -3471 -7_12 Henikoff JG, Thakur J, Kasinathan S, Henikoff S. 2015. A unique chromatin complex occupies young a-satellite arrays of human centromeres. Sci Adv 1 :e1400234. doi: 10.1126 / sciadv.1400234 https: / / www.ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.40 / . n.d. . NCBI. https: / / www.ncbi.nlm.nih.gov / data-hub / assembly / GCF_000001405.40 /
[0292] Jarvis ED, Formenti G, Rhie A, Guarracino A, Yang C, Wood J, Tracey A, Thibaud-Nissen F, Vollger MR, Porubsky D, Cheng H, Ash M, Logsdon GA, Carnevali P, Chaisson MJP, Chin C-S, Cody S, Collins J, Ebert P, Escalona M, Fedrigo O, Fulton RS, Fulton LL, Garg S, Gerton JL, Ghurye J, Granat A, Green RE, Harvey W, Hasenfeld P, Hastie A, Haukness M, Jaeger EB, Jain M, Kirsche M, Kolmogorov M, Korbel JO, Koren S, Korlach J, Lee J, Li D, Lindsay T, Lucas J, Luo F, Marschall T, Mitchell MW, McDaniel J, Nie F, Olsen HE, Olson ND, Pesout T, Potapova T, Puiu D, Regier A, Ruan J, Salzberg SL, Sanders AD, Schatz MC, Schmitt A, Schneider VA, Selvaraj S, Shafin K, Shumate A, Stitziel NO, Stober C, Torrance J, Wagner J, Wang J, Wenger A, Xiao C, Zimin AV, Zhang G, Wang T, Li H, Garrison E, Haussler D, Hall I, Zook JM, Eichler EE, Phillippy AM, Paten B, Howe K, Miga KH, Human Pangenome Reference Consortium. 2022. Semiautomated assembly of high-quality diploid human reference genomes. Nature 611 :519-531 . doi: 10.1038 / s41586-022-05325-5
[0293] Kuderna LFK, Ulirsch JC, Rashid S, Ameen M, Sundaram L, Hickey G, Cox AJ, Gao H, Kumar A, Aguet F, Christmas MJ, Clawson H, Haeussler M, Janiak MC, Kuhlwilm M, Orkin JD, Bataillon T, Manu S, Valenzuela A, Bergman J, Rouselle M, Silva FE, Agueda L, Blanc J, Gut M, de Vries D, Goodhead I, Harris RA, Raveendran M, Jensen A, Chuma IS, Horvath JE, Hvilsom C, Juan D, Frandsen P, Schraiber JG, de Melo FR, Bertuol F, Byrne H, Sampaio I, Farias I, Valsecchi J, Messias M, da Silva MNF, Trivedi M, Rossi R, Hrbek T, Andriaholinirina N, Rabarivola CJ, Zaramody A, Jolly CJ, Phillips-Conroy J, Wilkerson G, Abee C, Simmons JH, Fernandez-Duque E, Kanthaswamy S, Shiferaw F, Wu D, Zhou L, Shao Y, Zhang G, Keyyu JD, Knauf S, Le MD, Lizano E, Merker S, Navarro A, Nadler T, Khor CC, Lee J, Tan P, Lim WK, Kitchener AC, Zinner D, Gut I, Melin AD, Guschanski K, Schierup MH, Beck RMD, Karakikes I, Wang KC, Umapathy G, Roos C, Boubli JP, Siepel A, Kundaje A, Paten B, Lindblad-Toh K, Rogers J, Marques Bonet T, Farh KK-H. 2023. Identification of constrained sequence elements across 239 primate genomes. Nature 1-8. doi: 10.1038 / s41586- 023-06798-8
[0294] Liao W-W, Ash M, Ebler J, Doerr D, Haukness M, Hickey G, Lu S, Lucas JK, Monlong J, Abel HJ, Buonaiuto S, Chang XH, Cheng H, Chu J, Colonna V, Eizenga JM, Feng X, Fischer C, Fulton RS, Garg S, Groza C, Guarracino A, Harvey WT, Heumos S, Howe K, Jain M, Lu T-Y, Markello C, Martin FJ, Mitchell MW, Munson KM, Mwaniki MN, Novak AM, Olsen HE, Pesout T, Porubsky D, Prins P, Sibbesen JA, Siren J, Tomlinson C, Villani F, Vollger MR, Antonacci-Fulton LL, Baid G, Baker CA, Belyaeva A, Billis K, Carroll A, Chang P-C, Cody S, Cook DE, Cook-Deegan RM, Cornejo OE, Diekhans M, Ebert P, Fairley S, Fedrigo O, Felsenfeld AL, Formenti G, Frankish A, Gao Y, Garrison NA, Giron CG, Green RE, Haggerty L, Hoekzema K, Hourlier T, Ji HP, Kenny EE, Koenig BA, Kolesnikov A, Korbel JO, Kordosky J, Koren S, Lee H, Lewis AP, Magalhaes H, Marco-Sola S, Marijon P, McCartney A, McDaniel J, Mountcastle J, Nattestad M, Nurk S, Olson ND, Popejoy AB, Puiu D, Rautiainen M, Regier AA, Rhie A, Sacco S, Sanders AD, Schneider VA, Schultz Bl, Shafin K, Smith MW, Sofia HJ, Abou Tayoun AN, Thibaud-Nissen F, Tricomi FF, Wagner J, Walenz B, Wood JMD, Zimin AV, Bourque G, Chaisson MJP, Flicek P, Phillippy AM, Zook JM, Eichler EE, Haussler D, Wang T, Jarvis ED, Miga KH, Garrison E, Marschall T, Hall IM, Li H, Paten B. 2023. A draft human pangenome reference. Nature 617:312-324. doi:10.1038 / s41586-023-05896-x
[0295] Logsdon GA, Vollger MR, Hsieh P, Mao Y, Liskovykh MA, Koren S, Nurk S, Mercuri L, Dishuck PC, Rhie A, de Lima LG, Dvorkina T, Porubsky D, Harvey WT, Mikheenko A, Bzikadze AV, Kremitzki M, Graves-Lindsay TA, Jain C, Hoekzema K, Murali SC, Munson KM, Baker C, Sorensen M, Lewis AM, Surti U, Gerton JL, Larionov V, Ventura M, Miga KH, Phillippy AM, Eichler EE. 2021 . The structure, function and evolution of a complete human chromosome 8. Nature 593:101-107. doi: 10.1038 / s41586-021- 03420-7
[0296] Masumoto H, Masukata H, Muro Y, Nozaki N, Okazaki T. 1989. A human centromere antigen (CENP-B) interacts with a short specific sequence in alphoid DNA, a human centromeric satellite. J Cell Biol 109:1963-1973. doi: 10.1083 / jcb.109.5.1963
[0297] Masumoto H, Yoda K, Ikeno M, Kitagawa K, Muro Y, Okazaki T. 1993. Properties of CENP-B and Its Target Sequence in a Satellite DNA In: Vig BK, editor. Chromosome Segregation and Aneuploidy, NATO ASI Series. Berlin, Heidelberg: Springer, pp. 31-43. doi: 10.1007 / 978-3-642-84938-1 _3
[0298] Miga KH. 2015. Completing the human genome: the progress and challenge of satellite DNA assembly. Chromosome Res Int J Mol Supramol Evol Asp Chromosome Biol 23:421-426. doi: 10.1007 / s10577-015-9488-2
[0299] Nurk S, Koren S, Rhie A, Rautiainen M, Bzikadze AV, Mikheenko A, Vollger MR, Altemose N, Uralsky L, Gershman A, Aganezov S, Hoyt SJ, Diekhans M, Logsdon GA, Alonge M, Antonarakis SE, Borchers M, Bouffard GG, Brooks SY, Caldas GV, Chen N-C, Cheng H, Chin C-S, Chow W, de Lima LG, Dishuck PC, Durbin R, Dvorkina T, Fiddes IT, Formenti G, Fulton RS, Fungtammasan A, Garrison E, Grady PGS, Graves-Lindsay TA, Hall IM, Hansen NF, Hartley GA, Haukness M, Howe K, Hunkapiller MW, Jain C, Jain M, Jarvis ED, Kerpedjiev P, Kirsche M, Kolmogorov M, Korlach J, Kremitzki M, Li H, Maduro VV, Marschall T, McCartney AM, McDaniel J, Miller DE, Mullikin JC, Myers EW, Olson ND, Paten B, Peluso P, Pevzner PA, Porubsky D, Potapova T, Rogaev El, Rosenfeld JA, Salzberg SL, Schneider VA, Sedlazeck FJ, Shafin K, Shew CJ, Shumate A, Sims Y, Smit AFA, Soto DC, Sovic I, Storer JM, Streets A, Sullivan BA, Thibaud-Nissen F, Torrance J, Wagner J, Walenz BP, Wenger A, Wood JMD, Xiao C, Yan
[0300] SM, Young AC, Zarate S, Surti U, McCoy RC, Dennis MY, Alexandrov IA, Gerton JL, O’Neill RJ, Timp W, Zook JM, Schatz MC, Eichler EE, Miga KH, Phillippy AM. 2022. The complete sequence of a human genome. Science 376:44-53. doi: 10.1126 / science.abj6987
[0301] Rosandic M, Paar V, Basar I, Gluncic M, Pavin N, Pilas I. 2006. CENP-B box and pJalpha sequence distribution in human alpha satellite higher-order repeats (HOR). Chromosome Res Int J Mol Supramol Evol Asp Chromosome Biol 14:735-753. doi: 10.1007 / s10577-006-1078-x
[0302] Schneider VA, Graves-Lindsay T, Howe K, Bouk N, Chen H-C, Kitts PA, Murphy
[0303] TD, Pruitt KD, Thibaud-Nissen F, Albracht D, Fulton RS, Kremitzki M, Magrini V, Markovic C, McGrath S, Steinberg KM, Auger K, Chow W, Collins J, Harden G, Hubbard T, Pelan S, Simpson JT, Threadgold G, Torrance J, Wood JM, Clarke L, Koren S, Boitano M, Peluso P, Li H, Chin C-S, Phillippy AM, Durbin R, Wilson RK, Flicek P, Eichler EE, Church DM. 2017. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. Genome Res 27:849-864. doi: 10.1101 Zgr.213611.116
[0304] Volpe E, Corda L, Di Tommaso E, Pelliccia F, Ottalevi R, Licastro D, Formenti G, Capulli M, Guarracino A, Tassone E, Giunta S. 2023. The complete human diploid reference genome of RPE-1 identifies the phased epigenetic landscapes from multi-omics data, doi: 10.1101 / 2023.11 .01 .565049
[0305] Wevrick R, Willard HF. 1989. Long-range organization of tandem arrays of alpha satellite DNA at the centromeres of human chromosomes: high-frequency array-length polymorphism and meiotic stability. Proc Natl Acad Sci 86:9394-9398. doi: 10.1073 / pnas.86.23.9394
Claims
CLAIMS1 ) Use of a DNA motif as a marker for characterizing, identifying, isolating, assembling or validating the nucleotide sequence of the centromere of a human chromosome, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or *CCCG**T****CGAA*; or a pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2); said motif being repeated several times along the nucleotide sequence of the centromere according to a pattern of distances between subsequent motifs, said pattern being repeated several times along the nucleotide sequence of the centromere and being specific for each chromosome, wherein the distance is the number of nucleotides between the last nucleotide of a motif and the first nucleotide of the subsequent motif.2) Use according to claim 1 , wherein the motif *TTCG****A**CGGG* is the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4), and / or the motif *CCCG**T****CGAA* is the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO:6).3) Use according to any one of claims 1 -2, wherein the pattern of distances between subsequent CENP-B box protein binding motifs within the centromere is- chosen among (322 ± 1 )nnucleotides, (323 ± 1 )nnucleotides or (324 ± 1 )nnucleotides, preferably (323 ± 1 )nnucleotides, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 5, in the nucleotide sequence of the centromere of chromosome 1 ;- 321 ± 1 , 325 ± 1 nucleotides in the nucleotide sequence of the centromere of chromosome 2;- 497 ± 1 , 321 ± 1 , 324 ± 1 , 324 ± 1 , 322 ± 1 , 321 ± 1 and 663 ± 1 in the nucleotide sequence of the centromere of chromosome 3;- 664 ± 1 , 153 ± 1 , 323 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 491 ± 1 and 324 ± 1 in the nucleotide sequence of the centromere of chromosome 4;- 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 in the nucleotide sequence of the centromere of chromosome 5;- 665 ± 1 , 323 ± 1 , 323 ± 1 , 320 ± 1 , 153 ± 1 , 323 ± 1 and 832 ± 1 in the nucleotide sequence of the centromere of chromosome 6;- 325 ± 1 , 323 ± 1 , 323 ± 1 and 325 ± 1 in the nucleotide sequence of the centromere of chromosome 7;- 321 ± 1 , 154 ± 1 and 667 ± 1 in the nucleotide sequence of the centromere of chromosome 8;- 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 and 495 ± 1 in the nucleotide sequence of the centromere of chromosome 9;- 323 ± 1 , 323 ± 1 , 322 ± 1 and 321 ± 1 in the nucleotide sequence of the centromere of chromosome 10;- 321 ± 1 , 495 ± 1 in the nucleotide sequence of the centromere of chromosome 11 ;- 323 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 in the nucleotide sequence of the centromere of chromosome 12;- chosen between a first pattern 323 ± 1 , 324 ± 1 and 492 ± 1 and a second pattern 325 ± 1 , 321 ± 1 , 323 ± 1 , 324 ± 1 and 492 ± 1 in the nucleotide sequence of the centromere of chromosome 13;- 321 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 and 321 ± 1 n in the nucleotide sequence of the centromere of chromosome 14;- 154 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 and 325 ± 1 in the nucleotide sequence of the centromere of chromosome 15;- 322 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 in the nucleotide sequence of the centromere of chromosome 16;- 497 ± 1 , 496 ± 1 , 150 ± 1 , 151 ± 1 , 320 ± 1 , 149 ± 1 , 667 ± 1 and 149 ± 1 in the nucleotide sequence of the centromere of chromosome 17;- 321 ± 1 , 320 ± 1 , 321 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 in the nucleotide sequence of the centromere of chromosome 18;- chosen among (323 ± 1)nnucleotides, (322 ± 1)nnucleotides or (663 ± 1)nnucleotides, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 5, for example at least 6, at least 7, at least 8, at least 9 or at least 10, in the nucleotide sequence of the centromere of chromosome 19;- 663 ± 1 , 324 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 in the nucleotide sequence of the centromere of chromosome 20;- 492 ± 1 , 325 ± 1 , 321 ± 1 , 323 ± 1 , 324 ± 1 in the nucleotide sequence of the centromere of chromosome 21 ;- 321 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 in the nucleotide sequence of the centromere of chromosome 22;- chosen between 834 ± 1 , 494 ± 1 , 511 ± 1 , 150 ± 1 or 834 ± 1 , 150 ± 1 , 511 ± 1 , 494 ± 1 in the nucleotide sequence of the centromere of chromosome X. 4) Use according to anyone of claims 1 -3, wherein the pattern of distances between subsequent pJ-alpha (pJa) DNA motifs is:- 321 ± 1 and 325 ± 1 nucleotides in the nucleotide sequence of the centromere of chromosome 2;- 832 ± 1 , 664 ± 1 , 1173 ± 1 , 495 ± 1 in the nucleotide sequence of the centromere of chromosome 4;- chosen between 320 ± 1 , 1002 ± 1 or 320 ± 1 , 2195 ± 1 in the nucleotide sequence of the centromere of chromosome 8;- chosen between a first pattern (1173 ± 1 , 1853 ± 1 )n or a second pattern (1173 ± 1 )n, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively, in the nucleotide sequence of the centromere of chromosome 13;- 838 ± 1 , 1005 ± 1 , 663 ± 1 in the nucleotide sequence of the centromere of chromosome 15;- chosen between a first pattern (2702 ± 1 )n or a second pattern (1343 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 20;- (1853 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 21 , wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively;- chosen between a first pattern (498 ± 1 , 153 ± 1 )n or a second pattern (668 ± 1 , 2725 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 22, wherein n is the number of repetition of said distance valuesin tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively.5) Method for characterizing, identifying, isolating, assembling or validating the nucleotide sequence of a centromere of a human chromosome, said method comprising the following steps: a) identifying in an assembled genome or in raw sequencing long-reads or in contigs a motif, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or *CCCG**T****CGAA*; or pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2); b) identifying and isolating a centromere of a chromosome by selecting in the assembled genome or in the raw sequencing long reads or in the contigs, a nucleotide sequence wherein said motif is repeated along the nucleotide sequence according to a pattern of distances, between subsequent motifs, said pattern being specific for each chromosome, wherein the distance is the number of nucleotides between the last nucleotide of a motif and the first nucleotide of the subsequent motif; and optionally c) assembling the nucleotide sequences selected from the raw sequencing long reads or the contigs in step b) or, alternatively, validating the nucleotide sequences selected from the assembled genome in step b), wherein the patterns are repeated as defined in step b) for each chromosome.6) Method according to claim 5, wherein the motif *TTCG****A**CGGG* is the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4), and / or the motif *CCCG**T****CGAA* is the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO:6).7) Method according to any one of claims 5-6, wherein the pattern of distances of step b) between subsequent CENP-B box protein binding motifs within the centromere is chosen among (322 ± 1 )nnucleotides, (323 ± 1 )nnucleotides or (324 ± 1 )nnucleotides, preferably (323 ± 1)nnucleotides, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 5, in the nucleotide sequence of the centromere of chromosome 1 ;- 321 ± 1 , 325 ± 1 nucleotides in the nucleotide sequence of the centromere of chromosome 2;- 497 ± 1 , 321 ± 1 , 324 ± 1 , 324 ± 1 , 322 ± 1 , 321 ± 1 and 663 ± 1 in the nucleotide sequence of the centromere of chromosome 3;- 664 ± 1 , 153 ± 1 , 323 ± 1 , 321 ± 1 , 325 ± 1 , 495 ± 1 , 491 ± 1 and 324 ± 1 in the nucleotide sequence of the centromere of chromosome 4;- 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 in the nucleotide sequence of the centromere of chromosome 5;- 665 ± 1 , 323 ± 1 , 323 ± 1 , 320 ± 1 , 153 ± 1 , 323 ± 1 and 832 ± 1 in the nucleotide sequence of the centromere of chromosome 6;- 325 ± 1 , 323 ± 1 , 323 ± 1 and 325 ± 1 in the nucleotide sequence of the centromere of chromosome 7;- 321 ± 1 , 154 ± 1 and 667 ± 1 in the nucleotide sequence of the centromere of chromosome 8;- 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 and 495 ± 1 in the nucleotide sequence of the centromere of chromosome 9;- 323 ± 1 , 323 ± 1 , 322 ± 1 and 321 ± 1 in the nucleotide sequence of the centromere of chromosome 10;- 321 ± 1 , 495 ± 1 in the nucleotide sequence of the centromere of chromosome 11 ;- 323 ± 1 , 323 ± 1 , 323 ± 1 and 322 ± 1 in the nucleotide sequence of the centromere of chromosome 12;- chosen between a first pattern 323 ± 1 , 324 ± 1 and 492 ± 1 and a second pattern 325 ± 1 , 321 ± 1 , 323 ± 1 , 324 ± 1 and 492 ± 1 in the nucleotide sequence of the centromere of chromosome 13;- 321 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 and 321 ± 1 n in the nucleotide sequence of the centromere of chromosome 14;- 154 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 and 325 ± 1 in the nucleotide sequence of the centromere of chromosome 15;- 322 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 , 323 ± 1 in the nucleotide sequence of the centromere of chromosome 16;- 497 ± 1 , 496 ± 1 , 150 ± 1 , 151 ± 1 , 320 ± 1 , 149 ± 1 , 667 ± 1 and 149 ± 1 in the nucleotide sequence of the centromere of chromosome 17;- 321 ± 1 , 320 ± 1 , 321 ± 1 , 325 ± 1 , 325 ± 1 , 321 ± 1 in the nucleotide sequence of the centromere of chromosome 18;- chosen among (323 ± 1 )nnucleotides, (322 ± 1 )nnucleotides or (663 ± 1 )nnucleotides, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 5, for example at least 6, at least 7, at least 8, at least 9 or at least 10, in the nucleotide sequence of the centromere of chromosome 19;- 663 ± 1 , 324 ± 1 , 321 ± 1 , 325 ± 1 , 321 ± 1 , 321 ± 1 , 325 ± 1 in the nucleotide sequence of the centromere of chromosome 20;- 492 ± 1 , 325 ± 1 , 321 ± 1 , 323 ± 1 , 324 ± 1 in the nucleotide sequence of the centromere of chromosome 21 ;- 321 ± 1 , 325 ± 1 , 325 ± 1 , 325 ± 1 in the nucleotide sequence of the centromere of chromosome 22;- chosen between 834 ± 1 , 494 ± 1 , 511 ± 1 , 150 ± 1 or 834 ± 1 , 150 ± 1 , 511 ± 1 , 494 ± 1 in the nucleotide sequence of the centromere of chromosome X. 8) Method according to anyone of claims 5-7, wherein the pattern of distances of step b) between subsequent pJ-alpha (pJa) DNA motifs is:- 321 ± 1 and 325 ± 1 nucleotides in the nucleotide sequence of the centromere of chromosome 2;- 832 ± 1 , 664 ± 1 , 1173 ± 1 , 495 ± 1 in the nucleotide sequence of the centromere of chromosome 4;- chosen between 320 ± 1 , 1002 ± 1 or 320 ± 1 , 2195 ± 1 in the nucleotide sequence of the centromere of chromosome 8;- chosen between a first pattern (1173 ± 1 , 1853 ± 1 )n or a second pattern (1173 ± 1 )n, wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively, in the nucleotide sequence of the centromere of chromosome 13;- 838 ± 1 , 1005 ± 1 , 663 ± 1 in the nucleotide sequence of the centromere of chromosome 15;- chosen between a first pattern (2702 ± 1 )n or a second pattern (1343 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 20;- (1853 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 21 , wherein n is the number of repetition of said distance value in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively;- chosen between a first pattern (498 ± 1 , 153 ± 1 )n or a second pattern (668 ± 1 , 2725 ± 1 )n in the nucleotide sequence of the peri / centromere of chromosome 22, wherein n is the number of repetition of said distance values in tandem in said pattern and n is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10, i.e. the pattern is repeated at least 3 at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 times consecutively.9) Method according to any one of claim 5-8, wherein said method is a computer-implemented method and said steps a), b) and c) are carried out by a processor.10) Method for identifying possible pathologic changes in the centromere nucleotide sequence of a human chromosome of a subject, said method comprising the following steps: a) identifying in the chromosome nucleotide sequence of said subject a motif, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or *CCCG**T****CGAA*; or pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2); b) identifying the position of said motif in the chromosome nucleotide sequence; c) comparing the motifs, orientation and the positions thereof in the chromosome nucleotide sequence of the subject with the motifs, orientation and positions thereof in a reference nucleotide sequence wherein when a position of the motif, the orientation of the motif and / or the motif in the chromosome nucleotide sequence of the subject is different from the position and / orthe orientation of the same motif and / or from the motif itself in the reference chromosome nucleotide sequence a possible pathologic change of the centromere nucleotide sequence of the chromosome is present.11 ) Method according to claim 10, wherein the motif *TTCG****A**CGGG* is the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4), and / or the motif *CCCG**T****CGAA* is the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO:6).12) Method according to any one of claims 10-11 , wherein, when the change in the centromere nucleotide sequence of a chromosome of a subject is a structural inversion, said structural inversion is present when in the same position in the centromere nucleotide sequence of the subject the motif is *TTCG****A**CGGG*, whereas in the reference centromere nucleotide sequence the motif is *CCCG**T****CGAA* or vice versa, and / or when in the same position in the centromere nucleotide sequence of the subject the motif is SEQ ID NO:1 , whereas in the reference centromere nucleotide sequence the motif is SEQ ID NO:2 or vice versa.13) Method according to any one of claims 10-12, wherein said method is a computer-implemented method and said steps a), b) and c) are carried out by a processor.14) Method for calculating the size of the centromere nucleotide sequence of a human chromosome, preferably of the active centromere nucleotide sequence, said method comprising the step of: a) identifying a motif in the centromere nucleotide sequence, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or *CCCG**T****CGAA*; or pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2), said motif being repeated several times along the nucleotide sequence of the centromere according to a pattern of distances between subsequent motifs; b) identifying the pattern of distances of said motif, wherein said patterncomprises a number K of distance values, said pattern being repeated several times along the nucleotide sequence of the centromere and being specific for each chromosome, wherein the distance is the number of nucleotides between the last nucleotide of a motif and the first nucleotide of the subsequent motif; c) calculating the frequency with which the pattern is repeated in the centromere nucleotide sequence and d) calculating the size of the centromere nucleotide sequence by the following formula: centromere size (bp) = {(sum of distance values of the pattern (bp)) + [(17bp x k) + 17bp]} x Frequency.15) Method according to claim 14, wherein the motif *TTCG****A**CGGG* is the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4), and / or the motif *CCCG**T****CGAA* is the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO:6).16) Method according to any one of claims 14-15, wherein said method is a computer-implemented method and said steps a), b), c) and d) are carried out by a processor.17) Use of a DNA motif as a marker for characterizing, identifying, isolating, assembling or validating the nucleotide sequence of a human chromosome, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or *CCCG**T****CGAA*; or a pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2); said motif being repeated several times along the nucleotide sequence of the chromosome in conservated positions.18) Use according to claim 17, wherein the motif *TTCG****A**CGGG* is the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4), and / or the motif *CCCG**T****CGAA* is the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO:6).19) Method for characterizing, identifying, isolating, assembling or validatingthe nucleotide sequence and structural organization of a chromosome, said method comprising the following steps: a) identifying, in an assembled genome or in raw sequencing long-reads or in contigs, one or more motifs, wherein said one or more motifs are: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* or its reverse complement *CCCG**T****CGAA*; and / or> pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2), wherein said one or more motifs are repeated along the same strand of the nucleotide sequence of a chromosome according to a pattern of positions, said pattern being specific for each human chromosome; b) visually comparing the pattern of positions of said one or more motifs in said assembled genome or said raw sequencing long-reads or said contigs with the pattern of positions of the same motifs in a reference chromosome, wherein said assembled genome or said raw sequencing long-reads or said contigs belong to the chromosome having a pattern which is visually superimposable with their pattern; and optionally c) assembling the nucleotide the raw sequencing long reads or contigs or, alternatively, validating the nucleotide sequence of the assembled genome on the basis of the visual comparison of step b).20) Method according to claim 19, wherein the motif *TTCG****A**CGGG* is the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4), and / or the motif *CCCG**T****CGAA* is the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO:6).21 ) Method according to any one of claims 19-20, wherein said method is a computer-implemented method and said steps a), b), c) and d) are carried out by a processor.22) Method for identifying an alteration of human chromosome nucleotide sequence of a subject, said method comprising the following steps: a) identifying in the chromosome nucleotide sequence of said subject a motif, wherein said motif is: a CENP-B box protein binding motif chosen from: *TTCG****A**CGGG* orCCCG**T****CGAA*; or pJa protein binding motif chosen from: TTCCTTTT[C or T]CACC[A or G]TAG (SEQ ID NO:1 ) or CTA[C or T]GGTG[A or G]AAAAGGAA(SEQ ID NO:2); b) identifying the positions of the motifs identified in step a) in the chromosome nucleotide sequence; and c) comparing the positions of motifs identified in step b) with the positions of the same motifs in a reference chromosome nucleotide sequence, wherein when the positions of the motifs in the chromosome nucleotide sequence of the subject are different from the positions of the same motifs in the reference chromosome nucleotide sequence an alteration of chromosome nucleotide sequence is present.23) Method according to claim 22, wherein the motif *TTCG****A**CGGG* is the sequence [T or C]TTCGTTGGAA[A or G]CGGG[A or T] (SEQ ID NO:3), preferably the sequence TTTCGTTGGAAACGGGA (SEQ ID NO:4), and / or the motif *CCCG**T****CGAA* is the sequence [T or A]CCCG[T or C]TTCCAACGAA[A or G] (SEQ ID NO:5), preferably the sequence ACCCGTTTCCAACGAAA (SEQ ID NO:6).24) Method according to any one of claims 22-23, wherein step c) is carried out by visually comparing the positions of the motifs in a chromosome nucleotide sequence of a subject with the positions of the motifs in a reference chromosome nucleotide sequence.25) Method according to any one of claims 22-24, wherein said method is a computer-implemented method and said steps a), b) and c) are carried out by a processor.26) System for characterizing, identifying, isolating, assembling or validating the nucleotide sequence of a centromere of a chromosome, said system comprising:(A) storage media in which the following data are stored: an assembled genome or raw sequencing long-reads or contigs; and(B) control logic unit connected to said storage media and configured to perform steps a) to c) of the method as defined in any of claims 5-9.27) System for identifying possible pathologic changes in the centromere nucleotide sequence of a chromosome of a subject, said system comprising:(A) storage media in which the following data are stored: a chromosome nucleotide sequence of said subject; and(B) control logic unit connected to said storage media and configured to perform steps a) to c) of the method as defined in any of claims 10-13.28) System for calculating the size of the centromere nucleotide sequence of a chromosome, said system comprising:(A) storage media in which the following data are stored: a centromere nucleotide sequence of said subject; and(B) control logic unit connected to said storage media and configured to perform steps a) to d) of the method as defined in any of claims 14-16.29) System for characterizing, identifying, isolating, assembling or validating the nucleotide sequence of a chromosome, said system comprising: (A) storage media in which the following data are stored: assembled genome or raw sequencing long-reads or contigs; and(B) control logic unit connected to said storage media and configured to perform steps a) to c) of the method as defined in any of claims 19-21 .30) System for identifying an alteration of chromosome nucleotide sequence of a subject, said system comprising:(A) storage media in which the following data are stored: a chromosome nucleotide sequence of said subject; and(B) control logic unit connected to said storage media and configured to perform steps a) to c) of the method as defined in any of claims 22-25.