The present disclosure, among other things, provides
machine-learning technologies for identifying and localizing particular genomic elements (e.g.,
gene elements and / or regulatory elements) within
nucleotide sequences, such as
DNA and / or
RNA sequences. In certain embodiments, similar to the manner in which
image processing methods can be used to localize particular objects in images at pixel level resolution, referred to as "segmentation," systems and methods of the present disclosure predict presence and locations of certain genomic elements within
nucleotide sequences, thereby "segmenting"
nucleotide sequences. Accordingly, genomic element segmentation technologies described herein may be used to generate annotations that identify and
label portions of nucleotide sequences according to their predicted (e.g., via
machine learning models described herein) function – e.g., as
protein-coding genes, untranslated regions, splice sites, promotors, enhancers, etc. Among other things, these genomic annotations may be used to inform underlying biological processes driving diseases and facilitate development of new therapies.