Provided herein are technologies for framing and evaluating biological sequence-based analysis tasks in a unified, natural-language-based, text in and text out format. Among other things, methods and systems of the present disclosure provide
machine-learning technologies for combining biological sequence data, representing, for example,
DNA,
RNA, and
protein sequences, with
natural language, conversational style prompts that set out particular analysis tasks to be performed on the biological sequence data. This approach, for example, allows complex analysis tasks, including, but not limited to, identification of various sequence modifications, genes, and regulatory elements in
DNA sequences, and quantification of properties such as degradation propensity of
RNA and
protein stability, to be input to a
machine learning model in a uniform text-based format and for output to be generated in a same, unified, text-based format.