A computer-implemented method includes obtaining, from
text corpus including article-summary pairs in a plurality of languages, a plurality of article-summary pairs in a target language among the plurality of languages, to form an article-summary pairs dataset in which each article corresponds to a summary; inputting articles from the article-summary pairs to a
machine learning model; generating, by the
machine learning model, embeddings for sentences of the articles; extracting, by the
machine learning model, keywords from the articles with a probability that varies based on lengths of the sentences, respectively; outputting, by the
machine learning model, the keywords; applying a maximal marginal relevance
algorithm to the extracted keywords, to select relevant keywords; and generating a keyword-text pairs dataset that includes the relevant keywords and text from the articles, the text corresponding to the relevant keywords in each of keyword-text pairs of the keyword-text pairs dataset.