Among a great deal of other disclosure and scope, systems and methods are enclosed that enable automated
labelling of a subset of vectors in a given
problem space. For example, in some of many cases, a first
machine learning model pre-trained on a given
problem space makes predictions regarding fresh, unseen data. In addition to this prediction, the model can output a
confidence metric indicating its confidence regarding the prediction made. A subset of these vectors with the highest confidence may be selected. Relevant
heuristics assessing each vector in the subset may be computed. These
heuristics can be fed through a second
machine learning model, which identifies if the given prediction made by the first model is correct. If so, the vector is automatically annotated with the correct predicted
label, the vector is appended to the labeled set of data, and the first model is retrained with the new labeled set of data.