A method comprises obtaining video data from one or more data sources, and
processing the obtained video data in a
machine learning
system comprising an
inference stage and an anticipation stage. The
inference stage is configured to assign one or more labels to at least one of a group activity and an individual activity detected in the obtained video data. The anticipation stage is configured to predict one or more future actions relating to at least one of the group activity and the individual activity based at least in part on the one or more labels assigned in the
inference stage. The method further comprises generating at least one
control signal based at least in part on the predicted one or more future actions. The method is illustratively configured to implement role inference and action anticipation in team sports, although it is applicable to a wide variety of other contexts.