Systems, devices, and methods are provided for determining contextually relevant user-based product recommendations based on scene information. In at least one embodiment, techniques described herein may be used to determine, using a first
machine-learning model, first information associated with a first object within a first image of
digital content, determine, using a second
machine-learning model, similarity scores between the first object and a first plurality of products of an online
purchasing system, detect, in association with the first image of the
digital content, performance of a first computer-based action by a user, determine, using a third
machine-learning model and based on contextual data of the user, one or more affinity scores for the user, select a first product based on the one or more affinity scores, and present a recommendation to the user to perform a second computer-based action in association with the first product.