Disclosed is a computer-implemented method comprising obtaining captured video content that is generated by capturing an activity, wherein the activity is defined by a formation adjacent to the structure, or comprised in the 5 structure, the structure being located in a geographical location, identifying elements comprised in the captured video content, determining for each identified element a classification, wherein a first classification is a classification for essential elements and a second classification is for 10 replaceable elements, determining at least one participant of the activity as an element in the first classification, determining, based on a digital duplicate of the structure, at least one part of the structure as an element in the second classification, based on at least one parameter of a receiving 15 device, obtaining at least one additional element to replace the element in the second classification, generating modified video content comprising the at least one additional element, that replaces the element in the second classification, and the elements in the first classification, and providing the modified 20 video content for rendering by the receiving device