System and method for analyzing viewer engagement and providing targeted content recommendations

WO2026178669A1PCT designated stage Publication Date: 2026-09-03PROENVIROENERGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2026/050325
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2026-02-27
Publication Date
2026-09-03

Smart Images

  • Figure CA2026050325_03092026_PF_FP_ABST
    Figure CA2026050325_03092026_PF_FP_ABST
Patent Text Reader

Abstract

Systems, methods, and non-transitory computer-readable media are provided for analyzing viewer engagement within an environment and generating targeted content recommendations. In some embodiments, one or more processors receive, from one or more visual data capturing systems, viewer engagement parameters associated with a viewer, and generate performance metrics for one or more targets or points of interest based on the viewer engagement parameters. The one or more processors can generate recommendations for content based on the performance metrics and transmit the recommendations to a display device associated with one or more content providers.
Need to check novelty before this filing date? Find Prior Art

Description

CPST Ref: 40736 / 00041SYSTEM AND METHOD FOR ANALYZING VIEWER ENGAGEMENT AND PROVIDING TARGETED CONTENT RECOMMENDATIONS CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims priority to U.S. Provisional Patent Application No.63 / 765,472, filed February 28, 2025, the content of which is incorporated herein by reference in its entirety for all purposes.BACKGROUND

[0002] The present disclosure generally relates to systems, methods, and computer-readable media for monitoring viewers within an environment and, more particularly, to techniques for analyzing viewer engagement with one or more targets or points of interest in the environment and providing content recommendations based on the analyzed engagement.

[0003] In a variety of environments, including high-traffic locations such as stadiums, metro stations, galleries, and commercial spaces, it can be difficult to accurately assess whether viewers are attending to particular targets, as well as how intensely and for how long such attention is maintained. Environmental conditions, occlusions, viewing distance, and crowd dynamics can further complicate capture and analysis of engagement, and can reduce the usefulness of conventional approaches for evaluating content effectiveness.

[0004] In some scenarios, it may be desirable to generate and deliver content that is tailored to a viewer, a set of viewers, and / or a target based on observed engagement. However, effective tailoring in physical environments can be challenging when engagement data is not captured with sufficient fidelity, when attention is not evaluated in a timely manner, or when outputs useful to content providers are not produced in a manner that supports targeted delivery. Accordingly, there remains a need for improved techniques capable of detecting, tracking, and evaluating viewer engagement and generating recommendations for content based on performance metrics derived from viewer engagement parameters.SUMMARY

[0005] The following presents a simplified summary in order to provide a basic understanding of some aspects of the disclosed subject matter. This summary is not an extensive overview, and it is not intended to identify key / critical elements or to delineate the 1CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041scope thereof. Its sole purpose is to present some concepts in a simplified form as a prelude to the more detailed description that is presented later.

[0006] In some embodiments, a system for analyzing viewer engagement with one or more targets and points of interest within an environment includes a processor configured to receive, from one or more visual data capturing systems communicatively coupled to the processor, one or more viewer engagement parameters associated with a viewer within the environment. The processor can generate performance metrics for the viewer for the one or more targets and points of interest based on the viewer engagement parameters, generate recommendations for one or more content to be provided to the viewer based on the performance metrics, and transmit the recommendations to a display device associated with one or more content providers for providing targeted content to viewers in accordance with the recommendations.

[0007] In some implementations, the viewer engagement parameters can include heat signatures, posture, gaze direction, head orientation, and target engagement duration, associated with the viewer within the environment. In some implementations, the system further includes a video data capturing system configured to activate a first camera function to continuously scan the environment to identify one or more viewers showing interest in one or more targets within the environment by monitoring a first set of viewer engagement parameters. Upon identifying the one or more viewers, the processor can activate a second camera function of the video data capturing system to focus on the identified one or more viewers to monitor one or more second set of viewer engagement parameters associated with the identified one or more viewers. In some implementations, the processor can activate the second camera function to identify one or more viewers obscured from the first camera function.

[0008] In some implementations, the generated recommendations for the viewer can be associated with a unique identification associated with the viewer and stored in a database for future usage. In some implementations, the processor is configured to use an artificial intelligence (Al) model configured to process the monitored one or more viewer engagement parameters and generate the recommendations for one or more content to be provided to the viewer. In some implementations, the processor is further configured to use an iris detection estimation method for monitoring a gaze direction parameter associated with the viewer.2CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041

[0009] In some implementations, the viewer is moving inside a vehicle, and the processor is configured to detect a gaze direction as the viewer engagement parameter by capturing, using a video data capturing system, one or more video frames of the vehicle as the vehicle passes by the one or more targets or points of interest; analyzing, using one or more computer vision modules, each frame to detect a head pose of the viewer within the vehicle; and determining a gaze direction vector based on the detected head pose and one or more known camera parameters. In some implementations, the known camera parameters include one or more of focal length and lens distortion. In some implementations, the processor is configured to align one or more coordinate systems of a video data capturing system with one or more coordinates of the targets and the points of interest within the environment to map a location of the targets and the points of interest relative to a position of the video data capturing system. In some implementations, the processor is configured to adjust for gaze offset based on one or more behavioral parameters associated with the viewer to further refine the determined gaze direction vector.

[0010] Features from any of the above-mentioned embodiments may be used in combination with one another in accordance with the general principles described herein. These and other embodiments, features, and advantages will be more fully understood upon reading the following detailed description in conjunction with the accompanying drawings and claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The features of certain embodiments will become more apparent in the following detailed description in which reference is made to the appended figures wherein:

[0012] FIG. 1 is a simplified block diagram of an example system including a video data capturing system and an associated control system configured to process viewer engagement data using computer vision and AI / ML modules, store data in a database, and provide outputs via an output device.

[0013] FIG. 2 is a perspective view of an example environment depicting a deployment arrangement in which a video data capturing system is communicatively coupled to a control system to support viewer engagement analysis.

[0014] FIG. 3 is a perspective view of an example environment depicting a plurality of targets positioned as points of interest for monitoring viewer engagement.3CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041

[0015] FIG. 4 is a perspective view of a gallery environment depicting viewers and corresponding gaze directions relative to one or more targets to support engagement monitoring.

[0016] FIG. 5 is a perspective view of a metro station environment illustrating viewers and corresponding gaze directions toward one or more targeted points within the environment.

[0017] FIG. 6 is a plan view of an example stadium environment showing placement of video data capturing systems and corresponding fields of view for monitoring viewer engagement.

[0018] FIG. 7 is a perspective view of a stadium environment illustrating gaze directions of viewers for engagement monitoring.

[0019] FIG. 8 is a perspective view of an example highway billboard environment in which passing vehicles are monitored to determine viewer engagement with billboard content.

[0020] FIG. 9 is a perspective view of vehicles traveling along a highway in which one or more occupants are monitored for engagement with an external roadside target.

[0021] FIG. 10 is a perspective view of a roadway environment in which a vehicle travels along a highway past a roadside target for assessing viewer engagement associated with one or more occupants.

[0022] FIG. 11 is a perspective view of a highway billboard installation in which a video data capturing system is positioned relative to a highway and a billboard to capture visual data for determining viewer engagement of occupants of passing vehicles with content presented on the billboard.DETAILED DESCRIPTION

[0023] The following description is provided to exemplify one or more embodiments of the present subject matter. Although the embodiments illustrated herein may be described with reference to one or more specific features, it will be understood that, unless otherwise specified, all features described herein may be used interchangeably or in any combination with one or more other embodiments. Thus, any description of features relating one embodiment is not intended to limit the incorporation of such features to only that embodiment.4CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041

[0024] Various terms are used in the present description. Where appropriate, such terms are assigned meanings as indicated below or as they are introduced for the purpose describing one or more embodiments.

[0025] The term “embodiment” is used herein to describe one or more examples of representations or implementations of one or more features, elements, structures, or characteristics etc. (collectively “features”) of the present description. It will be understood that the features of a given embodiment of the description are not necessarily limited to such embodiment. In other words, any of the features described herein with respect to one embodiment may be used with, incorporated into, or combined with any of the described embodiments as would be understood by persons skilled in the art.

[0026] The terms “comprise”, “comprises”, “comprised” or “comprising” may be used in the present description. As used herein (including the specification and / or the claims), and unless stated otherwise, these terms are to be interpreted as open-ended terms and as specifying the presence of the stated features, integers, steps or components, but not as precluding the presence of one or more other feature, integer, step, component or a group thereof as would be apparent to persons having ordinary skill in the relevant art. Thus, the term "comprising" as used in this specification means "consisting at least in part of’. When interpreting statements in this specification that include that term, the features, prefaced by that term in each statement, all need to be present but other features can also be present. Related terms such as "comprise" and "comprised" are to be interpreted in the same manner.

[0027] The phrase “consisting essentially of’ or “consists essentially of’ will be understood as generally closed terms, with the exception of allowing inclusion of additional items, materials, components, steps, or elements, that do not materially affect the basic and novel characteristics or function of the item(s) used in connection therewith. For example, trace elements present in a composition, but not affecting the composition's nature or characteristics would be permissible if present under the “consisting essentially of’ language, even though not expressly recited in a list of items following such terminology. When using an open-ended term, such as “comprising” or “including”, it will be understood that direct support should be afforded also to “consisting essentially of’ language as well as “consisting of’ language as if stated explicitly and vice versa. In essence, use of one of these terms in the specification provides support for all of the others.5CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041

[0028] For the purposes of the present description and / or claims, and unless otherwise indicated, all numbers expressing quantities, percentages or proportions, and other numerical values used in the specification and claims, are to be understood as being modified in all instances by the term “about.” Accordingly, unless indicated to the contrary, the numerical parameters set forth herein are approximations that may vary depending upon the desired properties sought to be obtained by the presently described invention, inclusive of the stated value and has the meaning including the degree of error associated with measurement of the particular quantity. The term “about” generally refers to a range of numbers that one of ordinary skill in the art would consider as a reasonable amount of deviation to the recited numeric values (i.e., having the equivalent function or result). For example, as used herein, the term “about” can be construed as including a deviation of ±10 percent of the given numeric value provided such a deviation does not alter the end function or result of the value. Therefore, a value of about 1% can be construed to be a range from 0.9% to 1.1%.

[0029] The term "and / or" can mean "and" or "or".

[0030] Unless stated otherwise herein, the articles “a” and “the”, when used to identify an element, are not intended to constitute a limitation of just one and will, instead, be understood to mean “at least one” or “one or more”.

[0031] As used herein, the term “viewer” refers to a person located within an environment and potentially within a field of view of a visual data capturing system, including a person located within a vehicle.

[0032] As used herein, the term “environment” refers to a physical space in which viewer engagement is analyzed, including, by way of example, a stadium, a gallery, a room, a highway or roadway, a metro station, a commercial space, or a public space.

[0033] As used herein, the term “target” or “point of interest” refers to any object, person, equipment, product, advertisement, display, message, or other item and / or location within an environment for which viewer engagement is monitored.

[0034] As used herein, the term “viewer engagement parameter” refers to any measurable or inferable characteristic associated with a viewer that is usable to determine attention toward a target or point of interest, including, by way of example, gaze direction, head orientation, posture, facial expression, body language, heat signature, and target engagement duration.6CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041

[0035] As used herein, the term “performance metric” refers to any metric generated based on one or more viewer engagement parameters and associated with a viewer and / or a target or point of interest, wherein the performance metric is usable to evaluate engagement with the target or point of interest.

[0036] As used herein, the term “recommendation” refers to output data generated based on one or more performance metrics and indicative of one or more content items to be provided to a viewer and / or a set of viewers and / or via a content provider for targeted content delivery.

[0037] As used herein, the term “visual data capturing system” refers to one or more image capture devices configured to capture visual data representative of at least a portion of an environment, including one or more cameras that may include a wide-angle camera function and / or a zoom-enabled camera function.

[0038] As used herein, the term “computer vision module” refers to a software and / or hardware module configured to perform image and / or video analysis to determine, from captured visual data, one or more viewer engagement parameters, including detecting a face and estimating head pose.

[0039] As used herein, the term “AI / ML module” refers to a software and / or hardware module configured to apply artificial intelligence and / or machine learning to analyze one or more viewer engagement parameters and / or one or more performance metrics and to generate one or more recommendations.

[0040] As used herein, the term “head pose” refers to an estimated orientation of a viewer’s head in three-dimensional space, including one or more of yaw, pitch, and roll.

[0041] As used herein, the term “gaze direction vector” refers to an estimated direction of a viewer’s gaze determined based on at least head pose and one or more camera parameters.

[0042] As used herein, the term “camera parameters” refers to one or more parameters associated with a visual data capturing system that are usable for determining gaze direction or for geometric mapping, including, by way of example, focal length and lens distortion.

[0043] As used herein, the term “calibration” refers to determination and / or application of a mapping between a coordinate system associated with a visual data capturing system and one or more coordinates associated with a target or point of interest within an environment to enable estimation of the target location relative to a camera position.7CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041

[0044] As used herein, the term “gaze offset” refers to a difference between an estimated gaze direction based on head orientation and an actual gaze direction, wherein the gaze offset is adjustable based on one or more behavioral parameters associated with a viewer.

[0045] As used herein, the term “field of view” refers to a portion of an environment that is capturable by a visual data capturing system at a given time.

[0046] As used herein, the term “real time” or “near real time” refers to processing and / or output generation occurring during data capture or with a latency sufficiently low to support timely analysis and / or targeted content delivery based on the captured data.

[0047] As used herein, the term “unique identification” refers to a data value usable to distinguish a viewer from other viewers and usable to associate viewer engagement data and / or recommendations with a stored database record for future usage.

[0048] In one aspect, FIG. 1 is a simplified block diagram of an example system 100 for analyzing viewer engagement within an environment. The system 100 includes a video data capturing system 102 configured to capture visual data representative of at least a portion of the environment and a control system 200 coupled to the video data capturing system 102 to receive captured visual data and derive viewer engagement parameters from the captured visual data. The control system 200 is further coupled to a database 112 for persistence of engagement-related data and is coupled to an output device 110 to provide outputs indicative of performance metrics and recommendations associated with one or more targets or points of interest.

[0049] According to an embodiment, the video data capturing system 102 includes a first camera 104 and a second camera 106 arranged as respective camera functions that cooperate to support viewer detection and engagement analysis across different viewing conditions. In some embodiments, the first camera 104 provides a wide field-of-view for scanning the environment to identify one or more viewers and one or more candidate targets or points of interest associated with viewer attention. Additionally, the second camera 106 provides a variable field-of-view, including via optical zoom control, to obtain higher-resolution imagery for a selected region associated with an identified viewer, such as to support estimation of gaze direction, engagement duration, facial expression, and other viewer engagement parameters.

[0050] In some embodiments, the control system 200 includes an input unit 202, an output unit 204, a memory unit 206, and a processor 208 coupled via a local interface 212. The 8CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041local interface 212 is representative of one or more interconnects that carry data, address, and control signaling among the input unit 202, the output unit 204, the memory unit 206, and the processor 208 during operation of the control system 200. Furthermore, the input unit 202 is operable to receive operator input and / or configuration data associated with the system 100, the output unit 204 is operable to render one or more user interface elements for monitoring engagement results, the memory unit 206 stores executable instructions and data structures used by the processor 208, and the processor 208 executes the instructions to receive viewer engagement parameters, generate performance metrics, generate recommendations, and transmit the recommendations for targeted content delivery.

[0051] According to an embodiment, the processor 208 executes a computer vision module 214 to process image frames received from the video data capturing system 102 and to extract one or more viewer engagement parameters from the image frames. In some embodiments, the computer vision module 214 performs one or more of face detection, head pose estimation, gaze direction estimation, posture determination, and engagement duration tracking for one or more viewers appearing in the captured visual data. Additionally, outputs of the computer vision module 214 are provided to an AI / ML module 216 that is executable by the processor 208 to generate one or more performance metrics and to map the performance metrics to one or more recommendations for content to be provided to a viewer, a cohort of viewers, and / or a content provider workflow.

[0052] In some embodiments, the database 112 stores one or more of raw visual data, extracted viewer engagement parameters, derived performance metrics, generated recommendations, and associations between engagement-related data and a unique identification for a viewer. The database 112 is shown coupled to the control system 200 to support read and write operations performed by the processor 208, including retention of historical engagement data for subsequent reporting, model training, and recommendation refinement. Furthermore, the output device 110 is shown coupled to the control system 200 to receive, from the processor 208, output data indicative of the generated recommendations and to present such output data to a user interface associated with one or more content providers, thereby enabling targeted content to be selected, scheduled, and / or delivered based on the engagement analysis performed by the system 100.

[0053] In one aspect, FIG. 2 is a perspective view of an example deployment of system 100 within an environment, in which video data capturing system 102 is positioned to capture visual data associated with viewers and to provide the visual data, or data derived therefrom,9CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041to control system 200 for viewer engagement analysis. The illustrated arrangement represents an installation in which video data capturing system 102 is disposed at an elevated location relative to a viewing area of the environment to support monitoring of viewer behavior with respect to one or more targets or points of interest that may be present within the environment.

[0054] According to an embodiment, video data capturing system 102 is mounted to an environmental structure and oriented to obtain image frames of a portion of the environment that includes one or more viewers. As shown, video data capturing system 102 can include one or more camera devices arranged to capture visual data across a viewing zone, where the viewing zone is selected based on a target location, a viewer travel path, a seating arrangement, or another region in which viewer engagement with a target is assessed. Additionally, video data capturing system 102 can be positioned to support multi-range capture in which a first camera function provides a first field of view for detecting candidate viewers and a second camera function acquires higher-resolution imagery for one or more selected viewers, consistent with the dual-camera implementations described with respect to FIG. 1.

[0055] In some embodiments, control system 200 is communicatively coupled to video data capturing system 102 to receive captured visual data and to execute one or more processing workflows for determining viewer engagement parameters and generating outputs usable for content targeting. The coupling between video data capturing system 102 and control system 200 is represented schematically in FIG. 2 and can be implemented using a wired link, a wireless link, or a network connection. Furthermore, control system 200 can be located proximate to video data capturing system 102 for local processing, or can be located remotely to support centralized processing for multiple installations of system 100 across one or more environments.

[0056] According to an embodiment, system 100 operates in accordance with a capture-and-analyze sequence in which video data capturing system 102 acquires image frames for the environment and control system 200 processes the image frames to determine engagement associated with one or more targets or points of interest. In this operational context, control system 200 can determine, from the captured visual data, one or more viewer engagement parameters including gaze direction, head orientation, posture, heat signature, and engagement duration, and can generate corresponding performance metrics associated with a viewer and a target. Additionally, control system 200 can generate10CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041recommendations for content based on the performance metrics and can make the recommendations available for transmission to an output device associated with a content provider, as described with respect to FIG. 1, thereby enabling targeted content delivery based on engagement observed in the environment shown in FIG. 2.

[0057] In one aspect, FIG. 3 is a perspective view of an example environment 300 showing a set of targets 302 positioned as respective points of interest for which viewer engagement is analyzed by system 100. The example environment 300 is illustrated as an enclosed space having boundary surfaces, including walls and a floor, on which multiple targets 302 are disposed at different locations and heights, thereby representing a scenario in which multiple distinct points of interest are concurrently present within a field of view of one or more visual data capturing systems 102 described with respect to FIGS. 1-2.

[0058] According to an embodiment, the targets 302 include physical objects and / or displayed content elements located within the example environment 300, such as wall-mounted items, free-standing items proximate to a wall surface, and a target 302 positioned adjacent to a floor surface. In this context, each target 302 defines a candidate region for which control system 200 determines whether one or more viewers direct attention toward the target 302 based on one or more viewer engagement parameters, including one or more of gaze direction, head orientation, posture, heat signature, and target engagement duration, derived from visual data captured within the example environment 300.

[0059] In some embodiments, the arrangement of targets 302 within the example environment 300 supports associating viewer engagement with a selected target 302 by mapping a viewer-related line of sight, gaze direction vector, and / or head pose to a corresponding location of the target 302 in the environment coordinate space. Additionally, the presence of multiple targets 302 in the example environment 300 supports generation of performance metrics that distinguish engagement with different targets 302, such as by computing target-specific dwell time, frequency of attention events, or temporal sequences in which a given viewer transitions attention among targets 302, for use in generating recommendations for content to be provided to the viewer and / or to one or more content providers.

[0060] Furthermore, FIG. 3 illustrates that the geometric placement of the targets 302 within the example environment 300 can be used for calibration and reference during operation of the system 100. In some embodiments, control system 200 stores, in database 112, target location data for the targets 302 and associates the target location data with engagement- 11CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041related data derived from captured image frames, thereby enabling reporting and recommendation generation that is indexed by target 302 and by environment 300, including in deployments in which a plurality of environments are monitored by a shared control system 200.

[0061] In one aspect, FIG. 4 is a perspective view of the example environment 300 in which one or more viewers, shown as viewer 304, are positioned within a gallery space that includes multiple targets 302 disposed as points of interest. FIG. 4 illustrates viewer engagement as a spatial relationship between a viewer 304 and a selected target 302, where system 100 uses captured visual data to determine whether the viewer 304 is attending to the target 302 and to derive engagement-related data that is later used by control system 200 to generate performance metrics and content recommendations as described with respect to FIGS. 1-3.

[0062] According to an embodiment, FIG. 4 depicts that viewer engagement can be represented by a viewer gaze direction 306 extending from a head region of the viewer 304 toward a corresponding target 302. In operation, the visual data capturing system 102 captures image frames that include the viewer 304 and the targets 302, and the control system 200 processes the image frames to estimate the viewer gaze direction 306 as a viewer engagement parameter, including by determining head pose, facial landmarks, and / or eye-region features based on a viewing distance and an available image resolution.Additionally, the estimated viewer gaze direction 306 can be expressed as a gaze direction vector in a camera coordinate system and mapped, via calibration data for the example environment 300, to determine whether the gaze direction intersects, corresponds to, or is otherwise associated with a particular target 302.

[0063] In some embodiments, FIG. 4 further illustrates that a plurality of targets 302 can be concurrently present such that different viewers 304, or a same viewer 304 across time, can be associated with different targets 302 within the example environment 300. In this context, the control system 200 associates instances of the viewer gaze direction 306 with target identifiers for the targets 302 to support target-specific engagement determination, including engagement duration, attention event counts, and temporal sequences of attention among the targets 302. Furthermore, the control system 200 stores, in database 112, records linking such target-specific engagement data to a viewer identifier when available, thereby enabling subsequent reporting and recommendation generation that is indexed by viewer 304 and by target 302 within the example environment 300.12CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041

[0064] Additionally, FIG. 4 supports an operational interpretation in which viewer gaze direction 306 is evaluated in combination with other viewer engagement parameters to refine an engagement determination for a given target 302. In some embodiments, the control system 200 correlates the viewer gaze direction 306 with head orientation and posture of the viewer 304 to distinguish attentive viewing from incidental head turns, and tracks a persistence of a gaze-to-target association over successive frames to determine an engagement duration for the target 302. Furthermore, the engagement data derived for the viewers 304 and the targets 302 in the example environment 300 is provided to the AI / ML module 216 to generate performance metrics and to generate recommendations for content that may be transmitted to an output device associated with a content provider, thereby enabling selection or scheduling of content based on observed engagement within the gallery setting depicted in FIG. 4.

[0065] In one aspect, FIG. 5 is a perspective view of a metro station environment 500 in which a set of viewers, shown as viewer 502, are present within a transit space that includes one or more targeted points within the environment. FIG. 5 illustrates that system 100 derives viewer engagement parameters for the viewer 502 within the metro station environment 500 and uses the derived viewer engagement parameters to determine engagement with targets or points of interest, including for generation of performance metrics and recommendations as described with respect to FIGS. 1-4.

[0066] According to an embodiment, FIG. 5 depicts that viewer engagement within the metro station environment 500 can be represented by a gaze direction 504 associated with a given viewer 502 and oriented toward a target region located within the environment. In operation, one or more visual data capturing systems 102 capture image frames that include the viewer 502, and control system 200 processes the image frames to estimate the gaze direction 504 as a viewer engagement parameter, including by detecting a face region, estimating head pose, and mapping head orientation and / or eye-region features into a gaze direction vector in a camera coordinate system.

[0067] In some embodiments, the metro station environment 500 supports concurrent monitoring of multiple viewers 502 distributed along a travel path, platform region, or concourse region, such that gaze directions 504 associated with different viewers 502 can be evaluated with respect to different targeted points within the environment. Additionally, control system 200 associates a sequence of gaze direction 504 estimates for a viewer 502 over successive frames with a target identifier to determine target engagement duration and 13CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041other target-specific performance metrics, including attention event counts and temporal patterns that indicate transitions of attention among multiple targeted points.

[0068] Furthermore, FIG. 5 supports an operational context in which gaze direction 504 is evaluated in combination with other viewer engagement parameters to refine an engagement determination for a target within the metro station environment 500. In some embodiments, control system 200 correlates gaze direction 504 with head orientation and posture for the viewer 502 to distinguish target-focused attention from incidental scanning behavior associated with movement through the environment, and stores engagement-related data in database 112 for subsequent reporting and recommendation generation. Additionally, the stored engagement-related data is provided to AI / ML module 216 to generate performance metrics and to generate recommendations for content that is transmittable to output device 110 for use by a content provider in selecting, scheduling, or delivering targeted content associated with viewer engagement observed in the metro station environment 500.

[0069] In one aspect, FIG. 6 is a plan view of a stadium environment 600 in which a set of video data capturing systems 102 are positioned around a perimeter of the stadium environment 600 to capture visual data for monitoring viewer engagement. FIG. 6 depicts that each video data capturing system 102 is oriented toward an interior region of the stadium environment 600 such that the system 100, via control system 200 described with respect to FIG. 1, receives image frames from one or more vantage points and derives viewer engagement parameters associated with one or more viewers and one or more targets or points of interest located within the stadium environment 600.

[0070] According to an embodiment, FIG. 6 further depicts respective fields of view 601 corresponding to the video data capturing systems 102, where each field of view 601 represents a portion of the stadium environment 600 that is capturable by a given video data capturing system 102 during a given capture interval. In this arrangement, the fields of view 601 are defined by camera orientation, lens characteristics, and placement location of the video data capturing systems 102, and the fields of view 601 collectively cover seating regions and an event region to support measurement of viewer engagement parameters including gaze direction, head orientation, posture, heat signature, and engagement duration within the stadium environment 600.

[0071] In some embodiments, the spatial distribution of the video data capturing systems 102 around the stadium environment 600 supports multi-perspective capture in which a 14CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041viewer that is partially occluded in a first field of view 601 is observable in a second field of view 601 , thereby supporting continuity of engagement estimation across crowd dynamics and structural occlusions. Additionally, the control system 200 correlates viewer-related observations obtained from different video data capturing systems 102 based on time alignment and geometric relationships associated with the stadium environment 600, such as to associate a given viewer with a given target or point of interest over successive frames and to generate target-specific performance metrics that are usable for recommendation generation.

[0072] Furthermore, FIG. 6 supports implementations in which the fields of view 601 are used as inputs to calibration and monitoring workflows executed by the control system 200. In some embodiments, the control system 200 stores configuration data for respective video data capturing systems 102, including placement data and mapping data that relates a camera coordinate system for a given video data capturing system 102 to one or more locations within the stadium environment 600, thereby enabling mapping of a gaze direction vector derived from captured visual data to a selected region within the stadium environment 600. Additionally, the control system 200 uses such mappings to associate engagement events with corresponding targets or points of interest and to transmit recommendations, derived from performance metrics generated for the stadium environment 600, to an output device associated with one or more content providers as described with respect to FIG. 1.

[0073] In one aspect, FIG. 7 is a perspective view of a stadium environment 600 illustrating viewers 604 and respective gaze directions 602 within the stadium environment 600. FIG. 7 depicts an engagement-monitoring scenario in which system 100, via one or more visual data capturing systems 102 and control system 200 described with respect to FIGS. 1-6, derives gaze direction as a viewer engagement parameter for viewers 604 and associates such gaze direction information with one or more targets or points of interest in the stadium environment 600 for generation of performance metrics and content recommendations.

[0074] According to an embodiment, the viewers 604 are positioned within seating areas of the stadium environment 600, and the gaze directions 602 represent respective lines of sight extending from head regions of the viewers 604 toward an interior region of the stadium environment 600. In operation, captured image frames containing the viewers 604 are processed to estimate the gaze directions 602 based on one or more viewer engagement parameters, including head pose and facial landmark features, and the resulting gaze15CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041directions 602 are expressible as gaze direction vectors in a camera coordinate system for subsequent mapping to the stadium environment 600.

[0075] In some embodiments, the gaze directions 602 are evaluated across a plurality of viewers 604 to quantify audience attention toward one or more stadium-disposed targets or points of interest, such as a display surface, an event region, a scoreboard region, or another content presentation location. Additionally, the control system 200 correlates gaze directions 602 with other engagement-related observations for the viewers 604, including posture and engagement duration, to distinguish sustained attention events from transient head movements, and to generate target-associated performance metrics that quantify attention distribution within the stadium environment 600.

[0076] Furthermore, FIG. 7 supports implementations in which gaze directions 602 derived for the viewers 604 are aggregated overtime and across seating sections of the stadium environment 600 to produce outputs usable by content providers. In some embodiments, the control system 200 stores, in database 112, gaze-direction sequences and associated engagement duration values for the viewers 604 and uses such stored data to generate recommendations for content to be delivered via one or more display devices associated with content providers, including recommendations that are indexed by region within the stadium environment 600 and by temporal segments of an event occurring in the stadium environment 600.

[0077] In one aspect, FIG. 8 is a perspective view of a roadway monitoring scenario in which a highway 800 includes a billboard 802 positioned adjacent to a travel path for vehicles 804. FIG. 8 illustrates an example environment in which system 100, via one or more visual data capturing systems 102 and control system 200 described with respect to FIG. 1, monitors viewers located within passing vehicles 804 to determine engagement with content presented on the billboard 802 and to generate performance metrics and recommendations based on the determined engagement.

[0078] According to an embodiment, the billboard 802 is implemented as a target or point of interest disposed at a roadside location associated with the highway 800 such that an occupant of a vehicle 804 may have a line of sight to a display surface of the billboard 802 during vehicle travel. In operation, a visual data capturing system 102 is positioned at a location relative to the billboard 802 and oriented to capture image frames that include at least a portion of the vehicle 804 as the vehicle 804 traverses a monitored segment of the highway 800, and the control system 200 receives the image frames, or data derived16CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041therefrom, to determine one or more viewer engagement parameters for one or more occupants of the vehicle 804 with respect to the billboard 802.

[0079] In some embodiments, the control system 200 applies the computer vision module 214 to the captured image frames to detect one or more faces associated with vehicle occupants, estimate head pose for the detected faces, and determine a gaze direction vector based on the head pose and camera parameters for the visual data capturing system 102. Additionally, calibration data is applied to map a coordinate system associated with the visual data capturing system 102 to a location of the billboard 802, such that the gaze direction vector is evaluated relative to a plane or region associated with the billboard 802 to determine whether a corresponding occupant is attending to the billboard 802 during travel along the highway 800. Furthermore, the control system 200 tracks a persistence of a gaze-to-billboard association across successive frames to determine an engagement duration for one or more occupants of the vehicle 804 with respect to the billboard 802 and stores engagement-related data in database 112 for subsequent reporting and recommendation generation.

[0080] Additionally, FIG. 8 supports implementations in which engagement determinations for multiple vehicles 804 are aggregated to generate performance metrics associated with the billboard 802 and the highway 800 deployment location. In some embodiments, the AI / ML module 216 processes engagement-related data, including inferred attention events and engagement duration values, to generate performance metrics indicative of content effectiveness for content presented on the billboard 802, and to generate recommendations for content to be provided via the billboard 802 and / or via an associated content provider workflow. Furthermore, the generated recommendations are transmitted to an output device 110 associated with one or more content providers to support adjusting, selecting, scheduling, or otherwise providing targeted content based on engagement observed for vehicles 804 traveling along the highway 800.

[0081] In one aspect, FIG. 9 is a perspective view of a roadway monitoring scenario in which vehicles 804 travel along a highway 800 while one or more occupants within the vehicles 804 are monitored for engagement with an external roadside target. FIG. 9 illustrates a capture context in which system 100, via one or more visual data capturing systems 102 and control system 200 described with respect to FIGS. 1-8, acquires image frames that include a moving vehicle 804 and uses the captured visual data to derive viewer engagement17CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041parameters for determining attention toward a target located outside the passenger compartment.

[0082] According to an embodiment, the highway 800 defines a travel path that brings successive vehicles 804 through a monitored segment associated with a target region, such as a billboard or other point of interest positioned adjacent to the highway 800 as described with respect to FIG. 8. In operation, the visual data capturing system 102 is oriented to image vehicles 804 as they traverse the monitored segment, and the resulting image frames include views through vehicle glazing to expose facial regions of one or more occupants for subsequent analysis. Additionally, the perspective depiction in FIG. 9 represents that more than one vehicle 804 may be present within the monitored segment during a capture interval, enabling concurrent evaluation of engagement for occupants of different vehicles 804 traveling along the highway 800.

[0083] In some embodiments, control system 200 processes image frames captured for the vehicles 804 on the highway 800 to extract viewer engagement parameters associated with one or more occupants, including gaze direction, head orientation, posture, and engagement duration. Additionally, and consistent with the moving-vehicle processing described with respect to FIG. 8, control system 200 applies computer vision module 214 to perform face detection for an occupant visible within a cabin region of a vehicle 804, estimate head pose for the detected face, and determine a gaze direction vector based on the head pose and camera parameters. Furthermore, calibration data associating a coordinate system of the visual data capturing system 102 with a coordinate system for the roadway deployment enables evaluating whether the gaze direction vector corresponds to the external target region while the vehicle 804 travels along the highway 800, and enables tracking a persistence of a gaze-to-target association across successive frames to derive engagement duration for storage in database 112 and for subsequent generation of performance metrics and recommendations.

[0084] In one aspect, FIG. 10 is a perspective view of a roadway environment depicting a highway 800 along which a vehicle travels along a travel path associated with a roadside target used for viewer engagement analysis. FIG. 10 represents a capture context in which system 100, using a visual data capturing system 102 and control system 200 as described with respect to FIGS. 1-9, acquires visual data of a moving vehicle while the vehicle traverses a monitored segment of the highway 800 and uses the visual data to derive viewer engagement parameters associated with one or more occupants.18CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041

[0085] According to an embodiment, the highway 800 provides an environment in which relative motion between the vehicle and a roadside target produces a time-bounded opportunity to observe attention toward the roadside target. In this context, image frames captured while the vehicle traverses the monitored segment of the highway 800 are time-aligned by the control system 200 to support engagement duration tracking, including determining a sequence of attention events attributable to a given occupant as the vehicle approaches, passes, and departs from the roadside target region.

[0086] In some embodiments, the roadway environment of FIG. 10 further supports applying geometric calibration data that maps a camera coordinate system associated with the visual data capturing system 102 to an environment coordinate system associated with the highway 800 and the roadside target region. Additionally, the control system 200 determines one or more viewer engagement parameters for one or more occupants, including head pose and a gaze direction vector derived from head pose and camera parameters, and evaluates whether the gaze direction vector corresponds to the roadside target region during travel along the highway 800. Furthermore, engagement determinations and derived performance metrics associated with the highway 800 deployment are storable in database 112 and usable by the AI / ML module 216 to generate recommendations transmittable to an output device 110 for targeted content delivery associated with the roadside target.

[0087] In one aspect, FIG. 11 is a perspective view of a highway billboard installation in which a billboard 802 is disposed adjacent to a highway 800 and a video data capturing system 102 is positioned to capture visual data associated with occupants of vehicles traveling along the highway 800. FIG. 11 depicts a deployment arrangement in which the video data capturing system 102 is located at a roadside position that provides a viewing angle toward one or both of a display surface of the billboard 802 and window regions of passing vehicles, such that visual observations usable to derive viewer engagement parameters are obtainable while the vehicles traverse a monitored segment of the highway 800.

[0088] According to an embodiment, the video data capturing system 102 is mounted to, or positioned adjacent to, an environmental structure proximate to the billboard 802 to enable capture of image frames during a time interval in which vehicle occupants have visual access to the billboard 802. In this configuration, the video data capturing system 102 can include one or more camera devices configured to acquire image frames with sufficient spatial coverage to observe multiple traffic lanes of the highway 800 and with sufficient 19CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041temporal sampling to support engagement duration determination as vehicles approach, pass, and depart from the billboard 802. Additionally, the positional relationship between the video data capturing system 102 and the billboard 802 supports applying calibration data that maps a camera coordinate system associated with the video data capturing system 102 to a target coordinate representation associated with the billboard 802.

[0089] In some embodiments, FIG. 11 represents that the billboard 802 defines a target plane or target region for which attention is evaluated by mapping viewer-related gaze information into the target coordinate representation. In operation, image frames captured by the video data capturing system 102 are processed by the control system 200 described with respect to FIGS. 1-10 to detect and track one or more occupant faces visible through vehicle glazing and to estimate head pose for the detected faces. Additionally, the control system 200 derives a gaze direction vector using the estimated head pose and camera parameters associated with the video data capturing system 102, and evaluates whether the gaze direction vector intersects, or corresponds to, a plane associated with the billboard 802 to infer an attention event attributable to an occupant during travel along the highway 800.

[0090] Furthermore, FIG. 11 supports a workflow in which the inferred attention events are time-aligned across successive frames to compute target engagement duration values for the billboard 802 for respective occupants observed within the monitored segment of the highway 800. In some embodiments, such engagement duration values, along with other engagement-related data derived from the image frames, are stored in database 112 and aggregated across vehicles to generate performance metrics associated with content presented on the billboard 802 at the highway 800 location. Additionally, the performance metrics are usable by the AI / ML module 216 for generating recommendations for content to be presented via the billboard 802 and for transmitting recommendation outputs to an output device associated with a content provider to support targeted content selection and scheduling for the billboard 802 deployment.

[0091] In some embodiments, the system 100 is configured to collect, derive, and analyze viewer engagement parameters to determine viewer attention toward one or more targets 302 and / or points of interest within an environment, and to generate performance metrics and corresponding recommendations for content based on the determined viewer attention. In this regard, the system 100 can be implemented in a plurality of environments, including open spaces and closed spaces, and can process visual data captured by one or more visual data capturing systems 102 to track gaze direction and head orientation and to 20CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041estimate engagement levels associated with a target, such as an advertisement, an object, equipment, a product, a display, and / or a person.

[0092] In some implementations, the system 100 is configured to operate in real time or near real time to support monitoring of viewer engagement and generation of outputs usable by one or more content providers. For example, the system 100 can incorporate an artificial intelligence based model (e.g., AI / ML module 216) to process viewer engagement parameters during and / or shortly after capture of the visual data, thereby enabling dynamic selection and / or tailoring of content in physical environments. In some implementations, the system 100 is configured to provide reports including, for example, counts of viewers, demographic information inferred from facial features or vehicle type, and attention levels associated with one or more targets.

[0093] In some embodiments, the system 100 is configured to implement a dual-camera video data capturing system 102 in which a first camera 104 is configured as a wide-angle camera to monitor a broad field of view and a second camera 106 is configured with an automatically adjusting optical zoom to vary its field of view and focus for detailed tracking. In operation, the first camera 104 (or a first camera function) can continuously scan the environment to detect one or more viewers showing interest in a target based on a first set of viewer engagement parameters including, for example, heat signatures and postural cues such as head orientation. When the first set of viewer engagement parameters indicates viewer interest, the second camera 106 (or a second camera function) can be activated to monitor a second set of viewer engagement parameters associated with the viewer, such as gaze direction and engagement duration, and can further be used to identify one or more viewers obscured from the first camera function. In some implementations, for small and short-distance applications the system 100 can use a single stationary camera, and in some implementations the first and second camera functions can be combined into a single device.

[0094] In some embodiments, the system 100 is configured to link user identity and interest data to a central database 112 that stores user preferences. In some implementations, individuals can be identified through prior registration or subscription, and / or in ticketed or identity-recorded environments (e.g., stadiums, metro stations, museums). When identification is available, engagement analysis can be linked to a corresponding database record for future usage, for example for selecting relevant advertisements or product21CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041offerings for subsequent presentation based on previously recorded interests determined from one or more of gaze direction, head movement, and posture.

[0095] In some embodiments, the system 100 is configured for roadway or billboard environments in which a viewer is moving inside a vehicle 804. In such implementations, a camera can be positioned proximate to a target such as a billboard 802, and the system 100 can be configured to monitor occupants of passing vehicles 804 to determine whether, and for how long, the occupants focus on the billboard 802. In some implementations, captured data is transmitted to a control system 200 where one or more computer vision modules 214 and / or the AI / ML module 216 facilitate processing for deriving viewer engagement parameters and generating corresponding performance metrics and recommendations.

[0096] In some embodiments, the system 100 employs a calibration technique to enhance accuracy in real-world and billboard settings. For example, in environments where objects or individuals may be obscured from a primary camera view, a secondary camera can be used for semi-automated calibration. In some implementations, markers can be attached to one or more subjects, and the secondary camera can map subject positions to improve a main camera's performance and to support precise tracking of user behavior in the presence of visual occlusion.

[0097] In some embodiments, for gaze detection, the system 100 uses different methodologies based on distance from the camera 102. For example, when viewers are within a short distance (e.g., around 2 meters), iris direction estimation can be used to assess gaze more accurately, and when viewers are at longer distances (e.g., around 10 meters), head orientation can be used to estimate where the viewer is looking. These approaches can be applied in the billboard setting to monitor whether occupants of passing vehicles are engaging with the target and to support generation of performance metrics and reports.

[0098] In an example moving-vehicle implementation, the system 100 is configured to analyze a vehicle occupant's gaze direction by capturing video frames of a vehicle and vehicle occupants as the vehicle passes by a target (e.g., a billboard), and analyzing each frame to detect head pose using a computer vision module 214. In some implementations, the computer vision module 214 is configured to detect a face of the vehicle occupants using one or more facial recognition libraries, estimate head orientation using one or more 3D head pose estimation models, and calculate a gaze direction vector based on the head pose and one or more known camera parameters. In some implementations, the facial recognition 22CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041library includes OpenCV, Dlib, or combinations thereof, and the known camera parameters include focal length and lens distortion.

[0099] In some embodiments, the system 100 is configured to apply geometric calibration to align a camera coordinate system with a target position and to map a target location in 3D space relative to the camera. For example, the system 100 can align the camera coordinate system with the billboard position and implement one or more techniques to map the billboard location in 3D space relative to the camera, including using vanishing point detection and homography. The system 100 can then determine whether a gaze direction vector intersects a 3D plane associated with the billboard 802 to infer that an occupant is looking at the target, and can log instances of such gaze intersection over time to quantify viewer interest. In some implementations, the system 100 is configured to aggregate data over time to assess overall interest levels in the billboard advertisement, and to compare data from multiple billboard locations to evaluate effectiveness of each placement.

[0100] In some implementations, the system 100 is configured to adjust for gaze offset based on one or more behavioral parameters associated with a viewer to refine a gaze direction vector. For example, to mitigate inaccuracies that may arise when using head pose as a proxy for gaze at long distances, the system 100 can implement calibration to account for individual differences between head pointing direction and actual gaze, and / or can apply machine learning models to adjust gaze offsets based on behavioral patterns.

[0101] In some implementations, long-distance gaze estimation is performed by detecting facial landmarks and estimating head pose using a Perspective-n-Point (PnP) algorithm that maps 2D facial landmark image points to a pre-defined 3D head model. Facial landmark detection can be performed using Dlib's facial landmark predictor, MediaPipe Face Mesh, or combinations thereof. In some implementations, head pose is represented by yaw, pitch, and roll angles, and the head orientation angles are mapped into an estimated gaze direction, optionally with calibration to account for individual gaze offset.

[0102] In some implementations, short-distance gaze estimation is performed using iris detection. For example, the system 100 can detect an eye region using facial landmarks, localize an iris using deep learning-based models and / or traditional computer vision techniques, and estimate a gaze vector based on a position of the iris relative to eye contours. Example traditional computer vision techniques include edge detection and circular Hough transform. In some implementations, robust pre-processing can be applied to improve iris localization under challenging conditions, including adaptive thresholding.23CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041

[0103] In some embodiments, the system 100 is configured to recognize individuals in live video streams using a single reference photo per person. In some implementations, face detection is performed in each frame to isolate faces, feature extraction is applied to generate face embeddings for a reference photo and for faces in the live stream, and embeddings are compared using a similarity metric and threshold to identify an individual. In some implementations, continuity of identified individuals across frames is maintained using tracking algorithms such as SORT or Deep SORT. In some implementations, publicly available tools and / or libraries that can be used for facial recognition and / or related processing include OpenCV, Dlib, Facenet, DeepFace, OpenFace, and a Microsoft Azure Face API.

[0104] In some embodiments, the system 100 is configured to incorporate advanced methods for estimating 3D positions of individuals using at least one camera, including utilizing known environmental features (e.g., corners of a room or train, fixed objects, and / or a ground plane) as constraints for spatial localization. In some implementations, known 3D points in a scene and corresponding 2D image positions are used to calibrate camera intrinsic and extrinsic parameters, and a PnP algorithm is used to estimate 3D positions of points such as heads of people. In some implementations, depth ambiguity is mitigated by applying at least one constraint including a flat ground plane assumption and / or a head height assumption, and in some implementations both constraints are combined to constrain vertical and depth axes.

[0105] In some implementations, the system 100 accounts for seated individuals by selectively applying a sitting height assumption when a detected head position is lower than expected for a standing posture. In some implementations, sitting posture can be detected using pose estimation algorithms such as OpenPose or MediaPipe, and / or by detecting chairs via object detection and inferring a seated state based on chair detection. In some implementations, dynamic switching between standing and sitting assumptions is applied to improve 3D position estimation accuracy in mixed-posture environments.

[0106] Although the above description includes reference to certain specific embodiments, various modifications thereof will be apparent to those skilled in the art. Further, the specific embodiments described herein are not intended to be mutually exclusive. That is, one or more features of one embodiment may be included in another embodiment described herein unless specifically mentioned otherwise. Any examples provided herein are included solely for the purpose of illustration and are not intended to be limiting in any way. Any drawings 24CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041provided herein are solely for the purpose of illustrating various aspects of the description and are not intended to be drawn to scale or to be limiting in anyway. The scope of the claims appended hereto should not be limited by the preferred embodiments set forth in the above description but should be given the broadest interpretation consistent with the present specification as a whole.25CPST Doc: 1379-7987-3308.1

Claims

CPST Ref: 40736 / 00041CLAIMS1. A system for analyzing viewer engagement within an environment, the system comprising:a processor configured to:receive, from one or more visual data capturing systems communicatively coupled to the processor, one or more viewer engagement parameters associated with a viewer within the environment;generate, based on the one or more viewer engagement parameters, one or more performance metrics for the viewer, the one or more performance metrics being associated with one or more targets or points of interest within the environment; generate, based on the one or more performance metrics, one or more recommendations for content to be provided to the viewer; andtransmit the one or more recommendations to a display device associated with one or more content providers.

2. The system of claim 1 , wherein the one or more viewer engagement parameters include one or more of heat signatures, posture, gaze direction, head orientation, and target engagement duration.

3. The system of claim 1 , wherein the one or more visual data capturing systems include a video data capturing system configured to provide a first camera function configured to continuously scan the environment to identify one or more viewers showing interest in a target within the environment by monitoring a first set of viewer engagement parameters.

4. The system of claim 3, wherein the processor is configured to activate a second camera function of the video data capturing system to focus on the identified one or more viewers to monitor a second set of viewer engagement parameters associated with the identified one or more viewers.

5. The system of claim 4, wherein the second set of viewer engagement parameters includes gaze direction and an engagement duration associated with the identified one or more viewers.26CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 000416. The system of claim 3, wherein the first set of viewer engagement parameters includes heat signatures and head orientation.

7. The system of claim 3, wherein the processor is configured to activate a second camera function to identify one or more viewers obscured from the first camera function.

8. The system of claim 1 , further comprising a database, wherein the one or more recommendations are associated with a unique identification associated with the viewer and stored in the database for future usage.

9. The system of claim 1 , wherein the processor is configured to use an artificial intelligence model configured to process the one or more viewer engagement parameters and generate the one or more recommendations.

10. The system of claim 1 , wherein the processor is configured to determine gaze direction using an iris detection estimation method.

11. The system of claim 1 , wherein the viewer is located within a moving vehicle, and wherein the processor is configured to determine gaze direction as a viewer engagement parameter by:capturing, using the one or more visual data capturing systems, one or more video frames of the vehicle as the vehicle passes by a target; andanalyzing, using one or more computer vision modules, the one or more video frames to detect head pose of the viewer within the vehicle and to determine a gaze direction vector based on the detected head pose and one or more camera parameters.

12. The system of claim 11 , wherein the one or more camera parameters include one or more of focal length and lens distortion.

13. The system of claim 11 , wherein the processor is configured to align a coordinate system of a visual data capturing system with a coordinate representation of the target to map a location of the target relative to a position of the visual data capturing system.27CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 0004114. The system of claim 13, wherein the processor is configured to determine whether the gaze direction vector intersects a plane associated with the target.

15. The system of claim 14, wherein the processor is configured to log instances of intersection overtime to quantify viewer engagement with the target.

16. The system of claim 15, wherein the processor is configured to aggregate the logged instances overtime to assess an interest level associated with the target.

17. The system of claim 16, wherein the processor is configured to compare aggregated data from multiple target locations to evaluate effectiveness of respective target placements.

18. The system of claim 11 , wherein the processor is configured to adjust for gaze offset based on one or more behavioral parameters associated with the viewer to refine the gaze direction vector.

19. The system of claim 1 , wherein the processor is configured to generate a report including one or more of a count of viewers, demographics inferred from facial features or vehicle type, and attention levels.

20. The system of claim 1 , wherein the processor is configured to provide a graphical user interface via an output unit.

21. A method for analyzing viewer engagement within an environment, the method comprising:receiving, by one or more processors and from one or more visual data capturing systems, one or more viewer engagement parameters associated with a viewer within the environment;generating, by the one or more processors and based on the one or more viewer engagement parameters, one or more performance metrics for the viewer, the one or more performance metrics being associated with one or more targets or points of interest within the environment;28CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 00041generating, by the one or more processors and based on the one or more performance metrics, one or more recommendations for content to be provided to the viewer; andtransmitting, by the one or more processors, the one or more recommendations to a display device associated with one or more content providers.

22. The method of claim 21 , wherein receiving the one or more viewer engagement parameters includes receiving one or more of heat signatures, posture, gaze direction, head orientation, and target engagement duration.

23. The method of claim 21 , further comprising:continuously scanning the environment using a first camera function to identify one or more viewers showing interest in a target by monitoring a first set of viewer engagement parameters; andin response to identifying the one or more viewers, activating a second camera function to focus on the one or more viewers to monitor a second set of viewer engagement parameters.

24. The method of claim 23, wherein the first set of viewer engagement parameters includes heat signatures and head orientation, and wherein the second set of viewer engagement parameters includes gaze direction and engagement duration.

25. The method of claim 23, further comprising using the second camera function to identify one or more viewers obscured from the first camera function.

26. The method of claim 21 , further comprising associating the one or more recommendations with a unique identification associated with the viewer and storing the association in a database for future usage.

27. The method of claim 21 , further comprising using an artificial intelligence model to process the one or more viewer engagement parameters and generate the one or more recommendations.29CPST Doc: 1379-7987-3308.1CPST Ref: 40736 / 0004128. The method of claim 21 , further comprising determining gaze direction using an iris detection estimation method.

29. The method of claim 21 , wherein the viewer is located within a moving vehicle, and wherein receiving the one or more viewer engagement parameters includes:capturing one or more video frames of the vehicle as the vehicle passes by a target; andanalyzing the one or more video frames to detect head pose of the viewer within the vehicle and to determine a gaze direction vector based on the detected head pose and one or more camera parameters.

30. The method of claim 29, wherein the one or more camera parameters include one or more of focal length and lens distortion.

31. The method of claim 29, further comprising aligning a coordinate system of a visual data capturing system with a coordinate representation of the target to map a location of the target relative to a position of the visual data capturing system.

32. The method of claim 31 , further comprising determining whether the gaze direction vector intersects a plane associated with the target and logging instances of intersection overtime.

33. The method of claim 32, further comprising aggregating the logged instances over time to assess an interest level associated with the target and comparing aggregated data from multiple target locations to evaluate effectiveness of respective target placements.

34. The method of claim 29, further comprising adjusting for gaze offset based on one or more behavioral parameters associated with the viewer to refine the gaze direction vector.

35. The method of claim 21 , further comprising generating a report including one or more of a count of viewers, demographics inferred from facial features or vehicle type, and attention levels.30CPST Doc: 1379-7987-3308.1