Computer implementation methods, devices, and systems

The eye-tracking-based method addresses the lack of objective measures in reviewing digital content by generating scores based on gaze tracking and template comparison, ensuring subjects have adequately engaged with the content.

JP2026513526APending Publication Date: 2026-04-28XR SYNERGIES GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
XR SYNERGIES GMBH
Filing Date
2024-03-21
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing methods lack an objective measure to confirm that subjects have properly reviewed digital content, such as medical consent forms, leading to potential misunderstandings and consent to procedures without full understanding.

Method used

An eye-tracking-based method that tracks a subject's gaze while viewing digital content, generates a score by comparing viewed portions to a predetermined template, and determines if the subject has adequately engaged with the content using a sigmoid function to adjust for attention levels.

Benefits of technology

Provides an objective measure to ensure subjects have viewed and understood digital content, reducing the risk of misunderstandings and improving the reliability of consent processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026513526000001_ABST
    Figure 2026513526000001_ABST
Patent Text Reader

Abstract

This is a computer implementation method for determining whether digital content (104) has been viewed by a subject (102). The method includes: displaying the digital content to the subject (102) using a display device (106); tracking the subject's gaze and acquiring tracking data while displaying at least a temporary portion of the digital content (104); determining from the tracking data which portion of the display device (106) was viewed with respect to that temporary portion of the digital content (104); generating at least one score by comparing these portions with a predetermined display portion template for that temporary portion of the digital content (104); and determining, based on the at least one score, that the subject has viewed at least a temporary portion of the digital content. The display device (106) and system are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a device, system, and computer implementation method for determining whether or not digital content has been viewed by a subject. More specifically, the present invention provides an eye-tracking-based technique that can determine an objective measure indicating that a subject has viewed digital content. [Background technology]

[0002] There are many situations where information must be conveyed to subjects. For example, for various reasons such as safety or legal considerations, it is necessary to confirm that subjects have appropriately engaged with the content material, and to evaluate whether they have properly reviewed the information. One example is in the scenario of medical consent. Traditionally, patients consent to medical procedures based on explanations provided by clinicians. The problem with this approach is that there is no objective measure to infer whether patients have properly reviewed the medical procedures. As a result, patients may consent to medical procedures they do not fully understand.

[0003] Patent Document 1 discloses a method for obtaining informed consent from a patient for a medical procedure. In this method, a video explaining the medical procedure is displayed to the patient on a first display portion of a device, and images of the patient captured while the patient is watching the video are displayed on a second display portion of the device. The entire display is recorded to demonstrate that the patient has watched the video.

[0004] Patent Document 2 discloses a method for managing the electronic informed consent process in a clinical trial. This includes tracking the time participants spend reading content and comparing that time to a predetermined predicted time required for reading the content. If the difference between the time spent and the predetermined predicted time does not meet a time difference threshold, a deviation in the time spent is detected. This method includes triggering actions to mitigate the impact of such deviation on the electronic informed consent. In one example, each section of the consent form is compared to the predicted time for that section. If the deviation between these times does not meet a time difference threshold, an alert is triggered. Eye tracking is mentioned in the context of determining why a deviation in the time spent was determined. The time spent is tracked by either web tracking or content tracking (the content includes analytical code that tracks events performed by the patient). [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] U.S. Patent No. 11501875 [Patent Document 2] U.S. Patent Application Publication No. 2018 / 0102186 [Overview of the Initiative]

[0006] The present invention relates to a computer implementation method for determining whether digital content has been viewed by a subject. This method includes: displaying the digital content to the subject using a display device; tracking the subject's gaze and acquiring tracking data while displaying at least a temporary portion of the digital content; determining from the tracking data which portion of the display device was viewed for the temporary portion of the digital content; generating at least one score by comparing these portions with a predetermined display portion template for the temporary portion of the digital content; and determining, based on the at least one score, that the subject has viewed at least the temporary portion of the digital content.

[0007] Determining the display portion from the tracking data may include determining the display portion for each of multiple points in time (i.e., temporary portions) within the period.

[0008] Generating a score through comparison may involve comparing at least several points in the acquired tracking data with corresponding points in a given template.

[0009] The method may further include generating the predetermined template by tracking the gaze of at least one model subject while the digital content is being displayed and acquiring gaze tracking data.

[0010] The acquired tracking data may include positional data related to the subject's gaze for each of several timestamped points. Generating a score may include calculating an error value for each timestamped point by comparing the positional information at each timestamped point with corresponding positional data from a given template. Generating a score may also include calculating a cost value based on the magnitude of the error value. This method may also include generating a bounded score for each cost value. This score, i.e., the bounded score, may be relevant to the digital content as a whole or to a temporary portion thereof.

[0011] This method allows for the discarding of positional information related to a subject's gaze acquired at a specific point in time if corresponding positional information for the same point in time does not exist in a predetermined template.

[0012] In this method, the cost value may be converted to a bounded value using a sigmoid function, which may be the one shown in Equation 3. Advantageously, the adjustment value of the sigmoid function can be determined based on the cost value calculated for model subjects who were instructed not to carefully view the digital content prior to viewing it. Since it is unlikely that model subjects would carefully view the digital content without instruction, the accuracy of determining whether test subjects carefully viewed the digital content can be improved.

[0013] The method further includes the server transmitting a command to the display device to display the digital content in response to receiving a first identifier defining the subject and a second identifier defining the display device, and the method optionally further includes the server transmitting the digital content to the display device. The method may further include the steps of selecting the display device to display the digital content before the digital content is displayed by the display device, the subject device scanning a code provided to the display device, wherein the code includes or encodes the second identifier, and the subject device transmitting the second identifier to the server.

[0014] The method may further include the steps of generating a first identifier before the first identifier is transmitted to the display device, associating the digital content with the first identifier, and transmitting the first identifier, and at least the association between the digital content and the first identifier, to the server for storage.

[0015] The display device may be a virtual reality (VR) headset, an augmented reality (AR) headset, or a mixed reality (XR) headset, and the subject's gaze is tracked using sensors built into the VR, AR, or XR headset.

[0016] Provided the score is greater than a predetermined threshold, the method may include proving that the subject has the capacity to consent to a process related to the digital content. The process may be a medical procedure.

[0017] Provided the score is greater than a predetermined threshold, the method may include proving that the subject has successfully completed at least a portion of the training course to which the digital content is associated.

[0018] The present invention also provides a display device including a network interface for establishing a wireless connection, at least one processor coupled to the network interface, and a memory storing executable instructions. Here, the executable instructions are configured to operate the at least one processor so that the method of the present invention can be executed.

[0019] The present invention also provides a system including a server including a network interface for establishing a network connection, at least one processor coupled to the network interface, and a memory storing executable instructions. Here, the executable instructions are configured to operate the at least one processor so that the method of the present invention can be executed.

[0020] The present invention also provides a system. The system includes a server and a display device. The server includes a network interface for establishing a network connection, at least one processor coupled to the network interface, and a memory storing executable instructions. Here, the executable instructions are configured to operate the at least one processor to execute the method of the present invention. The display device includes a network interface for establishing a network connection, at least one processor coupled to the network interface, and a memory storing executable instructions. Here, the executable instructions are configured to operate the at least one processor to execute the method of the present invention.

[0021] The present invention further includes a computer program that causes a computer to execute the method of the present invention when executed on the computer.

Brief Description of the Drawings

[0022] Embodiments of the present invention will be described in detail below with reference to the drawings.

[0023] [Figure 1]This is a schematic diagram of a subject viewing digital content. [Figure 2] This is a schematic diagram illustrating the flow of digital information within the system. [Figure 3] This is a flowchart of the method. [Figure 4] This is a schematic diagram of the recorded eye movements and heatmaps of the subjects. It shows the parts of the digital content that the subjects viewed while the content was being displayed. [Figure 5] This is a flowchart of one example method. [Figure 6] This is a graph of the experimental results. [Figure 7] This is a graph of the experimental results. [Figure 8] This is a flowchart of the method. [Figure 9] This is a flowchart of the method. [Figure 10] This is a schematic diagram showing administrators and subjects viewing case reports. [Modes for carrying out the invention]

[0024] This disclosure relates to a device, system, and method for determining whether digital content has been viewed by a subject. In particular, this disclosure provides eye-tracking-based technology that can determine an objective measure indicating that a subject has viewed digital content. Various methods and means for eye tracking are known and fall within the scope of this disclosure. A brief introduction to eye tracking can be found at https: / / www.tobiidynavox.com / pages / what-is-eye-tracking, which is incorporated herein by reference.

[0025] In this method, digital content is displayed to subjects, and their gaze is tracked during this display. This information is used to determine which parts of the display device were viewed by the subjects. The viewed parts of the display can be stored in a map representing the subjects' gaze over the duration of the digital content. This map is then compared to a predetermined template for that digital content, and a score is generated based on the similarity between them. The predetermined template is a map of the subject's gaze over the duration of the digital content, i.e., a "model" subject, and is determined to represent that the subject viewed the digital content. A subject is determined to have viewed the digital content if this similarity score is greater than a predetermined threshold.

[0026] Figure 1 is a schematic diagram of a system 100 illustrating an aspect of the present invention. A subject 102 views digital content 104 displayed on a virtual reality (VR) headset 106. The digital content is in video format. However, other forms of digital content are also possible, such as still images or slideshows with or without audio content. In Figure 1, the projection of the digital content is shown for illustrative purposes only. Those skilled in the art will understand that the digital content is displayed on the display screen of the VR headset. The VR headset is fitted to the subject via fastening elements 108, such as a head strap.

[0027] Figure 2 is a schematic diagram showing the flow of digital information in system 200. The system includes a display device 202, a server 204, an administrator device 206, and a user device 208. The administrator device 206 and the user device 208 may be the same device, and these may generally be referred to as the user device, administrator device, or operator device. There may also be multiple user devices. The user device 208 and the administrator device 206 may consist of a mobile phone, tablet, computer, etc. As will be understood by those skilled in the art, the functions of the administrator and user devices may be performed by different devices through communication with the server 204, for example, by running the same application (app) locally or as a web application. Other devices, such as a supervisor device, may be connected to the server. In particular, it is intended that the system may include multiple display devices.

[0028] The display device 202, server 204, administrator device 206, and user device 208 each include a network interface for establishing a wireless connection (e.g., cellular, Wi-Fi, etc.), at least one processor coupled to the network interface, and memory for storing executable instructions. The executable instructions are configured to operate at least one processor so that at least one processor performs a specific method step. Multiple method steps, and in particular which of these devices performs a particular step, are described below in relation to Figures 3, 7, and 8.

[0029] Generally, System 200 provides a platform that facilitates a process of pairing a display device with a specific subject, configuring digital content on the display device for display to that subject, and determining, using an objective measure, whether or not the subject viewed the digital content. The determination results can be stored in Server 204, which functions as a repository for storing data. Such data can be uploaded or downloaded by an administrator or user device 208 using a web-based application. Examples of data that can be stored include multiple IDs, the association of these IDs with digital content, the digital content itself, and results from Figure 3 (e.g., a map, first and second scores).

[0030] Specifically, the server 204 may initiate the display of digital content by sending a command to the display device 202 to display the digital content, as shown in Figure 8. Alternatively, the display device may poll the server for the assigned digital content. The display device displays the digital content to the subject, records several results, and sends these results to the server (see Figure 3). The administrator and user devices 208 are used to pair the display device with the test patient, as shown in Figure 8. The administrator device is used to evaluate the results from Figure 3, as shown in Figure 9.

[0031] Figure 3 shows a method for determining whether or not digital content was viewed by a subject. This method provides an improved and objective scale based on eye tracking. From this scale, it is possible to determine whether or not a subject viewed the digital content.

[0032] In step 302, the digital content 104 is displayed to the subject using the display device 202. In the example shown in Figure 1, the digital content is displayed on a VR headset 106. Other display devices are also possible, such as AR headsets, XR headsets, smartphones, computers, tablets, and projector systems. The digital content may be displayed in an augmented reality, virtual reality, or mixed reality environment.

[0033] In one example, the digital content 104 includes a video. The video, for example, depicts a medical procedure being performed and explains the effects, risks, and benefits of the procedure. Optionally, audio content is played over the video during the display step 302. The display device 202 may include a speaker for this purpose. In one embodiment, the content 104 is a video, but the present invention is not limited thereto, and other forms of displayed content are also envisioned, including still images or slideshows with or without audio. In particular, the content may include an interactive virtual scene, if possible. The virtual scene may be recorded. Some or all of the virtual scene may be rendered in real time without pre-rendering, similar to a video. The virtual scene may allow a subject to move around or at least to move their head.

[0034] In step 304, the subject's gaze is tracked while the digital content 104 is displayed in step 302. Step 304 may be performed for the entire duration of step 302 or for a temporary portion thereof. That is, step 304 may be performed only for a portion of the duration of the digital content 104.

[0035] The process of eyeball or gaze tracking is known to those skilled in the art. Here, it is sufficient to state that a subject's gaze can be tracked over time by monitoring changes in light (e.g., infrared) reflected from specific features of the eye (e.g., corneal reflection or pupillary center). One or more light sources for illuminating the subject's eye and one or more camera sensors for detecting reflected light are arranged for the subject for this purpose. The light sources and / or camera sensors may be incorporated into the display device 202 or provided separately from the display device 202.

[0036] Optionally, eye-tracking of subject 102 includes monitoring the subject's head position while digital content 104 is displayed. Methods for monitoring head position and its effect on gaze direction are known to those skilled in the art.

[0037] The use of a VR, AR, or XR headset 106 for displaying digital content is particularly advantageous because head movements cause little to no relative movement to the display screen mounted on the display device 202. Therefore, generally, the portion of the display device, and thus the portion of the content, viewed by the subject at any given time depends solely on the position of the eyes. However, if the digital content spans multiple fields of view, the portion of the content viewed by the subject may change with head position, even if the portion of the display device viewed by the subject does not change. XR, VR, and AR headsets may include one or more built-in motion sensors, such as accelerometers or gyroscopes, to monitor head movements for this purpose.

[0038] XR, VR, and AR headsets 106 are particularly well-suited for performing steps 302 and 304 because they include a built-in light source and camera sensor positioned relatively close to the subject's eyes compared to other display devices such as smartphones. Therefore, the signal-to-noise ratio of the reflected signal captured in step 304 is relatively high for given light source and sensor specifications. Furthermore, when using a VR, XR, or AR headset, the positioning and orientation between the light source, the camera sensor, and each subject's eye are fixed. This simplifies the processing.

[0039] In step 306, the portion of the display device 202 viewed by the subject during display step 302 is determined through the subject's line of sight. The portion of the display includes one pixel or a cluster of pixels. The process of determining these portions of the display from the direction of line of sight is known to those skilled in the art. Step 306 is performed by the display device. Note that, depending on the frame rate of the digital content and the camera sensor, the subject may view multiple pixels or pixel clusters for each frame of the displayed digital content (the natural frame rate of the eye is approximately 30-60 Hz).

[0040] In one embodiment, the total number of times a subject views a particular display portion over the duration or a temporary portion thereof of display step 302 may be recorded, for example, in a matrix. The matrix may have dimensions equal to or less than the pixel dimensions of the display device and is initially empty (i.e., all 0s). The values ​​stored in the matrix are incremented when it is determined that the corresponding pixel index of the display device has been viewed by the subject. The resulting matrix can be visualized as a heatmap, which shows the total number of times the subject viewed a particular display portion of the display device.

[0041] Alternatively, the total number of times a subject views a particular display portion is recorded in a matrix for each frame or temporary portion of the digital content. The method of recording this information in a matrix is ​​substantially the same as described above. However, the difference is that a separate matrix is ​​generated for each frame of the digital content. Thus, multiple matrices form a set, and each matrix represents a time instance of the digital content. It will be understood that these matrices can be integrated into a higher-dimensional unitary matrix. The matrix set is a map that can determine which content portions were viewed in display step 302.

[0042] In step 308, a first score is generated by comparing the display portion determined in step 306 with a predetermined template. The predetermined template includes a map of a model subject's gaze over the duration of the digital content display. The predetermined template may represent a model subject or an ideal subject who viewed the digital content appropriately or attentively.

[0043] The generated template may be a heatmap showing the total number of times a model subject views a specific display portion of the display device over the duration or a temporary portion thereof of display step 302. Alternatively, the template may be a set of matrices showing the total number of times a model subject views a specific display portion for each frame (or a temporary portion thereof) displayed in step 302. Furthermore, results from multiple representative subjects may be used and “averaged” or otherwise combined to form a single ideal template result.

[0044] In one example, the display device 202 performs step 308, sending the first score and map to the server 204 for storage. For convenience, after step 304 or step 306 in Figure 3, the display device 202 sends or uploads the eye-tracking data acquired in step 304 to the server, and the subsequent steps are performed by the server 204. Any of steps 306 through 310 may be performed by either the display device or the server.

[0045] For example, the first score is expressed as a Sobolev norm, for example, h -1 It is calculated as the Sobolev norm. -1 Sobolevnorm d(A,B) 2 The formula for calculating this is shown below.

number

[0046] The Sobolev norm is particularly effective in this method because small differences in scale between maps have a far less impact on the calculated score than large differences in scale between maps. Therefore, the Sobolev norm is effective in removing variability in how different subjects view digital content. However, other methods are also possible for calculating the first score in step 308.

[0047] If the given template is a heatmap, step 308 includes generating a single Sobolev norm. The computational efficiency of step 308 is improved by integrating the viewed portion of the display into the heatmap in step 306. If the given template is a set of matrices, step 308 includes generating a Sobolev norm for each corresponding time instance; that is, a Sobolev norm is calculated for each frame of the displayed digital content. Using this approach, it is also possible to determine whether the subject viewed the appropriate portion of the content at the appropriate time.

[0048] Therefore, the first score generated in step 308 may be a single value or multiple values. However, a single value can also be derived from multiple values, for example, through averaging.

[0049] If the map generated in step 306 relates to a temporary portion of digital content, the corresponding temporary portion of a given template is used to generate the score in step 308.

[0050] In another embodiment, the subject's gaze is recorded from a start point to an end point and compared to a template record. The template record may be, for example, an actual record of a volunteer test subject, or it may be a composite template derived by "averaging" the results of previous actual records.

[0051] The subject's gaze is continuously recorded (e.g., frame by frame) during the playback of digital content. For each timestamp, a 2D vector (u,v) is stored along with the timestamp, where u represents the horizontal position (e.g., pixel column) and v represents the vertical position on the display device (e.g., pixel row). Below, ground truth (best) represents the optimal (model) recording of the gaze dataset for a particular video content, while measured (mes) gaze data represents the subject's recorded gaze for the same digital content (e.g., video). Thus, measurement relates to the subject's current results, while ground truth relates to template recordings.

[0052] Referring to Figure 5, the data processing scheme for generating the score in step 308 is shown.

[0053] As described above, the ground truth and measured gaze datasets contain a list of 2D vectors (u,v) for each timestamp of the displayed digital content. In the case of video, the timestamp corresponds to a specific frame. Data cleaning is performed before generating scores according to the scheme shown in Figure 5. Data cleaning is important because it has been found that the hardware of gaze tracking sensors (especially those built into XR, AR, and VR headsets) cannot reliably sample without errors. For example, over courses longer than one minute, current sensors tend to miss at least one data point in the measured or ground truth dataset. As a result, the measured and ground truth datasets may not sample the same time domain, and in particular, the measured dataset may sample the subject's gaze at timestamps that the ground truth dataset does not have.

[0054] To address this issue, data cleaning may include discarding measured gaze data acquired at timestamps for which the ground truth data does not have a corresponding (e.g., related to video frame j). That is, when the ground truth dataset does not have a 2D vector (u,v) at that timestamp. Furthermore, if the measured gaze dataset samples less than a predetermined percentage of the timestamps in the ground truth dataset, the score is set to 0 (i.e., the lowest value) to indicate that the subject did not view digital content.

[0055] In certain examples, the predetermined ratio is 0.05. For video content, this ensures that there is at least one sampling point per second (assuming the video frame rate is approximately 20 Hz). In some examples, the predetermined ratio may be larger, taking values ​​of 0.10, 0.20, 0.30, 0.40, 0.50, 0.60, 0.70, or 0.80, or generally in the range of 0.05 to 0.80. Those skilled in the art will understand that the value of the predetermined ratio can be adjusted according to the needs of the system. For example, the predetermined ratio can be a relatively large value because higher reliability of the eye-tracking sensor is expected to result in fewer discards.

[0056] Returning to Figure 5, the first step involves calculating an error value based on ground truth (i.e., a given template for a given temporal portion of digital content) and the measured gaze data set (i.e., the portion of that temporal portion viewed by the test subject). The error indicates the degree to which the ground truth and the measured gaze data differ. A larger error indicates a test subject who did not pay attention to the displayed digital content (or rather, did not pay attention to prominent features of the digital content), while a smaller error indicates a test subject who did pay attention to it. The error can be calculated as an L2 norm. For example, the error value can be calculated for each measured gaze data and ground truth data point in the data set (i.e., for each timestamp in the measurement data set). The error value would be a list of the respective error values ​​for each timestamp. Alternatively, the error value may be a single value obtained by summing all the calculated L2 norms.

[0057] In the second step, cost values ​​are calculated for the ground truth and the measured line-of-sight dataset. The cost value can be calculated by averaging all error values ​​in a list of error values, or, if the error value is a single value, by dividing it by the total number of timestamps in the measured line-of-sight dataset. The cost value output by the cost function is unbounded in that it can take any positive value.

[0058] In the third step, the cost values ​​are converted to a bounded range, such as a score between 0 and 10. Those skilled in the art will understand that the upper and lower bounds are arbitrary, in that the upper bound can take any value. The score can also be converted to other descriptions, such as a simple pass / fail or any other arbitrary grade. This is achieved by labeling different bounded score ranges. For example, if the bounded score range is from 0 to 10, the groups of scores can be named. For example, scores from 0 to 3 = "Fail", scores from 4 to 6 = "Partial Pass or Fail", and scores from 7 to 10 = "Pass". Different actions may be taken depending on where the subject's score falls on the scale.

[0059] An example of a cost function is the ground truth (u best ,v best ) Measured line of sight (u mes ,v mes This is the mean 2-norm (i.e., L2-norm) of the distance between the ground and the measured line-of-sight data over timestamps {n0, n1, n2, ..., N}. Specifically, the mean 2-norm between the ground and the measured line-of-sight data over timestamps {n0, n1, n2, ..., N} is calculated according to Equation 1. Formula 1:

number

[0060] The cost value (the higher, the worse, no upper limit) can be converted to a bounded score (the higher, the better) using a sigmoid function in the form shown in Equation 2 or 3. Equation 2:

Number

Number

[0061] The maximum score indicates that the test subject viewed the same display part as the model subject within the errors of the system and data processing scheme. That is, the display part viewed by the test subject is the same as that of a predetermined template.

[0062] Any convenient means may be used to convert the results of the cost analysis to a score. The cost function can be used in its raw form, but for example, it may be difficult to handle in terms of convenient display.

[0063] The score may be calculated for one or more temporal portions of the digital content. A bounded score allows the results to be easily presented in graph form and at any desired level of granularity. The format of the displayed content may suggest that different temporal portions of the displayed content have different levels of granularity. In this case, the final score can be calculated by taking a weighted average of the scores of each temporal portion. The weights can be determined based on the length (i.e., duration) of the temporal portion and / or the perceived importance or priority of each temporal portion. For example, if the digital content relates to a medical procedure, the temporal portion describing the risks of the medical procedure may be labeled as "high" importance. The weights increase as the relative length of the temporal portions increases and as the perceived importance of the temporal portions increases. For example, in the case of mixed content including static presentations and videos, the static portions require less detail, so the weight of the score corresponding to the static portions of the digital content will be relatively low. Alternatively, portions of the displayed content may have a higher priority and require a larger level of granularity.

[0064] When displaying or representing scores, raw data can be combined in any convenient way. For example, if a relatively low level of granularity is required, scores from multiple frames or timestamped points can be combined and averaged. Alternatively, the lowest or highest score can be taken to represent scores at multiple points in time over the represented period.

[0065] To generate a given template, robust ground truth is required so that the output provides a meaningful score. Multiple ground truth recordings for the same video can be combined to form robust ground truth. For example, they may be averaged (as shown in Equation 4) or used to define a bandwidth (range of viewing positions at each point in time) in which the line of sight trajectory may exist, resulting in zero cost within the bandwidth. Multiple acquired ground truth datasets can also be used to form a temporal bandwidth. The temporal bandwidth is the temporal domain from which the ground truth recordings are collectively sampled. For example, if two ground truth recordings are made in temporal domains {t0, t2, t4} and {t1, t3, t5}, the temporal bandwidth is temporal domain {t0, t1, t2, t3, t4, t5}. Formula 4:

number

[0066] Figure 6 shows the scores calculated by the sigmoid function represented by Equation 2, following the data processing scheme in Figure 5. In the experimental example, a total of six eye-tracking records were collected from a single male subject. All records with a score above a predetermined threshold (here set at 5) are considered to indicate that the subject was paying sufficient attention while viewing the digital content. Scores below this predetermined threshold indicate that the subject was not paying sufficient attention. As shown in Figure 6, one is labeled as ground truth (best case), two as examples of "poor" visual attention, two as examples of "intermediate" (int), and one is "unknown," but likely to be labeled as an "intermediate" attention case. As shown in Figure 6, one problem with the sigmoid function in Equation 2 is that it outputs scores close to the predetermined threshold. Therefore, the risk of false positives or false negatives is relatively high. This means that the digital content may need to be re-exposed to the test subject, wasting time and computational resources.

[0067] This is because the adjustment value is inherently highly dependent on the accuracy of the model subject's cost values. For example, if the model subject does not pay adequate attention to the displayed digital content, the adjustment value may be too low, potentially introducing false positives.

[0068] Figure 7 shows the scores calculated by the sigmoid function represented by Equation 3, according to the data processing scheme in Figure 5. A total of 37 eye-tracking records were collected in the experiment. Model subjects were divided into two groups: the first group (labeled Group A) was instructed to carefully view the digital content, and the second group (labeled Group B) was instructed not to carefully view the digital content. Group A corresponds to records with indices 0 through 13. Group B corresponds to records with indices 14 through 36.

[0069] The adjustment value thres in Equation 3 was determined as the average between the maximum cost value from group A and the minimum cost value from group B. In the specific example shown, the adjustment value is approximately 5.5. In other examples, the adjustment value could be set as the maximum cost value from group A (i.e., approximately 4.5 in Figure 7) or as the minimum cost value from group B (i.e., approximately 6.5 in Figure 7).

[0070] This approach is robust because model subjects are less likely to carefully browse digital content (uninstructed) than when they are instructed. This is advantageous because the adjustment thresholds depend on the lowest cost value from group B, allowing for more precise determination. The resulting sigmoid function is highly effective in distinguishing between test subjects who carefully browsed the digital content (i.e., those with a score greater than a predetermined threshold set to 5 in Figure 7) and those who did not.

[0071] Returning to Figure 3, in step 310, it is determined that the subject has (sufficiently) viewed the digital content (or a temporary portion thereof), provided that the first score is greater than a predetermined threshold.

[0072] The comparison between the first score and a predetermined threshold can be determined locally by the display device. Alternatively, the map determined in step 306 can be sent to the server 204 to perform steps 308 and 310.

[0073] If the first score includes multiple values, there may be a predetermined threshold corresponding to each score value, namely a predetermined threshold for each frame of the digital content. By comparing these respective score values ​​and thresholds, it is possible to infer which temporary portions of the digital content have not been viewed. These temporary portions can then be selected for redisplay as described in optional step 312.

[0074] If the first score generated in step 308 is based on the transient portion of the digital content, step 310 may further condition that the transient portion of the digital content is greater than 50%, 60%, 70%, 80%, 90%, or 95% of the duration of the digital content.

[0075] In an optional step 309a (not shown), one or more questions are displayed to the subject via the display device 202. One or more questions may be displayed concurrently with or consecutively with the execution of steps 302 to 308. The questions relate to the displayed digital content and test whether the subject has viewed and processed the displayed or displayed digital content. The display device provides means for inputting responses to the questions and means for recording these inputs. In one example, the input means is a microphone and the recording means is a storage medium. Other input means are also possible.

[0076] In an optional step 309b (not shown), a second score is generated based on the input received from the subject following step 309a. The second score indicates that the subject answered the questions correctly. In one example, the second score is the percentage of questions answered correctly. The second score can be calculated locally using a processing unit built into the display device 202. Alternatively, the input received by the display device is sent to a server 204 that calculates the second score. One or more questions to be displayed are pre-stored in the server 204.

[0077] Step 310 may be performed further by requiring that the second score be greater than a second predetermined threshold. As already mentioned, the second predetermined threshold indicates that the subject answered one or more questions correctly. In one example, the second predetermined threshold is in the range of 0.7 to 1, and the second score is the percentage of questions answered correctly.

[0078] Optionally, in step 312 (not shown), digital content or a temporary portion thereof is played to the subject. Following step 312, the method resumes in step 304. Step 312 is performed if the first or second score is below the first or second predetermined threshold, respectively. Playing a temporary portion of the digital content that the subject is determined not to have viewed, rather than the entire digital content, is advantageous because it reduces the computational cost of steps 304 through 308.

[0079] Figure 4 is a schematic diagram showing a visualization of the heatmap generated in step 306 for the digital content 104. The mapping shown is a heatmap 402 indicating how often the subject viewed a specific part of the display screen (in other words, the numbers represent the eye-tracking paths, i.e., the chronological order of recorded viewpoints). The heatmap is overlaid on the content portion of the digital content displayed in step 302. As with Figure 1, the projection of the digital content and heatmap is for illustrative purposes only.

[0080] Figure 8 is a flowchart of a method for pairing the display device 202 with a subject. Pairing the display device with the subject ensures that appropriate or correct digital content is subsequently displayed to them (e.g., in Figure 3). The method in Figure 8 may be implemented using a web application accessible by the administrator device 206 and / or the user device 208. The web application may be hosted by the server 204.

[0081] For example, when an administrator creates a new case, the administrator selects a display on which specific media content will be shown. The case ID (case or subject ID), content ID, and display device ID may be stored in server 204.

[0082] In one example, server 204 receives a request (e.g., from administrator device 206) and generates a case ID. In this example, the case ID is a random string containing alphanumeric characters. The case ID can be generated without receiving patient data or other forms of personal data related to the subject.

[0083] Server 204 sends the case ID to administrator device 206.

[0084] The administrator device 206 associates the case ID with a specific digital content item and a specific display device (using a display device ID such as a serial number).

[0085] The administrator device 206 transmits at least the case ID (related to the information), the associated content, and the serial number of the display device to the server 204. The associated content may be stored on another device, for example, and the server 204 only needs to know the associated information and location. However, the server 204 may store all the information. The display device 202 (which may be one of several display devices) may poll the server for the content associated with its serial number.

[0086] When media content display begins, for example, if server 204 is polled by display device 202 having the relevant content, server 204 sends the relevant content, or a command to retrieve the relevant content, to display device 202 for viewing by the subject on display device 202. Alternatively, server 204 may push the content to the display device after the initial handshake operation.

[0087] In step 502, an identifier (ID) (case ID) is generated on the server and associated with the digital content by the administrator device. The ID may uniquely define a subject, for example. In one example, the case ID is a random string containing alphanumeric characters. Preferably, for data security purposes, the case ID is generated without using personal data related to the subject. Step 502 may be performed by the administrator device 206 or the user device 208.

[0088] In step 504, the case ID and the association between the digital content and the case ID are sent to the server 204 for storage. Step 504 is performed by the administrator device 206. Optionally, step 504 further includes sending the digital content to the server 204. Alternatively, the server 204 may have the digital content stored in advance, and instead, a digital content identifier is sent to the server.

[0089] In step 506, the ID is sent, for example, to user device 208 or administrator device 206. Step 506 may be performed by server 204.

[0090] In one example, the ID is sent as a text message over a cellular network. The subject may provide or register their mobile phone number with the system for this purpose. Preferably, the subject's mobile phone number is deleted after the method steps in Figure 8 are completed, but this is not necessarily required.

[0091] In one example, the administrator device 206 is operated using a web-based application from steps 502 to 506. A text message optionally sent in step 506 may further include a link to the web-based application used by the administrator device in steps 502 to 506.

[0092] In step 508, the user device 208 sends the serial number that uniquely defines the display device and the ID received in step 506 to the server 204.

[0093] In one example, the user operates user device 208, and display device 202 is assigned to them. In another example, the user selects the display device. The serial number can be obtained by scanning a 2D code printed on the display device. In one example, the 2D code is a quick response (QR) code that encodes the serial number. Alternatively, the serial number itself is printed on the display device (unencoded), and the user can read it directly from the display device.

[0094] In one example, the user provides the serial number and ID to the server 204 through a web-based application. The user may also access the web application via a link provided to the user device 208 by the administrator device 206. The web-based application provides input fields for entering the serial number and ID of the display device. If the user scans a 2D code on the display device 202 using the camera, the serial number input field is automatically populated with the serial number after the 2D code has been processed.

[0095] In step 510, the server 204, in response to receiving the serial number and ID, determines which digital content 104 should be displayed to the subject using the ID. This is possible because, following step 504, the server has associated and stored the ID and the digital content (or digital content identifier) ​​with each other. After the method in Figure 8 is completed, any association between the serial number and ID of the display device can be removed.

[0096] In step 512, the server 204 sends a command to the display device 202 corresponding to the serial number to display the digital content 104 associated with the ID. Optionally, step 512 further includes sending the digital content to the display device. Alternatively, the digital content does not need to be sent as it is already stored in the display device. The server may maintain a record of the digital content stored by each display device for this purpose. This completes the pairing process. The method described in Figure 3 can be carried out after pairing is complete.

[0097] The numbering of the method steps is not intended to restrict the order in which those steps are performed. The steps can be performed in any other order. For example, step 504 may be performed after step 506, and so on.

[0098] The administrator 206 and user device 208 may be the same device, or any device running the web application. Both content assignment and display device 202 may be performed on the same administrator or user device.

[0099] In another embodiment, an administrator or user device is used for content allocation, while a subject device (not shown) is operated by the subject to allocate the display device.

[0100] As those skilled in the art will understand, the methods in Figures 3 and 5 are not limited to being applicable to any particular digital content. The method in Figure 3 has the advantage of objectively quantifying the extent to which a subject viewed digital content. As already mentioned, this score can be used to determine whether or not a subject viewed digital content. The score can also be used as a basis for performing other or additional steps or procedures.

[0101] One use case is for educational or training purposes. In the field of education, it is a well-known problem that children's attention spans vary. However, it is difficult to objectively determine which child has the shortest attention span, and more importantly, how long that attention span lasts on average. The approach of this disclosure provides a means to determine both which child has the shortest attention span (which is useful in itself for teachers) and their attention span. Scores generated according to the method of this disclosure can also be used to complement testing of the content being taught. For example, if a training course involves the use of machines, test subjects may only be allowed to practice using the machines after the method has determined that they have sufficiently viewed the digital content.

[0102] As those skilled in the art will understand, there are many other use cases, but another use case is medical consent procedures. In the scenario of medical consent procedures, the digital content relates to a medical procedure. The digital content may be a video depicting the medical procedure being performed, its effects, and the risks and benefits of the procedure. The digital content may also include audio content. The IDs referenced in Figure 8 are case IDs unique to both the subject and the medical procedure.

[0103] The subject may be the patient requiring the medical procedure, or the legal guardian of the person requiring the procedure. Before the medical procedure is performed, the subject and / or legal guardian must consent to the procedure. Therefore, following step 310, the subject and / or legal guardian may be provided with a medical consent form. The medical consent form can be displayed on the display device 202 or as a paper copy. This step is performed if the first score generated in step 308 is greater than a first predetermined threshold, and, if applicable, the second score generated in the optional step 309b is greater than a second predetermined threshold. The subject and / or legal guardian may then sign the medical consent form. The signature may be a digital signature entered into the display device 202, or a wet-ink signature. Thus, this process facilitates the medical consent procedure.

[0104] Figure 9 illustrates a method 600 in which a medical clinician or other qualified healthcare professional (user) verifies the results generated by the method in Figure 3 and proceeds with the medical treatment consent process based on those results. The user may or may not be the same person as the administrator.

[0105] In step 602, the case report is displayed by a user or administrator device. The user or administrator device may be the same as or different from user or administrator devices 206 and 208. For example, it may be a computer.

[0106] A case report may include the map generated in step 306, the first score generated in step 308, the optional second score generated in step 309b, the case ID generated in step 502, all answers to the questions given by the subject in the optional step 309a, and the medical consent form digitally signed by the subject obtained in step 312. For brevity, the details of these steps are not repeated here. If the medical consent form is a paper copy, it may be provided separately from the display device. Case report information can be downloaded from the server. Case reports can be viewed in the web-based application used in Figure 9.

[0107] In Step 604, the user evaluates the case report. The user may be a clinician or another qualified healthcare professional.

[0108] The evaluation in Step 604 includes one or more of the following: (i) verifying that the correct case ID is associated with the subject, and (ii) verifying that the first score and / or the second score are greater than their respective predetermined thresholds. Optionally, the user further evaluates the map displayed in the case report.

[0109] If the assessment in step 604 is positive, the user approves the subject's consent and proceeds to step 606 to sign the medical consent form. The medical consent form may be digitally signed, or, if a paper copy is provided, signed using a wet-ink signature.

[0110] In some cases, the method shown in Figure 9 is performed before step 312. In that case, both the supervisor and the user digitally sign the medical consent form in step 606.

[0111] Figure 10 is a schematic diagram showing clinician 702 and patient 102 viewing case report 704 on a computing device 706. The medical consent form 708 is a paper copy, which clinician 702 and test patient 102 sign following the verification procedure in Figure 9.

[0112] As will be apparent to those skilled in the art, various modifications are possible within the scope of the present invention. For example, processing steps performed locally by a display device may instead be performed by a server. Subjects do not necessarily have to be people receiving medical procedures. For example, a subject or test patient may be a parent or guardian of a child who requires such a medical procedure, or a person with guardianship rights over an individual who is legally unable to make such a decision. In such cases, the subject gives consent on behalf of the other person. The medical procedure may be related to a clinical trial for a drug or treatment arranged by a pharmaceutical company, and is not limited to surgery. Users and administrators may be staff of a pharmaceutical company.

Claims

1. A computer-implemented method for determining whether or not digital content has been viewed by a subject, The display device displays the digital content to the subject, Tracking the subject's gaze and acquiring tracking data while displaying at least a temporary portion of the digital content; From the aforementioned tracking data, determine the display portion of the display device that was viewed in the temporary portion of the digital content, By comparing these parts with a predetermined template of the display portion for the temporary portion of the digital content, at least one score is generated. Based on the aforementioned at least one score, it is determined that the subject viewed at least the temporary portion of the digital content. Methods that include...

2. The method of claim 1, wherein determining the display portion from the tracking data includes determining the display portion for each of a plurality of points in time within the period.

3. The method of claim 2, wherein generating a score by comparison includes comparing at least several points in time of the acquired tracking data with corresponding points in a predetermined template.

4. The method of claim 1, 2, or 3, further comprising tracking the gaze of at least one model subject while the digital content is being displayed, and generating the predetermined template by acquiring the gaze tracking data.

5. The method according to any one of claims 1 to 4, wherein the acquired tracking data includes, for each of a plurality of timestamped points, positional data related to the subject's gaze.

6. The method of claim 5, wherein generating a score includes calculating an error value for each timestamped point by comparing the positional information of each timestamped point with the corresponding positional data from the predetermined template.

7. Regarding each timestamped point that is present in the acquired tracking data but not in the predetermined template, The method of claim 6, wherein the positional information relating to the subject's line of sight corresponding to the timestamped point is discarded.

8. The method of claim 6 or 7, wherein generating the score includes calculating a cost value based on the error value.

9. The method of claim 8, comprising generating a bounded score from the cost value using a sigmoid function.

10. The method of claim 9, wherein the adjustment value of the sigmoid function is determined based on a cost value calculated for a model subject who was instructed not to carefully view the digital content prior to viewing the digital content.

11. The server further includes, in response to receiving a first identifier defining the subject and a second identifier defining the display device, transmitting a command to the display device to display the digital content, The method according to any one of claims 1 to 8, further optionally comprising the server transmitting the digital content to the display device.

12. Before the digital content is displayed by the display device, The steps include selecting a display device to display the aforementioned digital content, The subject device scans a code provided to the display device, wherein the code includes or encodes the second identifier. The subject device transmits the second identifier to the server. The method of claim 11, further comprising:

13. Before the first identifier is transmitted to the display device, The steps of generating the first identifier and The step of associating the digital content with the first identifier, The steps include transmitting the first identifier and at least the association between the digital content and the first identifier to the server for storage. The method of claim 9 or 10, further comprising the following:

14. The method according to any one of claims 1 to 13, wherein the display device is a virtual reality (VR), augmented reality (AR), or mixed reality (XR) headset, and the subject's gaze is tracked using sensors incorporated in the VR or AR headset.

15. Provided that the score is greater than a predetermined threshold, the method is as follows: The method according to any one of claims 1 to 14, comprising proving that the subject has the capacity to consent to a process relating to the said digital content.

16. The method of claim 15, wherein the process is a medical procedure.

17. Provided that the score is greater than a predetermined threshold, the method is as follows: The method according to any one of claims 1 to 16, comprising proving that the subject has successfully completed at least a portion of a training course to which the digital content relates.

18. A display device, A network interface for establishing a wireless connection, At least one processor coupled to the network interface, Memory that stores executable instructions and Includes, A display device wherein the executable instruction is configured to operate at least one processor so that the method of any one of claims 1 to 17 can be performed.

19. It is a system, Server and Display devices and Includes, The aforementioned server, A network interface for establishing a network connection, At least one processor coupled to the network interface, A memory for storing executable instructions configured to operate the at least one processor in order to perform the method of claim 1, Includes, The aforementioned display device is A network interface for establishing a network connection, At least one processor coupled to the network interface, A memory for storing executable instructions configured to operate the at least one processor in order to perform the method of any one of claims 1 to 17 A system that includes this.

20. A computer program that, when executed on a computer, causes the computer to perform the method described in any one of claims 1 to 17.

Citation Information

Patent Citations

  • Informed patient consent through video tracking

    US11501875B1

  • Method and system for managing electronic informed consent process in clinical trials

    US20180102186A1