Method and system for dynamically aligning picture-in-picture (PIP) window in a display unit
The electronic device dynamically aligns the PIP window using multimodal interaction and reinforcement learning to address alignment issues, improving usability and content legibility by optimizing placement based on user presence and habits.
Patent Information
- Application Number
- PCT/KR2024/006846
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-12
- Filing Date
- 2024-05-21
- Publication Date
- 2025-08-21
AI Technical Summary
Existing PIP window alignment methods in electronic devices often result in obstructed visuals, limited hardware compatibility, and compromised content legibility due to improper alignment, leading to reduced usability and inconsistent functionality across devices.
An electronic device employs multimodal interaction and reinforcement learning to detect user presence, determine user height and distance, divide the display into zones, and align the PIP window based on the user's line of sight, using sensors and AI models to optimize positioning over time.
Enhances user interaction and viewing experience by ensuring clear and readable content without compromising the main content, adapting to user habits for optimal PIP window placement.
Smart Images

Figure KR2024006846_21082025_PF_FP_ABST
Abstract
Description
METHOD AND SYSTEM FOR DYNAMICALLY ALIGNING PICTURE-IN-PICTURE (PIP) WINDOW IN A DISPLAY UNIT
[0001] The present disclosure generally relates to a method for an electronic device. In particular, the present disclosure discloses a method for dynamically aligning picture-in-picture (PIP) window in a display unit of the electronic device.
[0002] Currently, many electronic devices have an inbuilt picture-in-picture (PIP) feature. For example, the PIP feature can be found in electronic devices like television sets, computer monitors, and the like. Such electronic devices may be alternatively referred to as display devices. The PIP feature allows the viewer to concurrently watch two different sources of content on the screen at the same time. Typically, the main source of content is displayed on full screen, while a smaller PIP window shows a secondary source.
[0003] The PIP feature is commonly utilized for tasks such as watching a television (TV) show or a movie in the main window and monitoring a sports game or news broadcast in the smaller window. In computer monitors, the PIP feature allows users to work on a larger window while simultaneously keeping an eye on a smaller window (e.g., a PIP window) playing a video or displaying updates from a different application. Further, in smart home appliances, like smart refrigerators, the PIP feature allows the user to share calendars, photos, and notes, as well as stream music and videos. Some smart refrigerators also have built-in cameras to allow users to view contents inside their refrigerators remotely, and it is equipped with various smart home features. In some cases, smart home appliances may be further designed to be a central hub for family communication, organization, and entertainment in the kitchen. Thus, the PIP feature enhances multitasking and can be beneficial for activities that require monitoring multiple sources of content simultaneously.
[0004] Figure 1 illustrates an example of the PIP window in the television (TV). As depicted, the TV includes a display unit enabled with the PIP feature. The user may enable the PIP feature to display the PIP window. Further, the user may operate upon the PIP window for using various features as discussed in the above paragraphs.
[0005] At present, the user has the capability to drag and drop a PIP window to any location on the screen. However, when the user views the PIP window it usually does not align with the user's viewing angle. Accordingly, due to a fixed alignment of the PIP, the user has to manually adjust the PIP's window position to align with their viewing angle for optimal visibility.
[0006] To address the issue mentioned earlier, several solutions were proposed. These include software control-based solutions, customization options-based solutions, and user interaction-based solutions. In software control-based solutions, display devices are equipped with built-in software controls or settings that allow users to adjust the position, size, and content of the PIP window. These settings are typically accessible through the display device's menu, a remote control, or a software interface. Customization options-based solutions involve providing users with the capability to personalize the alignment of the PIP window according to their preferences, including choosing its position, size, transparency, and other visual aspects. Lastly, user interaction-based solutions enable users to interact with the PIP window using various input methods such as touch gestures, mouse controls, or remote control navigation, allowing them to move, resize, or toggle the PIP window on or off as needed.
[0007] The aforementioned solutions give rise to limitations such as obstructed visuals, limited hardware compatibility, unreliable support, and compromised content legibility. For instance, an improperly aligned PIP window can obscure essential elements of the main content, leading to reduced visibility and usability. Additionally, the restricted hardware compatibility hinders the implementation of advanced PIP alignment features, such as dynamic tracking or environmental sensing. The lack of necessary sensors or processing capabilities can also limit the range of PIP alignment options.
[0008] Moreover, not all devices or software applications support PIP window alignment or provide the same level of flexibility and control. Compatibility issues or restrictions in certain devices or platforms can constrain the availability and functionality of PIP window alignment features. When the PIP contains text or detailed visuals, improper alignment or small size can adversely impact legibility. Therefore, the placement of the PIP window should be carefully considered to ensure clear and readable content without compromising the user experience.
[0009] Thus, there is a need to provide a methodology to overcome the above-mentioned issues.
[0010] According to an embodiment of the disclosure, a method for dynamically aligning Picture-In-Picture (PIP) window in a display unit of a device is provided. According to an embodiment of the disclosure, a method performed by an electronic device may include detecting at least one of a user in a proximity of the device, and a user's engagement with the display unit of the device. According to an embodiment of the disclosure, a method performed by an electronic device may include determining information associated with the user including at least one of a height of the user or a distance of the user from the display unit. According to an embodiment of the disclosure, a method performed by an electronic device may include dividing, based on the information associated with the user, a zone of the display unit into the plurality of display zones. According to an embodiment of the disclosure, a method performed by an electronic device may include determining, using one or more image sensors, a user's line of sight with respect to the plurality of display zones of the display unit. According to an embodiment of the disclosure, a method performed by an electronic device may include determining, using a reinforcement learning network based on the user's line of sight, a first display zone among the plurality of display zones. According to an embodiment of the disclosure, a method performed by an electronic device may include aligning the PIP window over the determined first display zone.
[0011] According to an embodiment of the disclosure, an electronic device for dynamically aligning Picture-In-Picture (PIP) window in a display unit of the device is provided. According to an embodiment of the disclosure, the electronic device may include at least one processor and at least one memory storing computer executable instructions. According to an embodiment of the disclosure, at least one processor is configured to detect at least one of a user in a proximity of the device, and a user's engagement with the display unit of the device. According to an embodiment of the disclosure, at least one processor is configured to determine information associated with the user including at least one of a height of the user or a distance of the user from the display unit. According to an embodiment of the disclosure, at least one processor is configured to divide, based on the information associated with the user, a zone of the display unit into the plurality of display zones. According to an embodiment of the disclosure, at least one processor is configured to determine, using one or more image sensors, a user's line of sight with respect to the plurality of display zones of the display unit. According to an embodiment of the disclosure, at least one processor is configured to determine, using a reinforcement learning network based on the user's line of sight, a first display zone among the plurality of display zones. According to an embodiment of the disclosure, at least one processor is configured to align the PIP window over the determined first display zone.
[0012] These and other features, aspects, and advantages of the present invention will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:
[0013] Figure 1 illustrates an example of the PIP window in the television (TV);
[0014] Figure 2 illustrates an example environment of PIP window alignment in an electronic device, according to an embodiment of the present disclosure;
[0015] Figure 3 illustrates an exemplary architecture of the electronic device, according to an embodiment of the present disclosure;
[0016] Figure 4A illustrates a high-level architecture of the electronic device of Figure 3, according to an embodiment of the present disclosure;
[0017] Figure 4B illustrates an example architecture of the multimodal interaction module, according to an embodiment of the disclosure;
[0018] Figure 4C illustrates an example architecture of the display zone division module, according to an embodiment of the disclosure;
[0019] Figure 4D illustrates an example architecture of the gaze detection controller module, according to an embodiment of the disclosure;
[0020] Figure 4E illustrates an example architecture of the reinforcement learning agent module, according to an embodiment of the disclosure;
[0021] Figure 4F illustrates an example architecture of the PIP processing unit, according to an embodiment of the disclosure;
[0022] Figure 5 illustrates an operational flow for dynamically aligning the PIP window in the display unit of the electronic device, according to an embodiment of the present disclosure;
[0023] Figure 6 illustrates a flow chart of the operation flow of Figure 5, according to an embodiment of the present disclosure;
[0024] Figure 7 illustrates an example of dimensions and viewing distance (D) of the display panel, according to an embodiment of the present disclosure;
[0025] Figure 8 illustrates an example use case for dynamic PIP window alignment, according to an embodiment of the present disclosure; and
[0026] Figure 9 illustrates an example of a use case for dynamic PIP window alignment, according to an embodiment of the present disclosure.
[0027] Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help to improve understanding of aspects of the present invention. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present invention so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.
[0028] It should be understood at the outset that although illustrative implementations of the embodiments of the present disclosure are illustrated below, the present invention may be implemented using any number of techniques, whether currently known or in existence. The present disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary design and implementation illustrated and described herein, but may be modified within the scope of the appended claims along with their full scope of equivalents.
[0029] The term "some" as used herein is defined as "none, or one, or more than one, or all." Accordingly, the terms "none," "one," "more than one," "more than one, but not all" or "all" would all fall under the definition of "some." The term "some embodiments" may refer to no embodiments, to one embodiment or to several embodiments or to all embodiments. Accordingly, the term "some embodiments" is defined as meaning "no embodiment, or one embodiment, or more than one embodiment, or all embodiments."
[0030] The terminology and structure employed herein is for describing, teaching, and illuminating some embodiments and their specific features and elements and does not limit, restrict, or reduce the spirit and scope of the claims or their equivalents.
[0031] More specifically, any terms used herein such as but not limited to "includes," "comprises," "has," "consists," and grammatical variants thereof do NOT specify an exact limitation or restriction and certainly do NOT exclude the possible addition of one or more features or elements, unless otherwise stated, and furthermore must NOT be taken to exclude the possible removal of one or more of the listed features and elements, unless otherwise stated with the limiting language "MUST comprise" or "NEEDS TO include."
[0032] Whether or not a certain feature or element was limited to being used only once, either way, it may still be referred to as "one or more features" or "one or more elements" or "at least one feature" or "at least one element." Furthermore, the use of the terms "one or more" or "at least one" feature or element does NOT preclude there being none of that feature or element, unless otherwise specified by limiting language such as "there NEEDS to be one or more . . . " or "one or more element is REQUIRED."
[0033] Unless otherwise defined, all terms, and especially any technical and / or scientific terms, used herein may be taken to have the same meaning as commonly understood by one having ordinary skill in the art.
[0034] Embodiments of the present invention will be described below in detail with reference to the accompanying drawings.
[0035] According to an embodiment, the present disclosure discloses a method for dynamically aligning Picture-In-Picture (PIP) window in a display unit of an electronic device. According to an embodiment, the electronic device utilizes multimodal interaction and engagement to detect user presence in front of the electronic device and utilizes their activities as triggers. The electronic device further involves identifying the user's height and partitioning the display area of the display unit. The electronic device further determines an optimal screen view angle and adjusts the PIP window position to enhance the viewing experience of the user. Additionally, the electronic device employs reinforcement learning models to understand and adapt to the user's viewing habits over time, and gatherers behavioral data to notify the electronic device about the PIP window positioning. The disclosed methodology enhances the user interaction and the display of the PIP content.
[0036] The detailed methodology and architecture are explained in the following paragraphs.
[0037] Figure 2 illustrates an example environment of PIP window alignment in an electronic device, according to an embodiment of the present disclosure. Figure 2 illustrates an environment 200 including the electronic device 201. In a non-limiting example, the environment 200 may be an Internet of Things (IoT) environment having a plurality of IoT devices. As an example, the IoT devices may include but are not limited to, a smart television (TV), a smartphone, a smart refrigerator, smart home devices, and the like. In the example embodiment, the IoT device that is depicted here is the smart refrigerator 201 equipped with the PIP features. According to some embodiment, the IoT device may be any electronic device having a display unit adapted to display the PIP window 203. According to an example embodiment, the smart refrigerator 201, at block 205, detects the user's proximity with the smart refrigerator 201. In particular, the smart refrigerator 201 detects that the user is standing in front of the smart refrigerator 201. Further, at block 207, the smart refrigerator 201, detects the user's engagement with the display unit. Further, the smart refrigerator 201, based on the user's engagement, determines information associated with the user. The information associated with the user includes at least one of a height of the user or a distance of the user from the display unit. Further, at block 209, the smart refrigerator 201 divides, based on the information associated with the user, a zone of the display unit into the plurality of display zones. Further, at block 211, the smart refrigerator 201 aligns the PIP window 203 over a first display zone for displaying the PIP window 203. According to an embodiment of the disclosure, if the user moves first panel to second panel in case of the dual panel smart refrigerator 201, the PIP window 203 may be aligned to first panel to second panel dynamically. The first display zone is a zone by which the user could watch the PIP window most comfortably considering the user's line of sight. The smart refrigerator 201 is the IoT device. Further, the IoT device may be alternately referred to as the display device 201 or the electronic device 201 throughout the disclosure. A detailed implementation of the aligning of the PIP window will be explained in the forthcoming paragraphs.
[0038] Figure 3 illustrates an exemplary architecture of the electronic device, according to an embodiment of the present disclosure. The electronic device 201 includes a processor(s) 301, a memory 303, a module(s) 307, a database 309, an Audio / Video (AV) unit 305, and a network interface (NI) 311 coupled with each other. In an embodiment, the electronic device 201 may be further connected with a router (not shown) for internet connection. Further, the reference numerals are kept the same for similar components for ease of understanding.
[0039] For example, the processor 301 may be a single processing unit or a number of units, all of which could include multiple computing units. The processor 301 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, logical processors, virtual processors, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. Among other capabilities, the processor 301 is configured to fetch and execute computer-readable instructions and data stored in the memory 303 respectively.
[0040] The memory 303 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory, such as static random access memory (SRAM) and dynamic random access memory (DRAM), and / or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes.
[0041] As an example, the module(s) 307 may include a program, a subroutine, a portion of a program, a software component, or a hardware component capable of performing a stated task or function. As used herein, the module(s) 307 may be implemented on a hardware component such as a server independently of other modules, or a module can exist with other modules on the same server, or within the same program. The module(s) 307 may be implemented on a hardware component such as processor one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. The module(s) 307, when executed by the processor 301 respectively may be configured to perform any of the described functionalities.
[0042] As a further example, the database 309 may be implemented with integrated hardware and software. The hardware may include a hardware disk controller with programmable search capabilities or a software system running on general-purpose hardware. The examples of the database 309 include but are not limited to, in-memory databases, cloud databases, distributed databases, embedded databases, and the like. The database 309, amongst other things, serves as a repository for storing data processed, received, and generated by one or more of the processors, and the modules / engines / units.
[0043] In an embodiment, the module(s) 309 may be implemented using one or more AI modules that may include a plurality of neural network layers. Examples of neural networks include but are not limited to, Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Network (RNN), and Restricted Boltzmann Machine (RBM). Further, 'learning' may be referred to in the disclosure as a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to decide or prediction. Examples of learning techniques include but are not limited to supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. At least one of a plurality of CNN, DNN, RNN, RMB models and the like may be implemented to thereby achieve execution of the present subject matter's mechanism through an AI model. A function associated with an AI module may be performed through the non-volatile memory, the volatile memory, and the processor. The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). One or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.
[0044] As an example, the AV unit 305 outputs the audio / video content associated with the PIP window 203. As a further example, the NI unit 311 establishes a network connection with a network like a home network, a public network, a private network, or other IoT devices.
[0045] Figure 4A illustrates a high-level architecture of the electronic device of Figure 3, according to an embodiment of the present disclosure. In an embodiment, the module(s) 307 of the electronic device 201 further includes a multimodal interaction module 403, a display zone division module 405, a reinforcement learning agent 409, and a PIP processing unit 411 are coupled and collectively operate with each other.
[0046] According to a further embodiment, the electronic device 201 may be further coupled with at least one of an Artificial Intelligence (AI) engine 417, a graphical processing unit (GPU) 415, the database 309, and a display panel 413. According to an embodiment, the Artificial Intelligence (AI) engine 417 and the graphical processing unit (GPU) 415may be implemented within the electronic device 201.
[0047] According to an embodiment, various functions of the module(s) 307 can be performed by the processor 301 of Figure 3. However, for ease of understanding, an explanation is provided with respect to various modules. In embodiment module(s) may be a set of instructions that may be stored in memory. The processor executes the set of instructions thereby operating these modules.
[0048] According to an embodiment, the database 309 may further include at least one of data related to user recognition 419, user information 421, user's state 423, rules 425, statistics / usage information 427, training / testing data 429, and user's habit data of PIP alignment 431.
[0049] According to an embodiment, the display panel 413 includes at least a user interaction module 413-1, a camera 413-2, depth sensors 413-3, UWB sensor 413-4, Wi-Fi CSI 413-5, and I / O interfaces 413-6. The display panel 413 may be alternately referred to as the display unit throughout the disclosure. A brief working of each of the modules will be described in the forthcoming paragraphs.
[0050] Figure 4B illustrates an example architecture of the multimodal interaction module 403, according to an embodiment of the disclosure. The multimodal interaction module 403 may include multimodal sensors data fusion and processing module 4032. The multimodal interaction module 403 may include an engagement analysis module 4034. The multimodal interaction module 403 may include the user's height and distance measurement module 4036.
[0051] According to an embodiment, the multimodal interaction module 403 detects the user's presence using various sensors such as the cameras 413-2, the depth sensors 413-3, the microphones (not shown), the motion sensors (not shown), and the ultra wide-band (UWB) sensors 413-4 placed around the display panel 413. The display panel 413, captures using the aforesaid sensors, one or more images of one or more objects in a field of view of the one or more image sensors as an input 401. Further, the multimodal interaction module 403 detects the user's body or face by using the input 401. In particular, the multimodal interaction module 403 analyzes motion patterns in the input or uses computer vision techniques to detect the user's body or the face. Further, the multimodal interaction module 403 detects the user's engagement by tracking hand gestures, facial expressions, voice commands, or other actions performed by the user. Further, based on the data from the depth sensors 413-3 or the cameras 413-2 which are equipped with distance measurement capabilities or UWB sensors are used to, estimate the user's height and distance from the display panel 413 by analyzing a size and a position of the user's body in the captured image or depth maps by the cameras 413-2.
[0052] Figure 4C illustrates an example architecture of the display zone division module 405, according to an embodiment of the disclosure. The display zone division module 405 may include zone division algorithm module 4052. The display zone division module 405 may include zone mapping module 4054. The display zone division module 405 may include user height and posture tracking module 4056. The display zone division module 405 may include real time zone adjustment module 4058.
[0053] According to an embodiment, the display zone division module 405 determines display zones as per the user's height. The display zone division module 405 is implemented with a display zone division technique. The display zone division technique takes into account the user's height and the user's distance from the display panel 413. In an embodiment, the display zone division technique considers various factors such as a visual comfort of the user, legibility & ergonomic guidelines to determine the appropriate size and placement of each display zone. The display zone division module 405 further maps the determined display zones onto the display panel 413 indicating the boundaries and position of each zone. Further, the cameras 413-2, the depth sensors 413-3, the microphones, the motion sensors, and the ultra-wide-band (UWB) sensors 413-4 of the display panel 413 continuously track the user's height and posture change to dynamically adjust the display zone position in real-time. This ensures that the optimal zones align with the user's changing height and posture, providing a comfortable viewing experience to the user.
[0054] Figure 4D illustrates an example architecture of the gaze detection controller module 407, according to an embodiment of the disclosure. The gaze detection controller module 407 may include gaze tracking module 4072. The gaze detection controller module 407 may include image processing and eye tracking module 4074. The gaze detection controller module 407 may include user's angle of view estimation module 4076. The gaze detection controller module 407 may include calibration and feedback module 4078.
[0055] According to an embodiment, the gaze detection controller module 407 includes a camera or infrared sensor (not shown) to capture an image or a video feed of the user. In an embodiment, the gaze detection controller module 407 tracks the user's eye movements and gaze direction. The gaze detection controller module 407 captures a position of the user's eye relative to the display panel 413. In particular, the captured image or video feed is processed using computer vision algorithms to detect the position of the user's eyes and estimate the gaze vectors. By analyzing the gaze vectors, the gaze detection controller module 407 estimates the user's angle of view which further provides information about the portion in the display panel 413 that the user is currently looking at in the display panel 413. The gaze detection controller module 407 further incorporates a user feedback mechanism to validate and calibrate the gaze tracking accuracy. This can involve asking the user to focus on a specific point on the display panel 413 or performing calibration exercises to improve the tracking precision.
[0056] Figure 4E illustrates an example architecture of the reinforcement learning agent module, according to an embodiment of the disclosure. The reinforcement learning agent module 409 may include user interaction module 4091. The reinforcement learning agent module 409 may include state representation module 4093. The reinforcement learning agent module 409 may include learning and updating module 4095. The reinforcement learning agent module 409 may include habit module 4097. The reinforcement learning agent module 409 may include feedback and reward module 4099.
[0057] According to an embodiment, the reinforcement learning (RL) agent module 409, is implemented with an RL technique that focuses on enabling an agent to learn optimal actions or behaviors by interacting with an environment. It is inspired by the principles of behavioral psychology, where an agent receives feedback in the form of rewards or punishments based on its actions. In an embodiment, the RL agent module 409 collects data on the user's viewing behavior and utilizes the user's viewing behavior for the positioning of the PIP window. In an embodiment, the RL agent module 409 captures user inputs such as PIP window alignment preferences or actions performed by the user during the PIP window alignment. The RL agent module 409 processes the captured user inputs and the current state of PIP window alignment to create a state representation. The state representation may include features like a user's position, the PIP window position, a screen size, and other relevant parameters. Further, the RL agent module 409 learns and updates the PIP window alignment over time by optimizing its actions based on the observed user habits. Further, the RL agent module 409 maintains a habit model that captures the learned behavior and preferences of the user regarding PIP window alignment. This model helps agents in making more accurate decisions based on the user's habits. Furthermore, the RL agent module 409 provides feedback to the agent based on the user's preference response or a predefined reward signal. Positive feedback or rewards are given when the PIP alignment matches the user's preferences or habitual behavior, while negative feedback or penalties may indicate misalignment. The RL agent module 409 based on the analyzed user's parameters unit calculates the required adjustments / alignments for the PIP window to display over the display panel 413. Further, the RL agent module 409 is based on the user's line of sight, the user's behavior, and the user's preferences, which determines the first display zones.
[0058] Figure 4F illustrates an example architecture of the PIP processing module, according to an embodiment of the disclosure. The PIP processing module 411 may include coordinate conversion module 4112. The PIP processing module 411 may include alignment calculation and user parameter analysis module 4114. The PIP processing module 411 may include display control module 4116.
[0059] According to an embodiment, the PIP processing unit 411, based on the analyzed user's parameters, calculates the required adjustments / alignments for the PIP window to display on the determined first display zone of the display panel 413.
[0060] The user's parameter may include at least one of the user action, contextual information relating to content preferred by the user, or timestamp for determining the first display zone.
[0061] A detailed working explanation of the various modules of Figure 4 will be explained in detail through Figures 3 to 8 in the forthcoming paragraphs.
[0062] Figure 5 illustrates an operational flow for dynamically aligning the PIP window in the display unit of the electronic device, according to an embodiment of the present disclosure. The operation flow 500 is implemented in the electronic device 201 and will be explained through various operation steps 501 to 525. Further, Figure 6 illustrates a flow chart 600 of the operation flow 500 and hence will be explained collectively with the operation flow 500 for the sake of brevity and ease of reference. Accordingly, an explanation of the operation flow 500 will be explained in the forthcoming paragraphs and through Figures 3 to 8. Further, the reference numerals were kept the same for the similar components throughout the disclosure for ease of explanation and understanding.
[0063] According to an embodiment, at step 501, the display panel 413, captures using the one or more image sensors, one or more images of one or more objects in a field of view of the one or more image sensors as the input 401. In an embodiment, the input is provided to the multimodal interaction module 403. Further, at block 503, the multimodal interaction module 403 detects at least one of a user in a proximity of the electronic device 201 at block 505, and a user's engagement with the display unit of the electronic device 201. The detection of the user in the proximity of the electronic device 201, and the user's engagement with the display unit of the electronic device 201 will be explained in detail below.
[0064] According to an embodiment, the detection of the user in the proximity of the electronic device 201 can be determined in various ways. According to an embodiment, the detection of the user in the proximity of the electronic device 201 can be detected by using image data. According to an embodiment, the multimodal interaction module 403 first applies preprocessing techniques such as noise reduction, contrast adjustment, or image filtering to improve an image quality of the one or more images. In a non-limiting example, techniques like Gaussian smoothing, histogram equalization, or adaptive thresholding can be used. Further, the multimodal interaction module 403 uses a pre-trained Machine Learning (ML) model to detect the user from the one or more objects in the one or more images. In a non-limiting example, the pre-trained human detection ML model such as Haar Cascades, Histogram of Oriented Gradients (HOG), or a deep learning-based model (e.g. Faster R-CNN, YOLO) can be used as the pre-trained model to detect users in the one or more images. The multimodal interaction module 403 further uses a detection ML model to identify Regions of Interest (ROIs) where the users are likely to be present. The multimodal interaction module 403, by using the detection ML model, extracts the bounding box coordinates (x_min, y_min,x_max,y_max) or pixel-level segmentation masks for the detected users. Thereafter, the multimodal interaction module 403 defines a bounding box on the detected user in the one or more images by using the pre-trained ML model. Further, based on the defined bounding box on the detected user, the multimodal interaction module 403 estimates the proximity of the user. In particular, according to an embodiment, the multimodal interaction module 403 estimates the proximity of the user by calculating a centroid of the bounding box. As an example, the centroid of the bounding box having the coordinates (x_min, y_min,x_max,y_max) is given by equation 1.
[0065] x_c = (x_min + x_max) / 2, and y_c = (y_min + y_max) / 2 ----- (1)
[0066] Further, the user in proximity of the device is detected based on a Euclidean distance between the centroid and a reference point of the ROI. As an example, the Euclidean distance between the centroid and the reference point of the ROI is given by equation 2.
[0067] Euclidean distance = sqrt((x_c - x_ref)^2+ (y_c-y_ref)^2) --- (2)
[0068] In an embodiment, the multimodal interaction module 403 further compares the Euclidean distance with a threshold proximity value. If the multimodal interaction module 403 determines that the Euclidean distance is lower than the threshold proximity value, then the multimodal interaction module 403 detects that the user is in the proximity of the electronic device 201. In an embodiment, the threshold proximity value is determined based on a desired proximity range. For example, if the distance is below the threshold, consider the user to be in proximity, otherwise not.
[0069] According to an embodiment, the detection of the user in the proximity of the electronic device 201 can be detected by using sensors data. In an embodiment, the multimodal interaction module 403 uses the sensor data obtained from at least one of the infrared sensor (not shown), capacitive touch sensor (not shown), ultrasonic sensor (not shown) or UWB sensor 413-4. The infrared sensor, capacitive touch sensor, ultrasonic sensor may be implemented in the display panel 413. In an embodiment, the multimodal interaction module 403 calculates a first distance of the user from the electronic device 201 based on the reflection or absence of infrared light or radio signals of the infrared sensor or UWB sensor 413-4. Further, the multimodal interaction module 403 detects an intensity of a user's touch on the display panel 413 using the capacitive touch sensor. Further, the multimodal interaction module 403 determines a second distance of the user from the electronic device 201 using the ultrasonic sensor. Further, the multimodal interaction module 403 calculates a multi-sensors proximity value by merging the first distance, the intensity of the user touch, the second distance by assigning a corresponding weight to the first distance, the intensity of the user touch, the second distance. For example, the merging may use such as logical and, logical or, or any and / or combination. The assignment of the corresponding weight to the first distance, the intensity of the user touch, the second distance is performed by using a weighted fusion approach. For example, the weighted fusion may be used to incorporate the continuous distance information. For example, the weighted fusion may assign the weights w1, w2, and w3 to each sensor output (e.g., the first distance, the intensity of the user touch, and the second distance). Further, the multimodal interaction module 403 compares the multi-sensors proximity value with the threshold proximity value. If the multimodal interaction module 403 determines that the multi-sensors proximity value is lesser than the threshold proximity value, then the multimodal interaction module 403 detects the user in the proximity of the electronic device 201.
[0070] In an embodiment, the multimodal interaction module 403 uses the information from the user interaction module 413-1 to determine the user activity. Based on the user activity, the multimodal interaction module 403 detects the user's engagement with the display panel 413 of the electronic device 201 based on the user activity. The operation associated with the detection of the user proximity at block 505 explained in the above paragraphs corresponds to the step 601 of Figure 6.
[0071] Referring back to Figure 5, at block 507 of the block 503, the multimodal interaction module 403, at block 507, determines information associated with the user including at least one of a height of the user or a distance of the user from the display unit based on the detection of the at least one of the proximity of the user and the user's engagement. In an embodiment, the determination of the height of the user or the distance of the user from the display panel 413 at step 503 will be explained below.
[0072] According to an embodiment, the multimodal interaction module 403identifies two reference points on the image of the user's body that are visible within the detected bounding box, such as a top of the head and a bottom of the feet. Further, the multimodal interaction module 403extracts a pixel coordinates (e,g., x_head, y_head) and (e.g., x_feet, y_feet) of these reference points from the image. This can be done by accessing a relevant region within the bounding box or by applying object landmark detection algorithms to locate the specific point accurately. For example, the height may be calculated by using the known dimensions of the reference points, such as the average head-to-feet length or a reference height value. For example, the formula for height calculation may be height = (P*H_ref) / (p_head-p_feet) where P represents a physical distance between camera and display, H_ref represents known height of the reference points, p_head and p_feet represent the pixel coordinates of the head and feet reference points. The calculated height may be adjusted using calling factors or calibration parameters.
[0073] For example, the distance may be calculated by selecting a reference point on the user's body that may be visible within the detected bounding box. For example, a point that provide a reliable indication of the user's position may be selected (e.g., a center of the chest, a midpoint between eyes). The pixel coordinates (x_point, y_point) of the selected reference point may be extracted. This may be done by at least one of accessing the relevant region within the bonding box or applying object landmark detection algorithms to locate the specific point accurately.
[0074] Further, the multimodal interaction module 403may obtain the depth information corresponding to the pixel coordinates of the reference point. This can be obtained from a depth sensor, such as a Time-of-Flight camera or a structured light sensor or estimated using stereo vision techniques or other depth estimation algorithms. The depth sensor provides a depth value or a disparity map that represents the distance of each pixel from the sensor. Further, the multimodal interaction module 403 converts the depth value or the disparity map obtained from the depth sensor into real-world distance units. This requires knowledge of the camera's intrinsic parameters such as a focal length and principal point, obtained during the camera calibration. The multimodal interaction module 403 applies the camera calibration matrix to convert the pixel coordinates of the reference point to the corresponding 3D point in the camera coordinate system. The conversion formula is given by equation 3.
[0075] X = (x_point - c_x) * Z / f_x, Y = (y_point - c_y) * Z / f_y , Z = depth_value --- (3)
[0076] Where (X, Y, Z) represents the 3D coordinate in the camera coordinate system, (x_point, y_point) are the pixel coordinates of the reference point, depth_value is the depth value or disparity obtained from the depth sensor, and (f_x, f_y) and (c_x, c_y) are the focal length and principal point, respectively.
[0077] According to an embodiment, the multimodal interaction module 403 transforms the 3D point from the camera coordinate system to the world coordinate system. This transformation may involve applying camera extrinsic parameters, such as rotation and translation, to align the coordinate systems. The resulting 3D point represents the position of the reference point in the world coordinate system. Once the reference point's position in the world coordinate system is obtained, the multimodal interaction module 403 calculates the distance between the reference point and the camera or the position of the display device. This distance can be calculated using the Euclidean distance formula or any other appropriate distance metric.
[0078] Accordingly, by using the position of the reference points, such as the average head-to-feet length or a reference height value, the multimodal interaction module 403calculates the height of the user in real world units. The equation 4 provides an expression of the height calculation.
[0079] H = (Z / f)*h ---- (4)
[0080] Where, H = Height of user , Z = measured distance of user, f = focal length of the camera, h = projected height of person on the screen of the display panel, O = optical center of the camera.
[0081] The information associated with the user (e.g., the height, the distance) may change dynamically based on change of positions of the user.
[0082] According to an embodiment, the multimodal interaction module 403adjusts the calculated height using calibration parameters. The operation at block 507 corresponds to the step 603 of Figure 6.
[0083] In an embodiment, the information associated with the user e.g. the height of the user or the distance of the user from the display panel 413 is provided to the display zone division module 405. Referring back to Figure 4, at block 509, the display zone division module 405 calculates a number of display zones for dividing the display unit into the number of display zones to align the PIP window over at least one of the number of display zones based on the information associated with the user e.g. the height of the user or the distance of the user from the display panel 413.
[0084] According to an embodiment of the disclosure, the display zone division module 405 divides, based on the information associated with the user (e.g., the height of the user or the distance of the user from the display panel 413), the zone of the display unit into the plurality of display zones. According to an embodiment, the display zone division module 405, at first, calculates an angular field of view (FOV) of the user with respect to the display panel 413 in a horizontal direction (e.g., FOV_h) and a vertical direction (FOV_v). Figure 7 illustrates an example of dimensions and viewing distance (D) of the display panel, according to an embodiment of the present disclosure. In an embodiment, the display zone division module 405 calculates the angular field of view in horizontal direction (e.g. FOV_h) and the angular field of view in vertical direction (e.g. FOV_v) the based on the viewing distance (D) and the dimensions of the display. Further, the FOV_h and the FOV_v is given by equation 5.
[0085] FOV_h = 2 * arctan ((display_width / 2) / D)
[0086] FOV_v = 2 * arctan ((display_height / 2) / D)
[0087] ----- (5)
[0088] wherein, display_width and display_height are the dimensions of the display panel 413.
[0089] Further, based on the angular field of view, the display zone division module 405 determines an intermediate position of the PIP window with respect to the display panel 413. The intermediate position of the PIP window includes the horizontal position of the PIP (PIP_x) and the vertical position of the PIP (PIP_y) and given by equation 6.
[0090] PIP_x = (display_width - W) / 2
[0091] PIP_y = (display_height - H_PIP) / 2
[0092] ------ (6)
[0093] Further, the display zone division module 405, divides the display into different display zones based on the user's height (H), viewing distance (D), and the intermediate position of the PIP window. Accordingly, the display zone division module 405, at block 513, divides the display zone optimally. The display zone division module 405 further calculates the height and the width of each display zone to create proportional and visually accessible regions within the display zone. In particular, the display zone division module 405 divides the display panel 413 vertically into three zones e.g. a top zone, a PIP zone, and a bottom zone. The height of the top zone (H_top) is calculated with equation 7, the height of the bottom zone (H_bottom) is calculated with equation 8, the height of the PIP zone (H_PIP_zone) is equal to H_PIP and the width of each zone is equal to the display_width.
[0094] H_top = (H - H_PIP) / 2 --- (7)
[0095] H_bottom = (H - H_PIP) / 2 --- (8)
[0096] Further, the display zone division module 405 computes a dimension of each of the number of display zones based on the intermediate position of the PIP window and the information associated with the user. Furthermore, the display zone division module 405 calculates, at block 513, the number of display zones based on the dimension of each of the number of display zones. The method of calculating the number of display zones at block 513 and the display zone division module 405 corresponds to the step 605 of Figure 6.
[0097] According to an embodiment, incorporating a user state and adapting the user state at block 511 can further enhance the display zone division. In an embodiment, the display zone division module 405 continuously monitors at least one of the user's head movements, gaze direction, eye position or the user's line of sight to detect changes in the angle of view. Accordingly, the display zone division module 405 calculates the angular field of the view with respect to the display panel 413 and adjusts the size and position of the display zones to maintain an optimal visibility and usability.
[0098] According to an embodiment, considering a change in the user's height and the posture the display zone division module 405 further optimizes the division of the display zones. According to an embodiment, the display zone division module 405 utilizes sensors, such as the depth sensors or camera-based body tracking, to monitor the user's height and posture. The display zone division module 405 continuously tracks changes in the user's height or posture, such as standing up, sitting down, or leaning forward / backward. Accordingly, the display zone division module 405 detects changes in the user's height or posture, the display zone division module 405 considers adjusting the vertical alignment of the display zones. The display zone division module 405 may adapt display zones to user's state. Further, the display zone division module 405 detects changes in the user's height or posture, it considers adjusting the vertical alignment of the display zones. For example, if the user sits down, the display zones can be shifted upwards to align with the user's new eye level.
[0099] According to an embodiment, the display zone division module 405 calculates new dimensions of the display zones based on the updated angle of view, user's height, and desired proportion of content in each zone. Thus, taking into account the user's state, such as height and posture, ensures that the display zones are appropriately positioned for optimal visibility and usability. Thus, implementing smooth transitions and animations when adjusting the display zones provides a seamless and visually pleasing user experience. Further, by gradually resizing and repositioning the zones, abrupt changes that may disrupt the user's focus or cause discomfort are avoided. Further, gathering user feedback on the zone adjustment and adaption of the user state helps in providing effectiveness and usability for future actions.
[0100] Referring back to Figure 4, at block 519 of the block 515, the gaze detection controller module 407 is implemented with a model to track user gaze and AOV of the user. According to an embodiment, the model implements a coordinate framework for eye-tracking data, representing gaze coordinates as (x_eye, y_eye) within the eye-tracking coordinate system. Additionally, the model defines a coordinate system for the display screen, representing the display coordinates as (x_display, y_display). Further, the model undertakes a calibration process to map the eye-tracking coordinate system to the coordinate of the display panel 413. This involves collecting calibration points where the user looks at specific locations on the display panel 413, recording both the eye-tracking coordinates and the display coordinates for each calibration point. The model utilizes this data to compute a mapping function that transforms eye-tracking coordinates to display coordinates, possibly through a mathematical model such as polynomial regression or a transformation matrix. Further, the model continuously tracks the user's eye movements using the eye-tracking system and uses the mapping function to transform the eye-tracking coordinates to display coordinates, enabling computation of the display coordinates corresponding to the current gaze point. The model further incorporates the geometric properties of the display panel 413, such as size, position, and orientation, to calculate the angle of view based on the display geometry and the transformed gaze point coordinates. This calculation involves determining the angle between the user's line of sight to the display panel 413 and the display surface of the display panel 413. The model further offers visual feedback to the user by indicating the detected gaze point on the display, and optionally visualizing the calculated angle of view to help the user understand their viewing perspective.
[0101] Referring back to Figure 6, at block 519 of the block 515, the gaze detection controller module 407, determines a user's line of sight with respect to the number of display zones of the display panel 413 using the one or more image sensors. According to an embodiment, the gaze detection controller module 407 for determining the user's line of sight, at first identifies a co-ordinate on the display panel based on the user's line of sight using the model. The co-ordinate on the display panel 413 represents an interconnection point between the user's line of sight and the display panel. Further, the gaze detection controller module 407 continuously calibrates the coordinate of the display unit based on the user's line of sight using either the polynomial regression or the transformation matrix implemented in the model. Further, the gaze detection controller module 407 dynamically determines a change of the user's line of sight to identify the co-ordinate of the display panel 413 upon continuously calibrating the co-ordinate of the display panel 413 using the model. Further, the gaze detection controller module 407 displays a visual point on the display panel 413 indicating the co-ordinate of the display panel upon considering the display geometry of the display panel 413 and an angle of display with respect to the line of sight of the user. The method of determining the user gaze and angle of view at block 519 performed by the gaze detection controller module 407 corresponds to step 607 of Figure 6.
[0102] Further, referring back to Figure 4, at block 517 of the block 515, the reinforcement learning agent 409 is implemented with the reinforcement learning model to perform the reinforcement learning. According to an embodiment, the reinforcement learning agent 409 aligns the PIP window based on the user's habit. The reinforcement learning agent 409 is formulated with the problem as a Markov Decision Process (MDP), where the agent interacts with the environment to learn an optimal policy.
[0103] According to an embodiment, the reinforcement learning agent 409 defines the state space to capture pertinent information for Picture-in-Picture (PIP) alignment, including features like the PIP's current position (x, y), size, displayed content, and contextual data. Optionally, the reinforcement learning agent 409 includes additional features such as the user's gaze position or historical alignment data. Similarly, the reinforcement learning agent 409 defines the action space to encompass potential alignment adjustments the agent can execute, like horizontal or vertical shifting, resizing, rotation, or aspect ratio alterations. In an embodiment, the reinforcement learning agent 409 is designed with a reward function that guides the agent to align the PIP window according to the user's habits, encouraging alignment that matches the user's preferred position or alignment criteria. This involves defining a reward signal that quantifies alignment quality, potentially by minimizing the distance between the PIP window and the user's preferred alignment position.
[0104] According to an embodiment, the reinforcement learning agent 409 collects data on the user's past PIP alignment behavior and preferences to identify patterns or habits. Incorporate this information into the reward function or as prior knowledge for the reinforcement learning algorithm, leveraging statistical models or clustering techniques to pinpoint common alignment patterns or preferred positions.
[0105] According to an embodiment, the reinforcement learning agent 409 selects a suitable reinforcement learning algorithm, such as Q-learning, Deep Q-Networks (DQN), or Proximal Policy Optimization (PPO), to handle the defined state representation, action space, and reward function. Train the agent using the defined parameters and employ exploration-exploitation strategies like ε-greedy or Boltzmann exploration to balance the exploration of new alignment possibilities and exploitation of learned knowledge. The reinforcement learning agent 409 further periodically updates the user habit model based on the agent's interactions and user feedback and continuously trains and refines the reinforcement learning model using the updated user habit model to adapt to the user's evolving habits or preferences.
[0106] The reinforcement learning agent 409 further collects the user feedback on the PIP alignment and evaluates user satisfaction. The reinforcement learning agent 409 further assesses performance by assessing alignment accuracy, user preferences, and other relevant metrics such as time spent adjusting the PIP manually.
[0107] According to an embodiment, at block 517, the reinforcement learning agent 409 determines, using a reinforcement learning network (e.g. RL model) based on at least one of the user's line of sight and one or more behavioural parameters, an first display zone among the number of display zones. The behavioral parameters may include at least one of the user action, contextual information relating to content preferred by the user, and timestamp for determining the first display zone. According to an embodiment, the reinforcement learning agent 409 performs training of the reinforcement learning network. For doing so the reinforcement learning agent 409, at first, identifies one or more state space representations and one or more action spaces relating to the PIP window alignment in the display panel 413. As an example, the one or more state space representations correspond to at least one of different positions of the PIP window with respect to the display unit, a dimension of the PIP window, a type of content being displayed in the PIP window, and contextual information relating to the content being displayed. Further, the one or more state space representations may correspond to additional features such as the user's gaze position or historical alignment data. Further, the one or more action spaces correspond to at least one of a type of action performed by the user on the PIP window, the type of action relates to at least one of manually moving the PIP window over the display unit, resizing, rotating, and changing an aspect ratio of the PIP window on the display unit. Further, the reinforcement learning agent 409 identifies a reward signal to reduce differences in aligning the PIP window on the display unit with respect to the user's preferred position to view the PIP window. Further, the reinforcement learning agent 409 determines patterns and preferences of the user based on the PIP window alignment on the display unit based on the one or more state space representations, and the one or more action spaces. Furthermore, the reinforcement learning agent 409 determines the reinforcement learning network among one or more reinforcement learning networks to train the reinforcement learning network based on at least one of the one or more state space representations, the one or more action spaces, the reward signal, and the patterns and preferences of the user. Thereafter, the reinforcement learning agent 409 iteratively trains the determined reinforcement learning network to adapt to the user's preferences and determines the first display zone of the PIP window based on the trained reinforcement learning network. The operation performed at block 517 corresponds to the step 609 of Figure 6.
[0108] Referring back to Figure 5, at block 521, the PIP processing unit 411 aligns the PIP window over the first display zone for displaying the PIP window on the first display zone of the display panel 413. According to an embodiment, aligning the PIP window over the first display zone, the PIP processing unit 411 determines a dimension of the PIP window based on the information associated with the user. The dimension of the PIP window is given by w_nom, h_nom (width x height). Now for a particular viewing distance (Z), the size of the PIP window is recalculated and given by equation 9.
[0109] w = min((w_nom*Z) / d_nom, display_width)
[0110] h = min((h_nom*Z) / d_nom, (h_nom*display_width) / w_nom)
[0111] ----- (9)
[0112] The min. operator is needed to ensure that the projected PIP window is fully contained within the display panel 413. Further, the gaze detection controller module 407 estimates the position of the user's eyes (X,Y) in the camera coordinate system as explained above, these coordinates can be mapped to projected screen coordinates using equation 10.
[0113] x = (X*f) / Z
[0114] y = (Y*f) / Z
[0115] ------- (10)
[0116] The (x,y) is an ideal placement of the PIP center for the best viewing experience of the user. Further, the final placement depends on the actual bound of the display panel 413 (display_width, display_height). The PIP processing unit 411 calculates the final placement and is given by equations in Table 1.
[0117]
[0118] [Table 1]
[0119] Further, the PIP processing unit 411 analyzes the user's preferences by using statistical analysis and machine learning techniques. In statistical analysis, the PIP processing unit 411 calculates statistical measures such as the mean, median, mode, variance, and standard deviation of user parameters. Use histograms, scatter plots, or box plots to visualize the distribution of user preferences. Perform correlation analysis to identify relationships between different user parameters. Further, in machine learning the PIP processing unit 411 utilizes machine learning algorithms to predict user preferences based on historical data. The PIP processing unit 411 further uses techniques such as regression, classification, or clustering to build models that capture the relationships between user parameters and PIP alignment. Train the model using labeled data, validate its performance, and fine-tune the model based on user feedback. Additionally, the mathematical model can incorporate reinforcement learning algorithms to dynamically adapt the PIP alignment based on user feedback and optimize the alignment strategy over time.
[0120] According to an embodiment, the PIP processing unit 411 aligns the PIP window over the first display zone based on the user's line of sight, the dimension of the PIP window, and the user's preferences. The aligned the PIP window is then output at the display panel at block 525. The operation at block 521 corresponds to step 611.
[0121] Thus, the disclosed technique provides an efficient PIP placement that enhances the user experience while using the application.
[0122] Figure 8 illustrates an example use case for dynamic PIP window alignment, according to an embodiment of the present disclosure. At block 801, the user clicks on any video. Further, at block 803, as per the conventional method, the video is placed in a fixed position over the screen in the form of PIP window 804. Further, at block 805, as per the disclosed technique, the electronic device takes care of the user's angle of view and behavior of the aligning PIP window 804 on the screen. Thus, at block 807, the PIP window 804 is aligned on the user's preferred space on the screen dynamically based on user preference data and angle of view.
[0123] Figure 9 illustrates an example of a use case for dynamic PIP window alignment, according to an embodiment of the present disclosure. At block 901, the user is watching TV. Further, as per the related method, the video is placed in a fixed position over the screen in the form of PIP window 903. Further, at block 905, according to the disclsure, the electronic device takes care of the user's angle of view and behavior of the aligning PIP window 903 on the screen. Accordingly, the PIP window 903 is aligned on the user's preferred space on the screen dynamically based on user preference data and the angle of view.
[0124] While specific language has been used to describe the disclosure, any limitations arising on account of the same are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein.
[0125] The drawings and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein.
[0126] Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all of the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible. The scope of embodiments is at least as broad as given by the following claims.
[0127] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and any component(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as a critical, required, or essential feature or component of any or all the claims.
[0128] e processor may include various processing circuitry and / or multiple processors. For example, as used herein, including the claims, the term "processor" may include various processing circuitry, including at least one processor, wherein one or more of at least one processor, individually and / or collectively in a distributed manner, may be configured to perform various functions described herein. As used herein, when "a processor", "at least one processor", and "one or more processors" are described as being configured to perform numerous functions, these terms cover situations, for example and without limitation, in which one processor performs some of recited functions and another processor(s) performs other of recited functions, and also situations in which a single processor may perform all recited functions. Additionally, the at least one processor may include a combination of processors performing various of the recited / disclosed functions, e.g., in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.
[0129] According to an embodiment of the disclosure, the reinforcement learning network may be trained based on user's prior viewing preferences to optimize an alignment of the PIP window over the first display zone.
[0130] According to an embodiment of the disclosure, a method performed by the device may include multimodal sensors based on the detection. According to an embodiment of the disclosure, the multimodal sensors may correspond to at least one of the one or more image sensors, one or more depth sensors, one or more microphones, and one or more motion sensors.
[0131] According to an embodiment of the disclosure, the information associated with the user may change dynamically based on change of positions of the user.
[0132] According to an embodiment of the disclosure, a method performed by the device may include capturing, using the one or more image sensors, one or more images of one or more objects in a field of view of the one or more image sensors. According to an embodiment of the disclosure, a method performed by the device may include detecting, using a pre-trained Machine Learning model, the user from the one or more objects in the one or more images. According to an embodiment of the disclosure, a method performed by the device may include defining a bounding box on the detected user in the one or more images. According to an embodiment of the disclosure, a method performed by the device may include estimating the proximity of the user based on a calculation of a centroid of the bounding box. According to an embodiment of the disclosure, a method performed by the device may include detecting, upon a determination that the estimated proximity of the user is lower than a first threshold proximity value, the user in the proximity of the device.
[0133] According to an embodiment of the disclosure, a method performed by the device may include calculating, using an infrared sensor, a first distance of the user from the device. According to an embodiment of the disclosure, a method performed by the device may include detecting, using a capacitive touch sensor, an intensity of a user's touch on the display unit. According to an embodiment of the disclosure, a method performed by the device may include determining, using an ultrasonic sensor, a second distance of the user from the device. According to an embodiment of the disclosure, a method performed by the device may include calculating a multi-sensors proximity value by merging the first distance, the intensity of the user touch, the second distance by assigning a corresponding weight to the first distance, the intensity of the user touch, the second distance. According to an embodiment of the disclosure, a method performed by the device may include detecting, upon a determination that the multi-sensors proximity value is lesser than a first threshold proximity value, the user in the proximity of the device.
[0134] According to an embodiment of the disclosure, a method performed by the device may include calculating an angular field of view of the user with respect to the display unit in a horizontal direction and a vertical direction. According to an embodiment of the disclosure, a method performed by the device may include determining, based on the angular field of view, an intermediate position of the PIP window with respect to the display unit. According to an embodiment of the disclosure, a method performed by the device may include computing a dimension of each of the number of display zones based on the intermediate position of the PIP window and the information associated with the user. According to an embodiment of the disclosure, a method performed by the device may include calculating the number of display zones based on the dimension of each of the number of display zones.
[0135] According to an embodiment of the disclosure, a method performed by the device may include determining a change in the angular field of view based on continuous monitoring of the user's line of sight and the information associated with the user. According to an embodiment of the disclosure, a method performed by the device may include calculating, based on a determination of the change, the angular field of the view with respect to the display unit.
[0136] According to an embodiment of the disclosure, a method performed by the device may include identifying a co-ordinate on the display unit based on the user's line of sight, wherein the co-ordinate on the display unit represents an interconnection point between the user's line of sight and the display unit. According to an embodiment of the disclosure, a method performed by the device may include continuously calibrating, using either a polynomial regression or a transformation matrix, the co-ordinate of the display unit based on the user's line of sight. According to an embodiment of the disclosure, a method performed by the device may include dynamically determining, upon continuously calibrating the co-ordinate of the display unit, a change of the user's line of sight to identify the co-ordinate of the display unit. According to an embodiment of the disclosure, a method performed by the device may include displaying a visual point on the display unit indicating the co-ordinate of the display unit upon considering a display geometry of the display unit and an angle of display with respect to the line of sight of the user.
[0137] According to an embodiment of the disclosure, a method performed by the device may include identifying one or more state space representations and one or more action spaces relating to the PIP window alignment in the display unit. According to an embodiment of the disclosure, a method performed by the device may include identifying a reward signal to reduce differences in aligning the PIP window on the display unit with respect to the user's preferred position to view the PIP window. According to an embodiment of the disclosure, a method performed by the device may include determining patterns and preferences of the user based on the PIP window alignment on the display unit based on the one or more state space representations, and the one or more action spaces. According to an embodiment of the disclosure, a method performed by the device may include determining the reinforcement learning network among one or more reinforcement learning networks to train the reinforcement learning network based on the one or more state space representations, the one or more action spaces, the reward signal, and the patterns and preferences of the user. According to an embodiment of the disclosure, a method performed by the device may include iteratively training the determined reinforcement learning network to adapt to user's preferences. According to an embodiment of the disclosure, a method performed by the device may include determining the first display zone of the PIP window based on the trained reinforcement learning network.
[0138] According to an embodiment of the disclosure, the one or more state space representations may correspond to different positions of the PIP window with respect to the display unit, a dimension of the PIP window, a type of content being displayed in the PIP window, and contextual information relating to the content being displayed.
[0139] According to an embodiment of the disclosure, the one or more action spaces may correspond to a type of action performed by the user on the PIP window, the type of action relates to manually moving the PIP window over the display unit, resizing, rotating, and changing an aspect ratio of the PIP window on the display unit.
[0140] According to an embodiment of the disclosure, a method performed by the device may include determining a dimension of the PIP window based on the information associated with the user. According to an embodiment of the disclosure, a method performed by the device may include aligning the PIP window over the first display zone based on the user's line of sight, the dimension of the PIP window, and the user's preferences.
[0141] According to an embodiment of the disclosure, the using the reinforcement learning network may be based on the one or more behavioural parameters. According to an embodiment of the disclosure, the one or more behavioural parameters may correspond to at least one of a user action, contextual information relating to content preferred by the user, and timestamp for determining the optimal display zone.
[0142] According to an embodiment of the disclosure, the electronic device may include at least one of the multimodal sensors. According to an embodiment of the disclosure, the electronic device may include at least one of the image sensors.
[0143] According to an embodiment of the disclosure, the at least one processor may be configured to train the reinforcement learning network based on user's prior viewing preferences to optimize an alignment of the PIP window over the first display zone.
[0144] According to an embodiment of the disclosure, the device may include multimodal sensors based on the detection. According to an embodiment of the disclosure, the multimodal sensors may correspond to at least one of the one or more image sensors, one or more depth sensors, one or more microphones, and one or more motion sensors.
[0145] According to an embodiment of the disclosure, the information associated with the user may change dynamically based on change of positions of the user.
[0146] According to an embodiment of the disclosure, the at least one processor may be configured to capture, using the one or more image sensors, one or more images of one or more objects in a field of view of the one or more image sensors. According to an embodiment of the disclosure, the at least one processor may be configured to detect, using a pre-trained Machine Learning model, the user from the one or more objects in the one or more images. According to an embodiment of the disclosure, the at least one processor may be configured to define a bounding box on the detected user in the one or more images. According to an embodiment of the disclosure, the at least one processor may be configured to estimate the proximity of the user based on a calculation of a centroid of the bounding box. According to an embodiment of the disclosure, the at least one processor may be configured to detect, upon a determination that the estimated proximity of the user is lower than a first threshold proximity value, the user in the proximity of the device.
[0147] According to an embodiment of the disclosure, the at least one processor may be configured to calculate, using an infrared sensor, a first distance of the user from the device. According to an embodiment of the disclosure, the at least one processor may be configured to detect, using a capacitive touch sensor, an intensity of a user's touch on the display unit. According to an embodiment of the disclosure, the at least one processor may be configured to determine, using an ultrasonic sensor, a second distance of the user from the device. According to an embodiment of the disclosure, the at least one processor may be configured to calculate a multi-sensors proximity value by merging the first distance, the intensity of the user touch, the second distance by assigning a corresponding weight to the first distance, the intensity of the user touch, the second distance. According to an embodiment of the disclosure, the at least one processor may be configured to detect, upon a determination that the multi-sensors proximity value is lesser than a first threshold proximity value, the user in the proximity of the device.
[0148] According to an embodiment of the disclosure, the at least one processor may be configured to calculate an angular field of view of the user with respect to the display unit in a horizontal direction and a vertical direction. According to an embodiment of the disclosure, the at least one processor may be configured to determine, based on the angular field of view, an intermediate position of the PIP window with respect to the display unit. According to an embodiment of the disclosure, the at least one processor may be configured to compute a dimension of each of the number of display zones based on the intermediate position of the PIP window and the information associated with the user. According to an embodiment of the disclosure, the at least one processor may be configured to calculate the number of display zones based on the dimension of each of the number of display zones.
[0149] According to an embodiment of the disclosure, the at least one processor may be configured to determine a change in the angular field of view based on continuous monitoring of the user's line of sight and the information associated with the user. According to an embodiment of the disclosure, the at least one processor may be configured to calculate, based on a determination of the change, the angular field of the view with respect to the display unit.
[0150] According to an embodiment of the disclosure, the at least one processor may be configured to identify a co-ordinate on the display unit based on the user's line of sight, wherein the co-ordinate on the display unit represents an interconnection point between the user's line of sight and the display unit. According to an embodiment of the disclosure, the at least one processor may be configured to continuously calibrate, using either a polynomial regression or a transformation matrix, the co-ordinate of the display unit based on the user's line of sight. According to an embodiment of the disclosure, the at least one processor may be configured to dynamically determine, upon continuously calibrating the co-ordinate of the display unit, a change of the user's line of sight to identify the co-ordinate of the display unit. According to an embodiment of the disclosure, the at least one processor may be configured to display a visual point on the display unit indicating the co-ordinate of the display unit upon considering a display geometry of the display unit and an angle of display with respect to the line of sight of the user.
[0151] According to an embodiment of the disclosure, the at least one processor may be configured to identify one or more state space representations and one or more action spaces relating to the PIP window alignment in the display unit. According to an embodiment of the disclosure, the at least one processor may be configured to identify a reward signal to reduce differences in aligning the PIP window on the display unit with respect to the user's preferred position to view the PIP window. According to an embodiment of the disclosure, the at least one processor may be configured to determine patterns and preferences of the user based on the PIP window alignment on the display unit based on the one or more state space representations, and the one or more action spaces. According to an embodiment of the disclosure, the at least one processor may be configured to determine the reinforcement learning network among one or more reinforcement learning networks to train the reinforcement learning network based on the one or more state space representations, the one or more action spaces, the reward signal, and the patterns and preferences of the user. According to an embodiment of the disclosure, the at least one processor may be configured to iteratively train the determined reinforcement learning network to adapt to user's preferences. According to an embodiment of the disclosure, the at least one processor may be configured to determine the first display zone of the PIP window based on the trained reinforcement learning network.
[0152] According to an embodiment of the disclosure, the one or more state space representations may correspond to different positions of the PIP window with respect to the display unit, a dimension of the PIP window, a type of content being displayed in the PIP window, and contextual information relating to the content being displayed.
[0153] According to an embodiment of the disclosure, the one or more action spaces may correspond to a type of action performed by the user on the PIP window, the type of action relates to manually moving the PIP window over the display unit, resizing, rotating, and changing an aspect ratio of the PIP window on the display unit.
[0154] According to an embodiment of the disclosure, the at least one processor may be configured to determine a dimension of the PIP window based on the information associated with the user. According to an embodiment of the disclosure, the at least one processor may be configured to align the PIP window over the first display zone based on the user's line of sight, the dimension of the PIP window, and the user's preferences.
[0155] According to an embodiment of the disclosure, the using the reinforcement learning network may be based on the one or more behavioural parameters. According to an embodiment of the disclosure, the one or more behavioural parameters may correspond to at least one of a user action, contextual information relating to content preferred by the user, and timestamp for determining the optimal display zone.
[0156] According to an embodiment of the disclosure, a system for dynamically aligning Picture-In-Picture (PIP) window in a display unit of the device is provided. According to an embodiment of the disclosure, the system may include at least one processor and at least one memory storing computer executable instructions. According to an embodiment of the disclosure, at least one processor is configured to detect at least one of a user in a proximity of the device, and a user's engagement with the display unit of the device. According to an embodiment of the disclosure, at least one processor is configured to determine information associated with the user including at least one of a height of the user or a distance of the user from the display unit. According to an embodiment of the disclosure, at least one processor is configured to divide, based on the information associated with the user, a zone of the display unit into the plurality of display zones. According to an embodiment of the disclosure, at least one processor is configured to determine, using one or more image sensors, a user's line of sight with respect to the number of display zones of the display unit. According to an embodiment of the disclosure, at least one processor is configured to determine, using a reinforcement learning network based on the user's line of sight, a first display zone among the plurality of display zones. According to an embodiment of the disclosure, at least one processor is configured to align the PIP window over the determined first display zone.
Claims
1.A method for dynamically aligning Picture-In-Picture (PIP) window in a display unit of a device, the method comprising:detecting at least one of:a user in a proximity of the device, anda user's engagement with the display unit of the device;determining information associated with the user including at least one of a height of the user or a distance of the user from the display unit;dividing, based on the information associated with the user, a zone of the display unit into the plurality of display zones;determining, using one or more image sensors, a user's line of sight with respect to the plurality of display zones of the display unit;determining, using a reinforcement learning network based on the user's line of sight, a first display zone among the plurality of display zones; andaligning the PIP window over the determined first display zone.2.The method as claimed in claim 1, wherein the reinforcement learning network is trained based on user's prior viewing preferences to optimize an alignment of the PIP window over the first display zone.3.The method as claimed in any one of claims 1 and 2, wherein the determining the information comprises using multimodal sensors based on the detection, andwherein the multimodal sensors correspond to at least one of the one or more image sensors, one or more depth sensors, one or more microphones, and one or more motion sensors.4.The method as claimed in any one of claims 1 to 3, wherein the information associated with the user changes dynamically based on change of positions of the user.5.The method as claimed in any one of claims 1 to 4, wherein detecting the user in the proximity of the device comprises:capturing, using the one or more image sensors, one or more images of one or more objects in a field of view of the one or more image sensors;detecting, using a pre-trained Machine Learning model, the user from the one or more objects in the one or more images;defining a bounding box on the detected user in the one or more images;estimating the proximity of the user based on a calculation of a centroid of the bounding box; anddetecting, upon a determination that the estimated proximity of the user is lower than a first threshold proximity value, the user in the proximity of the device.6.The method as claimed in any one of claims 1 to 5, wherein detecting the user in the proximity of the device further comprises:calculating, using an infrared sensor, a first distance of the user from the device;detecting, using a capacitive touch sensor, an intensity of a user's touch on the display unit;determining, using an ultrasonic sensor, a second distance of the user from the device;calculating a multi-sensors proximity value by merging the first distance, the intensity of the user touch, the second distance by assigning a corresponding weight to the first distance, the intensity of the user touch, the second distance; anddetecting, upon a determination that the multi-sensors proximity value is lesser than a first threshold proximity value, the user in the proximity of the device.7.The method as claimed in any one of claims 1 to 6, wherein dividing the zone of the display unit into the plurality of display zones comprises:calculating an angular field of view of the user with respect to the display unit in a horizontal direction and a vertical direction;determining, based on the angular field of view, an intermediate position of the PIP window with respect to the display unit;computing a dimension of each of the number of display zones based on the intermediate position of the PIP window and the information associated with the user; andcalculating the number of display zones based on the dimension of each of the number of display zones.8.The method as claimed in claim 7, wherein calculating the angular field of view with respect to the display unit comprises:determining a change in the angular field of view based on continuous monitoring of the user's line of sight and the information associated with the user; andcalculating, based on a determination of the change, the angular field of the view with respect to the display unit.9.The method as claimed in any one of claims 1 to 8, wherein determining the user's line of sight with respect to the display unit comprises:identifying a co-ordinate on the display unit based on the user's line of sight, wherein the co-ordinate on the display unit represents an interconnection point between the user's line of sight and the display unit;continuously calibrating, using either a polynomial regression or a transformation matrix, the co-ordinate of the display unit based on the user's line of sight;dynamically determining, upon continuously calibrating the co-ordinate of the display unit, a change of the user's line of sight to identify the co-ordinate of the display unit; anddisplaying a visual point on the display unit indicating the co-ordinate of the display unit upon considering a display geometry of the display unit and an angle of display with respect to the line of sight of the user.10.The method as claimed in any one of claims 1 to 9, wherein a training of the reinforcement learning network comprises:identifying one or more state space representations and one or more action spaces relating to the PIP window alignment in the display unit;identifying a reward signal to reduce differences in aligning the PIP window on the display unit with respect to the user's preferred position to view the PIP window;determining patterns and preferences of the user based on the PIP window alignment on the display unit based on the one or more state space representations, and the one or more action spaces;determining the reinforcement learning network among one or more reinforcement learning networks to train the reinforcement learning network based on the one or more state space representations, the one or more action spaces, the reward signal, and the patterns and preferences of the user;iteratively training the determined reinforcement learning network to adapt to user's preferences; anddetermining the first display zone of the PIP window based on the trained reinforcement learning network.11.The method as claimed in claim 10, wherein the one or more state space representations correspond to different positions of the PIP window with respect to the display unit, a dimension of the PIP window, a type of content being displayed in the PIP window, and contextual information relating to the content being displayed, andwherein the one or more action spaces correspond to a type of action performed by the user on the PIP window, the type of action relates to manually moving the PIP window over the display unit, resizing, rotating, and changing an aspect ratio of the PIP window on the display unit.12.The method as claimed in claim 11, wherein aligning the PIP window over the first display zone comprises:determining a dimension of the PIP window based on the information associated with the user; andaligning the PIP window over the first display zone based on the user's line of sight, the dimension of the PIP window, and the user's preferences.13.An electronic device for dynamically aligning Picture-In-Picture (PIP) window in a display unit of the device, the electronic device comprising:at least one processor; andat least one memory storing computer executable instructions that, when executed by the at least one processor, cause the at least one processor configured to:detect at least one of:a user in a proximity of the device, anda user's engagement with the display unit of the device;determine information associated with the user including at least one of a height of the user or a distance of the user from the display unit;divide, based on the information associated with the user, a zone of the display unit into the plurality of display zones;determine, using one or more image sensors, a user's line of sight with respect to the plurality of display zones of the display unit;determine, using a reinforcement learning network based on the user's line of sight, a first display zone among the plurality of display zones; andalign the PIP window over the determined first display zone.14.The electronic device as claimed in claim 13, wherein the at least one processor is further configured to:capture, using the one or more image sensors, one or more images of one or more objects in a field of view of the one or more image sensors;detect, using a pre-trained Machine Learning model, the user from the one or more objects in the one or more images;define a bounding box on the detected user in the one or more images;estimate the proximity of the user based on a calculation of a centroid of the bounding box; anddetect, upon a determination that the estimated proximity of the user is lower than a first threshold proximity value, the user in the proximity of the device.15.The electronic device as claimed in any one of claims 13 and 14, wherein the at least one processor is further configured to:identify one or more state space representations and one or more action spaces relating to the PIP window alignment in the display unit;identify a reward signal to reduce differences in aligning the PIP window on the display unit with respect to the user's preferred position to view the PIP window;determine patterns and preferences of the user based on the PIP window alignment on the display unit based on the one or more state space representations, and the one or more action spaces;determine the reinforcement learning network among one or more reinforcement learning networks to train the reinforcement learning network based on the one or more state space representations, the one or more action spaces, the reward signal, and the patterns and preferences of the user;iteratively train the determined reinforcement learning network to adapt to user's preferences; anddetermine the first display zone of the PIP window based on the trained reinforcement learning network.
Citation Information
Patent Citations
Mobile terminal using eye tracking function and event indication method thereof
KR1020140002389A
Window placement based on user location
US11429263B1
User interface device, user interface method, and recording medium
US20100269072A1
Method and apparatus for display control using eye tracking
US20200004333A1
Children face distance alert system
US20200151432A1