Display of sign language videos through adjustable user interface (UI) elements

KR102999311B1Active Publication Date: 2026-08-03SONY KK (TAMBIEN COMERCIANDO COMO SONY CORP)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
KR1020237044249
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-04
Filing Date
2022-10-21
Publication Date
2026-08-03
Estimated Expiration
2042-10-21

Smart Images

  • Figure R1020237044249_ABST
    Figure R1020237044249_ABST
Patent Text Reader

Abstract

An electronic device and a method for displaying sign language video through adjustable user interface (UI) elements are provided. The electronic device receives a first media stream containing video. The electronic device determines the position of a receiver within the video. The electronic device extracts a portion of the video corresponding to the determined position of the video. The electronic device controls the playback of the video on a display device. The electronic device controls the display device based on the playback to render UI elements on the display device and to display the extracted portion of the video within the UI elements.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Cross-reference regarding related applications

[0002] This application claims priority to U.S. Patent Application No. 17 / 453553, filed with the U.S. Patent and Trademark Office on November 4, 2021. Each of the above-mentioned applications is hereby incorporated herein by reference in its entirety.

[0003] Technology field

[0004] Various embodiments of the present disclosure relate to the display of a signing video. More specifically, various embodiments of the present disclosure relate to an electronic device and a method for displaying a signing video through adjustable UI elements. Background Technology

[0005] Traditionally, display devices (such as televisions or mobile phones) receive broadcast or streaming media content that includes video files and audio files synchronized with the video files. Both the video and audio files are rendered simultaneously on the display device for viewing. In some cases, the media content (e.g., video) also includes a sign language speaker or interpreter using sign language (e.g., American Sign Language, ASL) to provide a service to deaf viewers. However, the video portion of the sign language speaker (sign language video) typically appears in the bottom corner of the video. The sign language video may be too small for comfortable viewing or may interfere with areas of the main video (such as important announcements in the main video). Existing systems do not provide simple user interface technologies to help deaf viewers conveniently view media content and sign language videos.

[0006] Further limitations and disadvantages of conventional and traditional approaches will become apparent to those skilled in the art through a comparison of the described systems and some aspects of the disclosure, as presented with reference to the remainder of this application and the drawings.

[0007] An electronic device and method for displaying sign language video through adjustable UI elements are substantially provided as illustrated in at least one of the drawings and / or described in connection with the drawings, as more fully described in the claims.

[0008] These and other features and advantages of the present disclosure can be understood from a review of the following detailed description of the present disclosure, together with the accompanying drawings in which similar reference numbers refer to similar parts throughout. Brief explanation of the drawing

[0009] FIG. 1 is a drawing illustrating an exemplary network environment for displaying sign language video through adjustable user interface (UI) elements according to an embodiment of the present disclosure. FIG. 2 is a block diagram of an exemplary electronic device for displaying sign language video through adjustable user interface (UI) elements according to an embodiment of the present disclosure. FIG. 3 is a drawing illustrating an exemplary scenario for displaying sign language video through user interface (UI) elements adjustable based on metadata, according to an embodiment of the present disclosure. FIG. 4 is a diagram illustrating an exemplary scenario for displaying sign language video through adjustable user interface (UI) elements for live video broadcasting, according to an embodiment of the present disclosure. FIG. 5 is a drawing illustrating an exemplary scenario for adjusting the position of a user interface (UI) element displaying a sign language video according to an embodiment of the present disclosure. FIG. 6 is a drawing illustrating an exemplary scenario for adjusting the size of a user interface (UI) element displaying a sign language video according to an embodiment of the present disclosure. FIG. 7 is a flowchart illustrating exemplary operations for displaying sign language video through adjustable user interface (UI) elements according to an embodiment of the present disclosure. Specific details for implementing the invention

[0010] The embodiments described below may be found in the disclosed electronic device and method for displaying sign language video (sign language video) through adjustable user interface (UI) elements. Exemplary aspects of the present disclosure provide an electronic device (e.g., a smart television or mobile device) that can be coupled to a display device. The electronic device may receive a media stream that may include video. The electronic device may determine the position of a sign language user within the video. The sign language user may be an animated character or a person who can perform using sign language in the video. The electronic device may further extract from the video a portion of the video corresponding to the determined position of the sign language user within the video. The electronic device may control the playback of the video on the display device. Based on the playback, the electronic device may control the display device to render user interface (UI) elements on the display device and to display the extracted portion of the video within the UI elements. The UI elements may be rendered as picture-in-picture (PiP) windows of adjustable size. Therefore, the electronic device can provide adjustable UI elements to conveniently view the recipient's video along with the main video.

[0011] In one embodiment, the electronic device may receive metadata associated with the video. The metadata may include information describing the location of the listener within the video at a plurality of timestamps. The electronic device may determine the location of the listener within the video based on the received metadata. In another embodiment, the electronic device may detect a region in the video based on the difference between the background of the region and the background of the rest of the video. The electronic device may detect the location of the listener within the video based on the detection of the region. In another embodiment, the electronic device may detect a boundary around the listener in the video. The electronic device may detect the location of the listener within the video based on the detection of the boundary. In some embodiments, the electronic device may be configured to detect hand signals associated with sign language within the video based on the application of a neural network model to one or more frames of the video (e.g., a live video broadcast). The electronic device may be configured to detect the location of the listener within the video based on the detection of the hand signals. The electronic device may also extract the video portion of the listener and control the display device to render UI elements (e.g., a PiP window) on the display device based on the detected location of the listener. Therefore, the electronic device can automatically detect the location of the receiver for live video broadcasting and create a PiP window for the receiver.

[0012] In an embodiment, the electronic device may provide the ability to customize UI elements (e.g., a PiP window) according to user preferences. The electronic device may be configured to adjust the size of the UI elements, the position of the UI elements, the theme or color scheme for the UI elements, the preference for hiding UI elements, and the schedule for rendering the UI elements based on user preferences. For example, the electronic device may receive a first input to change the current position of the PiP window to a first position different from the current position. The electronic device may control a display device to render the PiP window at the first position based on the first user input. In another example, the electronic device may receive a second input to change the current size of the PiP window to a first size different from the current size. The electronic device may control a display device to change the current size of the PiP window to match the first size based on the second input. Thus, the electronic device may provide a simple and easy-to-use UI technology for adjusting the size and position of the video portion of the listener based on the PiP window. Based on the adjustment of the position and size of the PiP window, the electronic device can provide a clear and enlarged screen of the video to the receiver, and can enable the main video (such as important announcements in the main video) to be viewed without obstruction.

[0013] FIG. 1 is a drawing illustrating an exemplary network environment for displaying sign language video through adjustable user interface (UI) elements according to an embodiment of the present disclosure. Referring to FIG. 1, a drawing of a network environment (100) is shown. In the network environment (100), an electronic device (102), a display device (104), and a server (106) are shown. The electronic device (102) may be communicably coupled to the server (106) via a communication network (108). The electronic device (102) may be communicably coupled to the display device (104) either directly or via the communication network (108). Referring to FIG. 1, a UI element (112) for displaying a main video (104A) and a video of a receiver (110) is additionally shown.

[0014] In FIG. 1, the electronic device (102) and the display device (104) are shown as two separate devices, but in some embodiments, the display device (104) may be integrated with the electronic device (102) without departing from the scope of the present disclosure.

[0015] The electronic device (102) may include appropriate logic, circuits, and interfaces that can be configured to render user interface (UI) elements (112) on a display device (104). The electronic device (102) may receive a first media stream containing video. The electronic device (102) may further receive metadata associated with the video. The electronic device (102) may further extract a portion of the video associated with the location of the receiver (110) within the video. The electronic device (102) may control the display device (104) to display the extracted portion of the video within the UI elements (112). The electronic device (102) may include appropriate middleware and codecs that can enable the reception and / or playback of media content (such as video). In one embodiment, the electronic device (102) may be associated with a plurality of user profiles. Each of the plurality of user profiles may include a collection of content items, settings or menu options, user preferences, etc. The electronic device (102) may allow the selection and / or switching of a user profile among a plurality of user profiles on a graphical user interface based on user input from a remote control or a touch screen interface. The electronic device (102) may include an infrared receiver or a Bluetooth® interface for receiving control signals transmitted from a remote control corresponding to a button pressed on the remote control.Examples of electronic devices (102) may include, but are not limited to, smart televisions (TVs), internet-protocol TVs (IPTVs), digital media players, micro-consoles, set-top boxes, over-the-top (OTT) players, streaming players, media extenders / regulators, digital media hubs, smartphones, personal computers, laptops, tablets, wearable electronic devices, head-mounted devices, or any other display device having the function of receiving, decoding, and playing content from broadcast signals, streaming content broadcasts, internet-based communication signals, etc., via cable or satellite networks. Examples of content may include, but are not limited to, images, animations (e.g., 2D / 3D animations or motion graphics), audio / video data, conventional television programs (provided via traditional broadcasting, cable, satellite, the Internet, or other means), pay programs, on-demand programs (such as in video-on-demand (VOD) systems), or Internet content (e.g., streaming media, downloadable media, webcasts, etc.).

[0016] In one embodiment, the electronic device (102) may be configured to generate UI elements (112) (e.g., picture-in-picture (PiP) windows) based on a received first media stream. The PiP windows may be resizable and repositionable based on user input. The PiP windows may display a portion of the main video (104A) including a sign language speaker (110) performing sign language in the main video. The ability to generate UI elements (112) (e.g., PiP windows) may be integrated with the electronic device (102) by the manufacturer of the electronic device (102), or may be available for download as an add-on application from a server (106) or an application store / marketplace.

[0017] The display device (104) may include appropriate logic, circuits, and interfaces that can be configured to render UI elements (112) that display the extracted video portion of the receiver. The display device (104) may be further configured to display the main video (104A) being played by the electronic device (102). In one embodiment, the display device (104) may be an external display device connected to the electronic device (102). For example, the display device (104) may be connected to the electronic device (102) (e.g., a digital media player or a personal video recorder) by a wired connection (e.g., a high-definition multimedia interface (HDMI) connection) or a wireless connection (e.g., Wi-Fi). In another embodiment, the display device (104) may be integrated with the electronic device (102) (such as a smart television). A display device (104), such as a display screen with integrated audio speakers, may include one or more controllable parameters such as brightness, contrast, aspect ratio, saturation, audio volume, etc. An electronic device (102) may be configured to control the parameters of the display device (104) by transmitting one or more signals via a wired connection (such as an HDMI connection). In one embodiment, the display device (104) may be a touch screen capable of receiving user input via touch input. The display device (104) may be realized through several known technologies, such as a Liquid Crystal Display (LCD), Light Emitting Diode (LED) display, plasma display, or Organic LED (OLED) display technology, or at least one of other display devices, but is not limited thereto.In at least one embodiment, the display device (104) may be a display unit of a smart TV, a head-mounted device (HMD), a smart glass device, a see-through display, a heads-up display (HUD), an in-vehicle infotainment system, a projection-based display, an electrochromic display, or a transparent display.

[0018] The server (106) may include suitable logic, circuits, and interfaces and / or code that can be configured to store one or more media streams. The server (106) may be further configured to store metadata for determining the position of a sign language speaker in one or more videos. In some embodiments, the server (106) may be configured to train a neural network model for detecting sign signals associated with sign language in a video. In some embodiments, the server may be configured to store a neural network model and a training dataset for training the neural network model. The server (106) may be further configured to store user profiles associated with the electronic device (102), preferences associated with UI elements (112) for each user profile, usage history of UI elements (112) for each user profile, etc. The server (106) may be implemented as a cloud server and may execute operations through web applications, cloud applications, HTTP requests, repository operations, file transfers, and similar things. Other exemplary implementations of the server (106) may include, but are not limited to, a database server, a file server, a web server, a media server, an application server, a mainframe server, or a cloud computing server. In at least one embodiment, the server (106) may be implemented as a plurality of distributed cloud-based resources using various techniques widely known to those skilled in the art. Those skilled in the art will understand that the scope of this disclosure may not be limited to implementing the server (106) and the electronic device (102) as two separate entities. In certain embodiments, the functions of the server (106) may be integrated wholly or at least partially within the electronic device (102) without departing from the scope of this disclosure.

[0019] A communication network (108) may include a communication medium through which an electronic device (102), a display device (104), and a server (106) can communicate with each other. The communication network (108) may be either a wired connection or a wireless connection, and examples of the communication network (108) may include, but are not limited to, the Internet, a cloud network, a cellular or wireless mobile network (e.g., Long-Term Evolution and 5G New Radio), a Wi-Fi (Wireless Fidelity) network, a Personal Area Network (PAN), a Local Area Network (LAN), or a Metropolitan Area Network (MAN). Various devices within a network environment (100) may be configured to connect to the communication network (108) according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of the Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Zigbee, EDGE, IEEE 802.11, light fidelity (Li-Fi), 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.

[0020] When operating, the electronic device (102) may receive a first media stream that may include video. The first media stream may be encoded in a standard digital container format for the transmission of video. In one embodiment, the electronic device (102) may receive the first media stream from a server (106). In one embodiment, the electronic device (102) may be further configured to receive metadata associated with the video. The metadata may include information describing the location of the signer (110) within the video at a plurality of timestamps. The signer (110) may be an animated character or person capable of acting using sign language in the video. For example, the signer (110) may be a person capable of interpreting spoken language within the video and converting the spoken language into sign language.

[0021] In one embodiment, the electronic device (102) may be further configured to determine the location of the receiver (110) within the video based on the reception of metadata. The determined location may include image coordinates of the corners of a rectangular area of ​​the video containing the receiver (110). In another embodiment, the electronic device (102) may be configured to apply a neural network model (shown in FIG. 2) to the video to determine the location of the receiver (110) within the video. In another embodiment, the electronic device (102) may determine the location of the receiver (110) in the video by image analysis based on the difference between the main video (104A) and the background of the area around the receiver (110), or based on the detection of boundaries around the receiver (110) in the video. Details of the application of the neural network model and image analysis are provided, for example, in FIG. 2 and FIG. 3b. In one embodiment, the electronic device (102) may be further configured to control the display device (104) to render a highlighted boundary around the receiver (110) in the displayed video. The boundary may be rendered based on a determined location. As an example, if the electronic device (102) identifies a number of candidates for a video portion of the receiver (110) from the video, the electronic device (102) may display the boundary to obtain user confirmation of the video portion of the receiver within the video.

[0022] The electronic device (102) may be further configured to extract a portion of the video corresponding to a determined location of the receiver (110) within the video. The portion of the video may be extracted from a rectangular area of ​​the video. In one embodiment, the electronic device (102) may receive a second media stream containing the extracted portion of the video. The second media stream may be different from the first media stream. The electronic device (102) may be further configured to control the playback of the video (e.g., main video (104A)) on the display device (104). The electronic device (102) may be further configured to control the display device (104) to render a UI element (112) on the display device (104) and to display the extracted portion of the video inside the UI element (112). The UI element (112) may be rendered as a picture-in-picture (PiP) window of adjustable size and position.

[0023] In one embodiment, the electronic device (102) may be configured to customize a UI element (112) (PiP window) according to user preferences. For example, the electronic device (102) may be configured to adjust, based on user preferences, the size of the UI element (112), the position of the UI element (112), the theme or color scheme for the UI element (112), the preference for hiding the UI element (112), and the schedule for rendering the UI element (112). For example, the electronic device (102) may receive a first input to change the current position of the PiP window to a first position different from the current position. Based on the first user input, the electronic device (102) may control the display device (104) to render the PiP window at the first position. In another example, the electronic device (102) may receive a second input to change the current size of the PiP window to a first size different from the current size. The electronic device (102) can control the display device to change the current size of the PiP window to match the first size based on the second input. The electronic device (102) can thereby provide a simple and easy-to-use UI technology to adjust the size and position of the video portion of the receiver based on UI elements (112) (such as the PiP window). Based on the adjustment of the position and size of the UI elements (112), the electronic device (102) can provide a clear and enlarged screen of the video of the receiver and allow the main video (104A) (such as important announcements in the main video (104A)) to be viewed without obstruction.

[0024] Modifications, additions, or omissions may be made to FIG. 1 without departing from the scope of the present disclosure. For example, the network environment (100) may include more or fewer elements than those exemplified and described in the present disclosure.

[0025] FIG. 2 is a block diagram of an exemplary electronic device for displaying sign language video through adjustable user interface (UI) elements according to an embodiment of the present disclosure. FIG. 2 is described together with the elements of FIG. 1. Referring to FIG. 2, a block diagram (200) of an electronic device (102) is shown. The electronic device (102) may include a circuit (202), a memory (204), an input / output (I / O) device (206), a network interface (208), and a neural network model (210). In at least one embodiment, the electronic device (102) may also include a display device (104). The circuit (202) may be communicably coupled to the memory (204), the I / O device (206), the network interface (208), the neural network model (210), and the display device (104).

[0026] The circuit (202) may include suitable logic, circuits, and interfaces that can be configured to execute program instructions associated with different operations to be executed by the electronic device (102). Different operations include determining the position of a receiver within a video, extracting a portion of the video corresponding to the determined position of the receiver within the video, and controlling the display device (104) to render a UI element (112) on the display device (104) and display the extracted portion of the video within the UI element (112). The circuit (202) may include one or more processing units, and the one or more processing units may be implemented as an integrated processor or a cluster of processors that collectively perform the functions of the one or more processing units. The circuit (202) may be implemented based on a number of processor technologies known in the art. Examples of implementations of the circuit (202) may be x86-based processors, Graphics Processing Units (GPUs), Reduced Instruction Set Computing (RISC) processors, Application-Specific Integrated Circuit (ASIC) processors, Complex Instruction Set Computing (CISC) processors, microcontrollers, central processing units (CPUs), and / or other computing circuits.

[0027] The memory (204) may include appropriate logic, circuits, and interfaces that can be configured to store program instructions to be executed by the circuit (202). In an embodiment, the memory (204) may store a received first media stream, a second media stream, received metadata, a determined location of the receiver (110), and an extracted video portion. The memory (204) may be further configured to store one or more user profiles associated with the electronic device (102), preferences associated with UI elements (112) for each user profile, usage history of UI elements (112) for each user profile, sign language preferences for each user profile (e.g., American Sign Language or British Sign Language), etc. In some embodiments, the memory (204) may further store one or more preset locations and one or more preset sizes of UI elements (112). Memory (204) may store one or more preset locations and one or more preset sizes of UI elements (112) as defaults for all users, or may store one or more preset locations and one or more preset sizes of UI elements (112) for each user profile. Memory (204) may also be configured to store predefined templates for image analysis, a neural network model (210), and a training dataset received from the server (106).Examples of implementations of memory (204) may include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Hard Disk Drive (HDD), Solid-State Drive (SSD), CPU cache, and / or Secure Digital (SD) card.

[0028] The I / O device (206) may include appropriate logic, circuits, and interfaces that can be configured to receive one or more input(s) and provide one or more output(s) based on the received one or more input(s). The I / O device (206), which includes various input and output devices, may be configured to communicate with the circuit (202). In one example, the electronic device (102) may receive user input through the I / O device (206) indicating a change in the current position of a UI element (112) rendered on the display device (104). In another example, the electronic device (102) may receive user input through the I / O device (206) indicating a change in the current size of a UI element (112) rendered on the display device (104). Examples of I / O devices (206) may include, but are not limited to, a remote console, a touch screen, a keyboard, a mouse, a joystick, a microphone, a display device (e.g., a display device (104)), and a speaker.

[0029] The network interface (208) may include suitable logic, circuits, and interfaces that can be configured to facilitate communication between the circuit (202) and the server (106) or display device (104) through the communication network (108). The network interface (208) may be implemented using various known technologies to support wired or wireless communication between the electronic device (102) and the communication network (108). The network interface (208) may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, a Bluetooth® receiver, an infrared receiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODE) chipset, a subscriber identity module (SIM) card, or a local buffer circuit. The network interface (208) can be configured to communicate wirelessly with networks such as the Internet, intranet, or wireless networks, such as a cellular telephone network, a wireless local area network (LAN) and a metropolitan area network (MAN).Wireless communication uses one or more of multiple communication standards, protocols, and technologies such as the Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), Long Term Evolution (LTE), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, or IEEE 802.11n), Voice over Internet Protocol (VoIP), Light Fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), email protocols, instant messaging, and Short Message Service (SMS). It is possible.

[0030] The neural network model (210) may be a system of artificial neurons arranged in multiple layers as nodes or a computational network. The multiple layers of the neural network model may include an input layer, one or more hidden layers, and an output layer. Each layer of the multiple layers may include one or more nodes (or, for example, artificial neurons represented by circles). The outputs of all nodes in the input layer may be coupled to at least one node in the hidden layer(s). Similarly, the inputs of each hidden layer may be coupled to the outputs of at least one node in other layers of the neural network model. The outputs of each hidden layer may be coupled to the inputs of at least one node in other layers of the neural network model. The node(s) of the final layer may receive inputs from at least one hidden layer and output a result. The number of layers and the number of nodes in each layer may be determined from the hyper-parameters of the neural network model. Such hyper-parameters can be set before, during, or after training the neural network model (210) on the training dataset.

[0031] Each node of the neural network model (210) may correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) having a set of parameters that can be tunable during the training of the network. The set of parameters may include, for example, weight parameters, regularization parameters, and similar ones. Each node may use the mathematical function to calculate an output based on one or more inputs from nodes of other layer(s) (e.g., previous layer(s)) of the neural network model (210). All or some of the nodes of the neural network model (210) may correspond to the same or different mathematical functions.

[0032] According to one embodiment, the circuit (202) can train a neural network model (210) on one or more features related to a video, one or more features related to the background of the signer (110) in the video, one or more features related to the hand movements of the signer (110) in the video, and obtain a trained neural network model (210). The neural network model (210) can be trained to detect hand signals associated with sign language in the video and to detect the location of the signer (110) in the video based on the detection of the hand signals. In another embodiment, the neural network model (210) can be trained to distinguish the background of the signer (110) in the video from other parts of the video and to detect the location of the signer (110) in the video based on the background. For example, the circuit (202) can train the neural network model (210) by inputting a video, a predetermined hand signal of sign language (e.g., American Sign Language or British Sign Language), etc.

[0033] In training a neural network model (210), one or more parameters of each node of the neural network model may be updated based on whether the output of the final layer for a given input (from the training dataset) matches the correct result based on the loss function for the neural network model. The process may be repeated for the same or different inputs until a minimum value of the loss function can be achieved and the training error can be minimized. Several methods for training, such as gradient descent, stochastic gradient descent, batch gradient descent, gradient boost, meta-heuristics, and similar methods are known in the art.

[0034] The neural network model (210) may include electronic data that can be implemented, for example, as a software component of an application executable on an electronic device (102). The neural network model (210) may rely on libraries, external scripts, or other logic / instructions for execution by a processing device such as a circuit (202). The neural network model (210) may include code and routines configured to enable a computing device such as a circuit (202) to perform one or more operations for detecting hand signals associated with sign language in a video. Additionally or alternatively, the neural network model (210) may be implemented using hardware including a processor, a microprocessor (e.g., one that performs or controls one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the neural network model (210) may be implemented using a combination of hardware and software.

[0035] Examples of neural network models (210) may include, but are not limited to, deep neural networks (DNN), convolutional neural networks (CNN), R-CNN, Fast R-CNN, Faster R-CNN, artificial neural networks (ANN), (You Only Look Once) YOLO networks, CNN+ANN, fully connected neural networks, deep Bayesian neural networks, and / or combinations of these networks. In some embodiments, the neural network model (210) may be based on a hybrid architecture of multiple deep neural networks (DNN).

[0036] The image analysis processor (212) may include suitable hardware and software algorithms configured to perform one or more image analysis techniques, such as object detection, object recognition, image segmentation, motion detection, pose estimation, edge detection, and template matching. For example, the image analysis processor (212) may perform template matching based on predefined features or templates associated with the shape and size of parts of the listener (110) in the video. In another example, the image analysis processor (212) may perform edge detection to detect visible boundaries around the listener (110). In yet another example, the image analysis processor (212) may perform motion detection to distinguish a stationary (static) area behind the listener (110) from a moving (dynamic) background of the main video (104A).

[0037] As described in FIG. 1, functions or operations executed by the electronic device (102) can be performed by the circuit (202). Operations executed by the circuit (202) are described in detail, for example, in FIGs. 3, 4, 5, 6 and 7.

[0038] FIG. 3 is a drawing illustrating an exemplary scenario for displaying sign language video through user interface (UI) elements adjustable based on metadata, according to an embodiment of the present disclosure. FIG. 3 is described together with the elements of FIG. 1 and FIG. 2. Referring to FIG. 3, an exemplary scenario (300) is illustrated. In the exemplary scenario (300), a block of received data (302) is illustrated. The received data (302) may include a first media stream (304) that may include one or more frames of video, and metadata (306) associated with the video. The video may include a sign language recipient (308). One or more frames may include a first frame (304A), a second frame (304B), and up to an Nth frame (304N). Referring to FIG. 3, a display device (104) associated with an electronic device (102) is also illustrated. The display device (104) may be configured to display one or more frames of a video (e.g., main video (310)).

[0039] The first media stream (304) may include one or more frames of video that can be rendered on a display device (104). For example, a main video (310) may be encapsulated within the first media stream (304). In one embodiment, the circuit (202) may receive the first media stream (304) from a server (106). In another embodiment, the circuit (202) may receive the first media stream (304) from a server associated with a broadcast network. In such a scenario, the first media stream (304) may include text information, such as an electronic program guide (EPG) associated with a broadcast channel.

[0040] Metadata (306) may include information associated with the video and may be embedded in the video's analog or digital signal. As an example, metadata (306) may include the duration of the video, the title of the video, the resolution of the video, the type of codec and / or container associated with the video, information associated with one or more characters (312) or people within the video (e.g., main video (310)) and similar information. In one embodiment, metadata (306) may include information associated with a sign language user (308) present in the video. The sign language user (308) may be an animated character or a person capable of acting using sign language in the video. In an embodiment, the sign language user (308) may translate a language (e.g., English) spoken by one or more characters (312) within the video into sign language (e.g., American Sign Language, ASL). Metadata (306) may include information that can describe the location of the sign language user (308) within the video at a plurality of timestamps. The location of the receiver (308) may include image coordinates (314) that correspond to the corners of a rectangular area of ​​the video containing the receiver (308). In an embodiment, the image coordinates (314) may be labeled with respect to pixels or image coordinates of one or more frames of the video. Examples of received metadata (306) at different timestamps are presented in Table 1 as follows:

[0041]

[0042] In one embodiment, the circuit (202) may be configured to determine the location of the receiver (308) at a plurality of time stamps based on the received metadata (306). For example, the circuit (202) may be configured to determine the location of the receiver (308) by parsing the received metadata (306). Referring to Table 1, the location of the receiver (308) in the video at a plurality of time stamps mentioned in the first column of Table 1 may be determined from the third column of Table 1. Table 1 (the third column) includes a set of four image coordinates representing the rectangular boundary of the receiver (308) in the video, but the present disclosure may not be so limited. The metadata (306) may include any number of coordinates according to the shape of the boundary of the receiver (308) (e.g., polygonal shape).

[0043] The circuit (202) may be further configured to extract a video portion (316) from the video. The extracted video portion (316) may correspond to a determined location in the video. For example, the video portion (316) may be extracted from a rectangular area of ​​the video that corresponds to the area between image coordinates (314) and includes the listener (308). The circuit (202) may be further configured to control the playback of the video on the display device (104). In one embodiment, the circuit (202) may be configured to control the playback of the video on the display device (104) based on user input. In one embodiment, the circuit (202) may be configured to control the display device (104) to render a boundary around the listener (308) within the video based on the extracted video portion (316). The display device (104) may render the boundary based on a determined location of the listener (308) within the video. A boundary around the receiver (308) may be rendered to distinguish the receiver (308) from one or more characters (312) within the video. In some embodiments, the circuit (202) may highlight the boundary around the receiver (308) with a bright color (e.g., bright green) to help the user find the receiver (308) in the video.

[0044] The circuit (202) may be configured to control the display device (104) to render a user interface (UI) element (318) to the display device (104) based on playback. The UI element (318) may be rendered as a picture-in-picture (PiP) window. The PiP window may be of adjustable size. For example, the circuit (202) may be configured to resize the PiP window based on one or more user inputs. The circuit (202) may be further configured to display an extracted video portion (316) inside the UI element (318) (e.g., the PiP window). The extracted video portion (316) may include a listener (308). In an embodiment, the circuit may connect image coordinates (314) with a line and copy the video portion (316) inside such a line to the PiP window in real time.

[0045] In one embodiment, the circuit (202) may be configured to receive a second media stream (e.g., a second signal) from the server (106) which may be different from the first media stream (304). For example, the second media stream may include a video portion (316). In this scenario, the circuit (202) may be configured to control the playback of the video from the first media stream (304) on the display device (104). The circuit (202) may be further configured to control the display device (104) to render a UI element (318) on the display device (104). The circuit (202) may control the display device (104) to display the video portion (316) extracted from the received second stream within the UI element (318) by time-synchronizing it with the playback of the video from the first media stream (304). The circuit (202) can be configured to adjust the size of the PiP window in both cases, namely when the video portion (316) is extracted based on the received metadata (306) and when the video portion is received from the server (106) as a second media stream.

[0046] In one embodiment, the circuit (202) may be further configured to receive a first input for changing the current position of a rendered UI element (318) (e.g., a PiP window) to a first position (320) that may be different from the current position. The circuit (202) may also receive a second input for changing the current size of the rendered UI element (318) (e.g., a PiP window) to a first size that is different from the current size. The circuit (202) may be further configured to control the display device (104) to render the UI element (318) at the first position (320) and at the first size based on the first input and the second input. Detailed descriptions of the position and size adjustment of the UI element (318) are provided, for example, in FIGS. 5 and 6. In one embodiment, the circuit (202) may be configured to control the display device (104) to blur the receiver (308) in the main video (310) when a UI element (318) containing the receiver (308) is displayed. In another embodiment, the circuit (202) may be configured to control the display device (104) to replace the video portion of the receiver (308) in the main video (310) with the background pixels of the main video when a UI element (318) containing the receiver (308) is displayed.

[0047] FIG. 4 is a diagram illustrating an exemplary scenario for displaying sign language video through adjustable user interface (UI) elements for live video broadcasting according to an embodiment of the disclosure. FIG. 4 is described together with the elements of FIG. 1 through 3. Referring to FIG. 4, an exemplary scenario (400) is illustrated. In the exemplary scenario (400), a block of received data (402) is illustrated. The received data (302) may include a third media stream (404) containing one or more frames of the live video broadcast. In the exemplary scenario (400), a listener (406) that may be present in the live video broadcast is further illustrated. One or more frames may include a first frame (404A), a second frame (404B), and up to an Nth frame (404N). In the exemplary scenario (400), a neural network model (210) associated with an electronic device (102) and a display device (104) are further illustrated.

[0048] The third media stream (404) may include a live video broadcast that can be broadcast by various media such as terrestrial or over-the-air broadcasting, streaming broadcasting, satellite television broadcasting, etc. For example, the live video broadcast may be encapsulated within the third media stream (404). The live video broadcast may include a main video (408) depicting one or more characters (410). In one embodiment, the third media stream (404) may be received from a server (106) or a server associated with a broadcasting network. If the third media stream (404) is a live video broadcast, the metadata to which the third media stream (404) is inserted may not include the location of the listener (406) within the live video broadcast. In this case, the circuit (202) may be configured to analyze frames of the live video broadcast to determine the location of the listener (406) in the live video broadcast. For example, the circuit (202) may be configured to perform image analysis on one or more frames of a live video broadcast to detect the area of ​​the receiver (406) within the video (e.g., an invariant background area). In another example, the circuit (202) may be configured to apply a neural network model (210) to one or more frames of a live video broadcast to detect hand signals associated with sign language within the video. For example, the circuit may be configured to apply the neural network model (210) to the first frame (404A) to detect hand signals in the first frame (404A). The circuit (202) may be configured to apply the neural network model (210) to the second frame (404B), and continue up to the Nth frame (404N), to detect hand signals in each frame of the video. In another embodiment, the circuit (202) may be configured to apply a neural network model (210) to one or more frames of a live video broadcast to distinguish the background of a portion of the video where a receiver (406) may be present from other portions of the video.The neural network model (210) may be configured to predict bounding boxes for the location of the receiver (406) in the video at multiple time stamps based on the detection of background or reception signals corresponding to the location of the receiver (406) in the video.

[0049] The neural network model (210) may be a pre-trained model that can be trained to detect sign signals associated with sign language (such as American sign language (ASL) or British sign language (BSL)) and output image coordinates (412) corresponding to the corners of a rectangular area of ​​a video portion containing the signer (406). The circuit (202) may be further configured to detect the location of the signer (406) within the video based on the image coordinates (412). The electronic device (102) can thereby identify the location of the signer (406) within the video even if the metadata embedded in the third media stream (404) does not include the location of the signer (406), or even if the metadata is not present in the third media stream (404).

[0050] In another embodiment, the circuit (202) may be configured to perform image analysis (such as object detection) using an image analysis processor (212) to detect a listener (406) in a video and output image coordinates (412). The circuit (202) may detect an area around the listener (406) in the video based on a difference in background color of an area compared with the main video (408), a difference in background shade of an area compared with the main video (408), or a predefined boundary around the listener. For example, the circuit (202) may be configured to detect the location of the listener (406) within a video (such as the main video (408)) when the video portion containing the listener has a background color (or shade) different from the background color (or shade) of other portions of the video. In another example, the circuit (202) may be configured to detect a background area that is not moving (static) and different from the background of the main video (408) to detect the location of the listener (406). In another example, the circuit (202) may use edge detection or template matching techniques to detect boundaries around the listener (406) when the video contains visible boundaries of a predefined shape and color around the listener (406), and may detect the location of the listener (406) within the video based on the detected boundaries. In these scenarios, the circuit (202) may rely on image analysis techniques that require less computational power than the execution of the neural network model (210).

[0051] In one embodiment, the circuit (202) may be further configured to control the display device (104) to render a highlighted boundary (406A) around the receiver (406) in the displayed video. The boundary (406A) may be rendered based on the image coordinates (412) of the predicted bounding box. As an example, the circuit (202) may display the boundary (406A) to obtain user confirmation of the video portion of the receiver (406) in the video when the neural network model (210) identifies multiple candidates as video portions of the receiver (406) in the video. For example, the circuit (202) may receive user confirmation for the highlighted displayed candidate by a prompt displayed on the display device (104) (Press OK to confirm; press the right arrow ▶ to see the next candidate). The circuit (202) can obtain user identification of the receiver (406) in the video when the confidence score of the detection and / or bounding box prediction of the receiver (406) may be lower than the threshold score. In another embodiment, the circuit (202) can obtain user identification of the receiver (406) in the video when the image analysis processor (212) outputs a number of candidates for the video portion of the receiver (406) in the video.

[0052] The circuit (202) may be further configured to extract a video portion (414) from a live video broadcast. In one embodiment, the circuit (202) may be further configured to extract a video portion (414) from a live video broadcast based on user verification of an highlighted candidate. The extracted video portion (414) may correspond to a determined location in the live video broadcast. For example, the video portion (414) may be extracted from a rectangular area of ​​the video. The rectangular area may correspond to an area between image coordinates (412) and may include a listener (406).

[0053] The circuit (202) may be further configured to control the playback of video on the display device (104). In one embodiment, the circuit (202) may be configured to control the playback of video on the display device (104) based on user input. The circuit (202) may be configured to control the display device (104) based on playback. The circuit (202) may control the display device (104) to render a user interface (UI) element (416) on the display device (104). For example, the UI element (416) may be rendered as a picture-in-picture (PiP) window of adjustable size. The circuit (202) may be configured to display an extracted video portion (414) containing a receiver (406) inside the UI element (416).

[0054] FIG. 5 is a drawing illustrating an exemplary scenario for adjusting the position of a user interface (UI) element for displaying a sign language video according to an embodiment of the present disclosure. FIG. 5 is described together with the elements of FIG. 1 through 4. Referring to FIG. 5, an exemplary scenario (500) is illustrated. In the exemplary scenario (500), an electronic device (102) and a display device (104) associated with the electronic device (102) are illustrated. The electronic device (102) can control the display device (104) to display a main video (502) within a display area (506). Referring to FIG. 5, a user (508) associated with the electronic device (102) is further illustrated.

[0055] In one embodiment, the circuit (202) may be configured to receive user input including a selection of a user profile associated with the user (508). Based on the selected user profile, the circuit (202) may retrieve one or more user preferences associated with a user interface (UI) element (510) on which an extracted video portion (512) of the receiver (516) may be displayed. In some embodiments, the circuit (202) may retrieve one or more user preferences from memory (204). The UI element (510) may be rendered based on the one or more retrieved user preferences. For example, one or more user preferences may include one or more of a position preference for a UI element (510) within a display area (506) of a display device (104), a theme or color scheme for a UI element (510), a size preference for a UI element (510), a show / hide preference for a UI element (510), a schedule for rendering a UI element (510), and a sign language preference (e.g., American Sign or British Sign).

[0056] Location preferences may include preferred locations where the UI element (510) can be displayed. The circuit (202) may retrieve a user preference for a first location (514) from a set of locations. The first location (514) may be a preferred location for displaying the UI element (510) according to the user profile of the user (508). As an example, the first location (514) may correspond to the bottom right corner within the display area (506) of the display device (104). The theme or color scheme for the UI element (510) may correspond to the design or color preference of the selected user profile for the UI element (510). As an example, the theme or color scheme for the UI element (510) may include a green background behind the receiver (516) or a green color boundary for the UI element (510). Size preferences for the UI element (510) may include a default size predefined by the manufacturer of the electronic device (102). The preference for hiding UI elements (510) may correspond to the preference of the user (508) regarding whether to hide or show the UI elements (510). The schedule for rendering the UI elements (510) may correspond to a first period during which the UI elements (510) may be rendered and a second period during which the UI elements (510) may not be rendered. For example, the user preference for the schedule may indicate that the UI elements (510) may be rendered between 10:00 AM and 4:00 PM, and that the UI elements (510) may be hidden between 4:01 AM and 10:00 PM. In another embodiment, the user preference may include a command to show the UI elements (510) when one of the characters (504) in the main video (502) is speaking, and to hide the UI elements (510) when there is no voice in the main video (502).

[0057] At time T1, the circuit (202) may receive a first media stream that may include a video. The video may include a main video (502) depicting a character (504). The circuit (202) may further receive metadata associated with the video. The circuit (202) may further determine the location of the receiver (516) within the video based on the received metadata. The metadata may include information describing the location of the receiver (516) within the video at multiple timestamps. In another embodiment, the circuit (202) may determine the location of the receiver (516) based on image analysis by an image analysis processor (212) or based on the application of a neural network model (210). The circuit (202) may further extract a video portion (512) associated with the determined location of the receiver (516) within the video. Based on the extracted location, the circuit (202) may control the playback of the video on the display device (104). The circuit (202) may further control the display device (104) to render a UI element (510) (e.g., a PiP window) on the display device (104) at a first location (514) based on a retrieved user preference. The circuit (202) may control the display device (104) to display a video portion (512) extracted inside the UI element (510). In some embodiments, if a user preference for the location of the UI element (510) is not available in memory (204), the UI element (510) may overlap with a determined location of the receiver (516) in the main video (502) according to a default location predefined by the manufacturer of the electronic device (102). As illustrated in FIG. 5, the circuit (202) may control the display device (104) to render the UI element (510) at the bottom right corner of the display area (506) of the display device (104).

[0058] The circuit (202) may receive a first input (518) for changing the current position (or first position (514)) of a rendered UI element (510) to a second position (520). The second position (520) may be different from the first position (514). If the electronic device (102) is a television controlled by a remote control, the display device (104) may display a pop-up menu (510A) (e.g., a context menu) when the UI element (510) is selected. The pop-up menu (510A) may include "Resize" and "Move" options. When the "Move" option (a selection highlighted in gray) is selected, the display device (104) may display sub-options of "Move to preset position" and "Drag". When the "Move to preset location" (selection highlighted in gray) option is selected, the display device (104) may display sub-options such as "preset location 1", "preset location 2", etc., based on the stored preferences of the selected user profile and / or default locations set by the manufacturer of the electronic device (102). For example, "preset location 1" may correspond to the lower left corner of the display area (506), and "preset location 2" may correspond to the upper left corner of the display area (506). These preset locations may be stored in memory (204) based on the set preferences of the selected user profile and / or default locations set by the manufacturer of the electronic device (102). When one of the sub-options is selected, the circuit (202) may control the display device (104) to display a UI element (510) at a second location (520) at time T2. As an example, the second position (520) may correspond to the lower left corner within the display area (506) of the display device (104).When the “Drag” option is selected, the display device (104) may highlight the UI element (510) to indicate that the UI element (510) has been selected and may display a prompt to drag the UI element (510) to any arbitrary location within the display area (506) using the arrow buttons (▶◀▼▲) on the remote control. If the electronic device (102) is a smartphone with touchscreen input, when the UI element (510) is selected, the display device (104) may display a prompt to drag and move the UI element (510) to any arbitrary location within the display area (506) by touch input. The circuit (202) may control the display device (104) to display the UI element (510) at a second location (520) (e.g., bottom left corner) at time T2 based on the first input (518). The display device (104) can smoothly continue the playback of the extracted video portion (512) of the receiver (516) by time-synchronizing with the main video (502) before, during, and after the movement of the UI element (510).

[0059] FIG. 6 is a drawing illustrating an exemplary scenario for adjusting the size of a user interface (UI) element for displaying a sign language video according to an embodiment of the present disclosure. FIG. 6 illustrates an exemplary scenario (600). In the exemplary scenario (600), an electronic device (102) and a display device (104) associated with the electronic device (102) are illustrated. The electronic device (102) can control the display device (104) to display a main video (602) within a display area (606).

[0060] In an embodiment, the circuit (202) may be configured to receive user input including a selection of a user profile. Based on the selected user profile, the circuit (202) may retrieve one or more user preferences associated with a UI element (610) on which an extracted video portion of the receiver (608) may be displayed. In some embodiments, the circuit (202) may retrieve one or more user preferences from memory (204). The UI element (610) may be rendered based on the retrieved one or more user preferences. For example, one or more user preferences may include size preferences for the UI element (610).

[0061] At time T1, the circuit (202) may receive a first media stream that may include a video. The video may include a main video (602) depicting a character (604). The circuit (202) may further receive metadata associated with the video. Based on the received metadata, the circuit (202) may further determine the location of the receiver (608) within the video. In another embodiment, the circuit (202) may determine the location of the receiver (608) based on image analysis by an image analysis processor (212) or based on the application of a neural network model (210). The circuit (202) may further extract a portion of the video corresponding to the determined location within the video. Based on the extracted location, the circuit (202) may be configured to control the playback of the video on the display device (104). The circuit (202) may be further configured to control the display device (104) to render a user interface (UI) element (610), such as a PiP window, to the display device (104) in a first size (e.g., height H1, width W1) based on a retrieved user preference associated with a selected user profile. In some embodiments, if a user preference for the size of the UI element (610) is not available in memory (204), the UI element (610) may be displayed based on a default size predefined by the manufacturer of the electronic device (102). As illustrated in FIG. 6, the circuit (202) may control the display device (104) to render the UI element (610) in a first size (H1, W1) in a display area (606) of the display device (104).

[0062] The circuit (202) may receive a second input (612) to change the current size (or first size) of the rendered UI element (610) to a second size. The second size may be different from the first size. If the electronic device (102) is a television controlled by a remote control, the display device (104) may display a pop-up menu (610A) when the UI element (610) is selected. The pop-up menu (610A) may include "Resize" and "Move" options. When the "Resize" option (selected in gray) is selected, the display device (104) may display "Resize to preset sizes" and "Zoom in / out" sub-options. When the "Resize to preset sizes" (selection highlighted in gray) option is selected, the display device (104) may display sub-options such as "preset size 1", "preset size 2", based on the retrieved preferences of the selected user profile and / or default positions set by the manufacturer of the electronic device (102). For example, "preset size 1" and "preset size 2" may correspond to different sizes having a fixed aspect ratio so that the extracted video portion of the receiver (608) has optimal resolution. When one of the sub-options is selected, the circuit (202) may control the display device (104) to display the UI element at time T2 in a second size (height H2, width W2). When the “zoom in / out” option is selected, the display device (104) can highlight the UI element (610) to indicate that the UI element (610) has been selected, and can display a prompt to resize the UI element (610) to any arbitrary size within the display area (606) using the arrow buttons (▶◀▼▲) on the remote control.If the electronic device (102) is a smartphone having touchscreen input, when a UI element (610) is selected, the display device (104) may display a prompt to resize the UI element (610) to any arbitrary size within the display area (606) using touch-based actions (e.g., pinch-open or pinch-close actions of fingers). The circuit (202) may control the display device (104) to display the UI element at a second size (H2, W2) at time T2 based on the second input (612). For example, the circuit (202) may control the display device (104) to change the current size of the UI element (610) to match the second size (H2, W2). In an embodiment, the circuit (202) may be configured to upscale or downscale a video portion to match a second size (H2, W2) of the UI element (610) before the video portion is displayed inside the UI element (610). The circuit (202) may upscale or downscale to change the resolution of the extracted video portion of the listener (608) according to the modified size of the UI element (610). As illustrated in FIG. 6, the second size (H2, W2) of the UI element (610) may be larger than the first size (H1, W1) of the UI element (610). In such a case, the circuit (202) may upscale the extracted video portion of the listener (608) from a lower resolution (e.g., 720p) to a higher resolution (e.g., 1080p).

[0063] FIG. 7 is a flowchart illustrating exemplary operations for displaying sign language video through adjustable user interface (UI) elements according to an embodiment of the present disclosure. FIG. 7 is described together with the elements of FIG. 1 through 6. Referring to FIG. 7, a flowchart (700) is illustrated. The operations of 702 through 712 may be implemented by any computing system, such as the electronic device (102) of FIG. 1 or the circuit (202) of FIG. 2. The operations may start at 702 and proceed to 704.

[0064] In 704, a first media stream including video can be received. In at least one embodiment, the circuit (202) may be configured to receive a first media stream including video, for example, as described in FIGS. 1, 3 and 4.

[0065] In 706, the position of the receiver (110) within the video can be determined, wherein the receiver (110) may be an animated character or person performing using sign language in the video. In at least one embodiment, the circuit (202) may be configured to determine the position of the receiver (110) within the video. A detailed description of the determination of the position of the receiver (110) is provided in FIGS. 1, FIGS. 3, and FIGS. 4.

[0066] In 708, a portion of the video corresponding to a determined location within the video can be extracted from the video. In at least one embodiment, the circuit (202) may be configured to extract a portion of the video corresponding to a determined location within the video from the video. A detailed description of the extraction of the video portion is provided, for example, in FIGS. 1, 3, and 4.

[0067] In 710, playback of video in the display device (104) can be controlled. In at least one embodiment, the circuit (202) can be configured to control playback of video in the display device (104).

[0068] In 712, the display device (104) may be controlled based on playback to render a user interface (UI) element (112) to the display device (104) and to display an extracted video portion inside the UI element (112). In at least one embodiment, the circuit (202) may be configured to control the display device (104) based on playback to render the UI element (112) to the display device (104) and to display an extracted video portion inside the UI element (112). A detailed description regarding rendering the UI element (112) is provided, for example, in FIGS. 1, 3, 4, and 5. Control may be terminated.

[0069] Various embodiments of the present disclosure may provide a non-transient computer-readable medium and / or storage medium storing instructions executable by a machine and / or computer to operate an electronic device such as an electronic device (102). The instructions may cause the machine and / or computer to perform operations including receiving a first media stream containing video. The operations may further include determining the position of a receiver (such as a receiver (110)) within the video. The receiver may be an animated character or person capable of acting using sign language in the video. The operations may further include extracting a portion of the video from the video that corresponds to the determined position of the video. The operations may further include controlling the playback of the video on a display device (such as a display device (104)). The operations may further include controlling the display device (104) based on playback to render a user interface (UI) element (such as a UI element (112)) on the display device (104).

[0070] An exemplary embodiment of the present disclosure may include an electronic device (such as the electronic device (102) of FIG. 1) comprising a circuit (such as the circuit (202)) that can be communicably coupled to a display device (such as the display device (104)). In one embodiment, the electronic device (102) may be configured to receive a first media stream including video. The receiver (110) may be an animated character or person performing using sign language in the video. The electronic device (102) may be configured to determine the position of the receiver (110) in the video. The determined position may include image coordinates corresponding to the corners of a rectangular area of ​​the video containing the receiver (110).

[0071] According to one embodiment, an electronic device (102) may receive metadata associated with a video. The metadata includes information describing the location of a receiver (110) within the video at a plurality of timestamps. The electronic device (102) may determine the location of a receiver (110) within the video based on the received metadata.

[0072] According to one embodiment, the electronic device (102) may be configured to detect hand signals associated with sign language within the video based on the application of a neural network model (such as the neural network model (210)) on the frames of the video. The electronic device (102) may be further configured to detect the location of the receiver (110) in the video based on the detection of the hand signals. In this embodiment, the video may correspond to a live video broadcast.

[0073] According to one embodiment, the electronic device (102) can detect a region within the video based on the difference between the background of the region and the background of the rest of the video. The electronic device (102) can detect the location of the receiver (110) in the video based on the detection of the region. In another embodiment, the electronic device (102) can detect a boundary around the receiver (110) in the video. The electronic device (102) can detect the location of the receiver within the video based on the detection of the boundary.

[0074] According to one embodiment, the electronic device (102) may be configured to extract a portion of video corresponding to a determined location within the video from the video. The portion of video is extracted from a rectangular area of ​​the video. The electronic device may be further configured to control the playback of the video on the display device (104). The electronic device (102) may be further configured to control the display device based on playback to render a user interface (UI) element (such as a UI element (112)) on the display device (104) and to display the extracted portion of video inside the UI element (112). The UI element (112) may be rendered as a picture-in-picture (PiP) window of adjustable size. In an embodiment, the electronic device (102) may be further configured to control the display device (104) to render a boundary around the receiver (110) in the displayed video based on a determined location.

[0075] According to one embodiment, an electronic device (102) may receive a first user input including one or more user preferences associated with a UI element. The UI element may be rendered based on the received first user input. One or more user preferences may include a position preference for the UI element (112) within a display area (such as a display area (506)) of a display device (104), a theme or color scheme for the UI element (112), a size preference for the UI element (112), a hiding preference for the UI element (112), and a schedule for rendering the UI element (112).

[0076] According to one embodiment, the electronic device (102) may be configured to receive a first input (such as a first input (518)) for changing the current position of a rendered UI element (112) to a first position different from the current position. The electronic device (102) may be further configured to control a display device (104) based on the first input to render the UI element (112) at the first position. The first position may be within the display area (506) of the display device (104).

[0077] According to an embodiment, the electronic device (102) may be configured to receive a second input (such as a second input (612)) for changing the current size of a rendered UI element (112) to a first size different from the current size. The electronic device (102) may be configured to control a display device (104) to change the current size of the rendered UI element (112) to match the first size based on the received second input. The electronic device (102) may be further configured to upscale or downscale a video portion to match the first size of the UI element (112) before the video portion is displayed inside the UI element (112).

[0078] According to an embodiment, the electronic device (102) may be configured to receive a second media stream including an extracted video portion. The second media stream may be different from the first media stream.

[0079] The present disclosure may be realized in hardware, or in a combination of hardware and software. The present disclosure may be realized in a centralized manner in at least one computer system, or in a distributed manner in which different elements may be distributed across multiple interconnected computer systems. A computer system or other device adapted to perform the methods described herein may be suitable. A combination of hardware and software may be a general-purpose computer system having a computer program capable of controlling the computer system to perform the methods described herein when the computer program is loaded and executed. The present disclosure may be realized in hardware comprising a part of an integrated circuit that also performs other functions.

[0080] The present disclosure may also be embedded in a computer program product capable of executing the methods described herein when loaded into a computer system, and may include all features that enable the implementation of the methods described herein. In this context, a computer program means a set of instructions expressed in any language, code, or notation to enable a system having information processing capabilities to perform a specific function immediately, or after one or both of a) conversion into another language, code, or notation; or b) reproduction into another material form.

[0081] Although the present disclosure has been described with reference to specific embodiments, it will be understood by those skilled in the art that various modifications and substitutions may be made without departing from the scope of the present disclosure. Furthermore, many modifications may be made to adapt specific situations or materials to the teachings of the present disclosure without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific embodiments disclosed, and is intended to include all embodiments falling within the scope of the appended claims.

Claims

Claim 1 An electronic device comprising a circuit communicably coupled to a display device, wherein the circuit comprises: receiving a first media stream comprising video; determining the position of a signer within the video — the signer is an animated character or person performing using sign language in the video —; extracting a portion of video corresponding to the determined position of the signer within the video from the video; controlling the playback of the video on the display device; rendering user interface (UI) elements on the display device; and further configured to control the display device based on playback to display the extracted portion of video within the UI elements, and further configured to control the display device to render a boundary around the signer in the displayed video based on the determined position. Claim 2 An electronic device according to claim 1, wherein the UI element is rendered as a picture-in-picture (PiP) window of adjustable size. Claim 3 An electronic device according to claim 1, wherein the circuit comprises: receiving metadata associated with the video—the metadata comprising information describing the location of a receiver within the video at a plurality of timestamps—; and additionally configured to determine the location of the receiver within the video based on the received metadata. Claim 4 An electronic device according to claim 1, wherein the determined position includes image coordinates corresponding to the corners of a rectangular area of ​​the video including the receiver, and the video portion is extracted from the rectangular area of ​​the video. Claim 5 An electronic device according to claim 1, wherein the circuit is further configured to: detect hand signals associated with sign language in the video based on the application of a neural network model to the frames of the video; and detect the position of the receiver in the video based on the detection of the hand signals. Claim 6 In paragraph 5, the above video is an electronic device corresponding to a live video broadcast. Claim 7 An electronic device according to claim 1, wherein the circuit is further configured to: detect the region within the video based on the difference between the background of the region and the background of the rest of the video; and detect the location of the receiver in the video based on the detection of the region. Claim 8 An electronic device according to claim 1, wherein the circuit is further configured to: detect a boundary around the receiver in the video; and detect the position of the receiver in the video based on the detection of the boundary. Claim 9 delete Claim 10 An electronic device according to claim 1, wherein the circuit receives a first input for changing the current position of a rendered UI element to a first position different from the current position; and further configured to control the display device based on the first input to render the UI element to the first position within the display area of ​​the display device. Claim 11 An electronic device according to claim 1, wherein the circuit comprises: receiving a second input for changing the current size of a rendered UI element to a first size different from the current size; controlling the display device to change the current size of the rendered UI element to match the first size based on the received second input; and further configured to control the display device to upscale or downscale the video portion to match the first size of the UI element before the video portion is displayed inside the UI element. Claim 12 In claim 1, the circuit is further configured to receive a second media stream including the extracted video portion, and the second media stream is different from the first media stream, an electronic device. Claim 13 An electronic device according to claim 1, wherein the circuit is further configured to receive a first user input including one or more user preferences associated with the UI element, and the UI element is rendered based on the received first input. Claim 14 An electronic device according to claim 13, wherein one or more user preferences include one or more of a position preference for the UI element within a display area of ​​the display device, a theme or color scheme for the UI element, a size preference for the UI element, a hiding preference for the UI element, and a schedule for rendering the UI element. Claim 15 A method comprising: receiving a first media stream containing a video; determining the position of a receiver within the video — the receiver is an animated character or person performing using sign language in the video —; extracting a video portion from the video corresponding to the determined position of the receiver within the video; controlling the playback of the video on a display device; and rendering user interface (UI) elements on the display device; and controlling the display device based on playback to display the extracted video portion within the UI elements, and further comprising the step of controlling the display device to render a boundary around the receiver in the displayed video based on the determined position. Claim 16 In paragraph 15, the method wherein the UI element is rendered as a picture-in-picture (PiP) window of adjustable size. Claim 17 A method according to claim 15, further comprising the step of receiving metadata associated with the video—the metadata including information describing the location of the receiver within the video at a plurality of timestamps—and the step of determining the location of the receiver within the video based on the received metadata. Claim 18 In paragraph 15, the determined position comprises image coordinates corresponding to the corners of the rectangular area of ​​the video including the receiver, and the video portion is extracted from the rectangular area of ​​the video. Claim 19 A method according to claim 15, further comprising: a step of detecting hand signals associated with sign language within the video based on the application of a neural network model to the frames of the video; and a step of detecting the position of the signer within the video based on the detection of the hand signals. Claim 20 A non-transient computer-readable medium storing computer-executable instructions that, when executed by an electronic device, cause the electronic device to execute operations, wherein the operations include: receiving a first media stream containing video; determining the position of a receiver in the video — the receiver is an animated character or person performing using sign language in the video —; extracting a portion of the video corresponding to the determined position of the receiver in the video from the video; controlling the playback of the video on a display device; and rendering user interface (UI) elements on the display device; and controlling the display device based on playback to display the extracted portion of the video within the UI elements, and further comprising the step of controlling the display device to render a boundary around the receiver in the displayed video based on the determined position.