Multi-source multimedia output and synchronization

By implementing multi-source multimedia output and synchronization on a computing device, the problem of difficulty in rendering audio and video components of multimedia content in the prior art is solved, and the effect that users can listen and watch multimedia content at the same time is achieved.

CN120226366APending Publication Date: 2025-06-27QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380066296.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-22
Filing Date
2023-08-01
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to render audio and video components synchronously in the presentation of multimedia content, especially under environmental noise or venue limitations, resulting in consumers being unable to listen and view multimedia content at the same time.

Method used

By implementing multi-source multimedia output and synchronization on a computing device, receiving user input, detecting and identifying multimedia content, obtaining identified multimedia content from the source of multimedia content, and rendering audio or video components in synchronization with rendering of a remote media player.

Benefits of technology

Synchronous rendering of audio and video components of multimedia content rendered on a remote media player within a user-aware distance is realized, ensuring that users can listen and watch multimedia content at the same time, improving users' multimedia experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226366A_ABST
    Figure CN120226366A_ABST
Patent Text Reader

Abstract

Various embodiments may include a method for multi-source multimedia output and synchronization in a computing device. The method may include receiving user input that selects one of an audio component, a video component, or other perceptible media component associated with multimedia content being rendered by a remote media player within a perceptible distance of a user of the computing device. The user input indicates that the user wishes to render the selected audio component, video component, or other perceptible media component on the computing device. The method also identifies the multimedia content and obtains the identified multimedia content from a source of the multimedia content. The method also renders, by the computing device, a selected one of the audio component, video component, or other perceptible media component from the obtained multimedia content in synchronization with the rendering by the remote media player within the perceptible distance of the user.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims the benefit of priority of Greek Patent Application No. 20220100777, filed on September 22, 2022; the entire content of the Greek patent application is incorporated herein by reference. Background Art

[0003] The presentation of multimedia content has become ubiquitous on media players deployed in many bars, restaurants, stadiums, airports, event venues, homes, and other private and / or public locations. However, the ability to perceive both the audio and video components of the multimedia content being presented is often limited. For example, although monitors displaying different multimedia content can be found in a bar, the ambient noise typically means that people can see the video but not hear the audio. Whether it is the proximity to the speakers of the media player, the ambient noise in the area, or a combination of both, people are often restricted to only watching and not listening to the multimedia content. In cases where there are multiple monitors streaming different content, typically only the video component of the corresponding multimedia content is rendered, which means that consumers cannot hear the audio component intended to accompany the rendered video component. Additionally, venues like bars sometimes only provide speakers that broadcast the audio component of multimedia content (such as a competition or a sports event), leaving the listener to imagine what the event looks like. Consumers could benefit from an augmented reality (XR) device configured to provide multi-source multimedia output and synchronization. Summary of the Invention

[0004] Aspects of the present disclosure include methods, systems, and devices for rendering an audio component, a video component, or other perceivable media component of a multimedia stream being rendered on a remote multimedia device (such as on a computing device or more specifically an augmented reality (XR) device in synchronization with the multimedia stream) in a manner perceivable by a user. Aspects may include a method for multi-source multimedia output and synchronization in a computing device. The method may include receiving a user input that selects one of an audio component, a video component, or other perceivable media component associated with multimedia content being rendered by a remote media player within a perceivable distance of a user of the computing device. The user input indicates that the user desires to render the selected audio component, video component, or other perceivable media component on the computing device. The method also identifies the multimedia content and obtains the identified multimedia content from a source of the multimedia content. The method also renders, by the computing device, in synchronization with the rendering performed by the remote media player within the perceivable distance of the user, the selected one of the audio component, video component, or other perceivable media component from the obtained multimedia content.

[0005] In some aspects, the multimedia content may be selected by the user from among a plurality of multimedia contents observable by the user. In some aspects, receiving user input may include: detecting a gesture performed by the user; interpreting the detected gesture to determine whether it identifies the multimedia content being rendered within a threshold distance of the computing device; and identifying one of the audio component, video component, or other perceivable media component of the identified multimedia content that the user desires to be rendered on the computing device.

[0006] In some aspects, identifying the multimedia content being rendered on a display within a perceivable distance of the user of the computing device may include: detecting the user's gaze direction; and identifying the multimedia content being rendered on the display in the user's gaze direction.

[0007] In some aspects, identifying the multimedia content being rendered on the display within a perceivable distance of the user of the computing device may include: receiving user input indicating the direction from which the user is perceiving the multimedia content; and identifying the multimedia content based on the received user input.

[0008] In some aspects, obtaining the identified audio component, video component, or other perceivable media component of the multimedia content from the source of the multimedia content may include: obtaining metadata about the multimedia content; using the obtained metadata to identify the source of the multimedia content; and obtaining the audio component, video component, or other perceivable media component from the identified source of the multimedia content.

[0009] In some aspects, obtaining the identified audio component, video component, or other perceivable media component of the multimedia content from the source of the multimedia content may include: sending a query about the multimedia content to a remote computing device; requesting an identification of the source of the multimedia content; and obtaining the audio component, video component, or other perceivable media component from the identified source of the multimedia content.

[0010] Some aspects may include sampling one of the audio component or video component being rendered by the remote media player. The query sent may include at least a portion of one of the sampled audio component or video component. At least one of the identification of the source of the multimedia content or the synchronization with the rendering performed by the remote media player may be based on information received in response to the query sent.

[0011] In some aspects, obtaining the identified multimedia content from the source of the multimedia content may include: obtaining a subscription access to the multimedia content from the source of the multimedia content; and receiving the identified audio component, video component, or other perceivable media component of the multimedia content based on the obtained subscription access.

[0012] In some aspects, rendering, by the computing device, a selected one of the audio component, the video component, or other perceivable media components synchronously with the rendering performed by the remote media player within the perceivable distance of the user may include: sampling one or more of the audio component, the video component, or other perceivable media components of the multimedia content being rendered by the remote media player within the perceivable distance of the user. Additionally, a timing difference between samples of one or more of the audio component, the video component, or other perceivable media components of the multimedia content being rendered and the audio component, the video component, or other perceivable media components obtained from the source of the multimedia content may be determined. Moreover, the computing device may render a selected one of the audio component, the video component, or other perceivable media components such that the user will perceive the selected one of the audio component, the video component, or other perceivable media components so rendered as being synchronous with the multimedia content rendered by the remote media player.

[0013] In some aspects, the computing device may be an augmented reality (XR) device.

[0014] Other aspects include a computing device configured with a processor for performing one or more operations of any of the methods outlined above. Additional aspects may include a non-transitory processor-readable storage medium having stored thereon processor-executable instructions configured to cause a processor of a computing device to perform the operations of any of the methods outlined above. Additional aspects include a computing device having components for performing the functions of any of the methods outlined above. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings incorporated herein and constituting a part of this specification illustrate exemplary embodiments of the claims and, together with the general description given above and the detailed description given below, serve to explain the features of the claims.

[0016] Figure 1A is a block diagram of components of a multi-source multimedia environment suitable for implementing various embodiments.

[0017] Figure 1B is a block diagram of components of another multi-source multimedia environment suitable for implementing various embodiments.

[0018] Figure 1C is a block diagram of components of another multi-source multimedia environment suitable for implementing various embodiments.

[0019] Figure 1D is a block diagram of components of another multi-source multimedia environment suitable for implementing various embodiments.

[0020] Figure 2A It is a schematic diagram of a gesture-based user input technology applicable to implementing various embodiments.

[0021] Figure 2B It is a schematic diagram of a gaze-based user input technology 201 applicable to implementing various embodiments.

[0022] Figure 2C It is a schematic diagram of a screen-based user input technology 202 applicable to implementing various embodiments.

[0023] Figure 2D It is a schematic diagram of another screen-based user input technology 203 applicable to implementing various embodiments.

[0024] Figure 2E It is a schematic diagram of another XR overlay-based user input technology 204 applicable to implementing various embodiments.

[0025] Figure 3 It is a component block diagram illustrating an example computing and wireless modem system on a chip applicable to use in a computing device for implementing any of various embodiments.

[0026] Figure 4A It is a communication flowchart illustrating a method for multi-source multimedia output and synchronization in a computing device according to various embodiments.

[0027] Figure 4B It is a communication flowchart illustrating an example method 401 for multi-source multimedia output and synchronization in a computing device according to various embodiments.

[0028] Figure 5A It is a process flowchart illustrating a method for multi-source multimedia output and synchronization in a computing device according to various embodiments.

[0029] Figure 5B It is a process flowchart illustrating additional operations executable by a processor of a computing device according to various embodiments.

[0030] Figure 6 It is a component block diagram of a user mobile device applicable to various embodiments.

[0031] Figure 7 It is a component block diagram of an example of smart glasses applicable to various embodiments.

[0032] Figure 8 It is a component block diagram of a server applicable to various embodiments. Detailed Description

[0033] Various embodiments will be described in detail with reference to the accompanying drawings. Wherever possible, the same reference numerals will be used throughout the drawings to refer to the same or like parts. References to specific examples and specific implementations are for illustrative purposes only and are not intended to limit the scope of the claims.

[0034] Various embodiments provide a user device (e.g., a computing device) that is configured to: receive a user input requesting to render an audio component, a video component, or other perceivable media component of a multimedia content stream being rendered on a nearby media player, detect and identify the multimedia content, obtain the requested multimedia content component from a source of the multimedia content, and render a selected audio component, video component, or other perceivable media component from the obtained multimedia content in a manner synchronized with the multimedia content being rendered. Various embodiments enable a user who can see but not hear, can hear but not see, can feel but not hear, or can feel but not see the multimedia being rendered nearby to receive an audio component, a video component, or other perceivable media component in the user's mobile device (such as an XR device). Various embodiments include user interaction methods that enable a user to identify multimedia content that the user desires to hear, see, or feel on the user's mobile device.

[0035] Generally speaking, various embodiments may include receiving a user input that selects one of an audio component, a video component, or other perceivable media component associated with multimedia content being rendered by a remote media player within a perceivable distance of the user of the computing device. The user input may indicate that the user desires to render the selected audio component, video component, or other perceivable media component on the computing device. The method may further include identifying the multimedia content and obtaining the identified multimedia content from a source of the multimedia content. Additionally, the computing device may render a selected one of the audio component, video component, or other perceivable media component from the obtained multimedia content synchronously with the rendering being performed by the remote media player within the perceivable distance of the user.

[0036] A solution for a venue with numerous monitors displaying different and changing multimedia content is to provide the system with separate audio and video components, such as wireless devices (e.g., headphones or mobile video players) that can receive media from a local network. Such a separate audio / video system can allow customers to select from different radio stations or channels that broadcast the audio component of the multimedia content (e.g., via a wireless router), while displaying the video component on a separate monitor. Such a separate audio / video system does not require identification of the multimedia content being played, as the local wireless network or preset radio stations or channels make content identification unnecessary. Additionally, speakers or media players that receive broadcasts on preset radio stations or channels only output the audio stream when it is received. Therefore, synchronization of audio output is not required or possible, as the equipment only receives the audio component of the separate audio / video components. Additionally, the separate audio / video components prevent the use of many enhanced features for computing devices used in conjunction with multimedia content, such as haptic feedback, augmented reality, augmented reality overlays, etc. Additionally, such a separate audio / video system requires custom on-site hardware that demultiplexes the multimedia content and sends the audio and video components separately. Additionally, such a separate audio / video system is limited to specific venues equipped with such capabilities.

[0037] Various embodiments provide solutions for scenarios where a user of a mobile device (such as an XR device) may be unable to hear the audio component of multimedia content being rendered on a nearby remote display. Specifically, a computing device can detect multimedia content streams on one or more nearby devices and identify or recognize the multimedia content. Once the multimedia content is identified, the computing device can obtain a copy of its own multimedia content stream and provide the user of the computing device with dynamically synchronized audio and video components of the multimedia content. The multimedia streams discussed can be live broadcasts (e.g., sports, news, etc.), scheduled programs (e.g., major television broadcasts), or on-demand replays (e.g., Netflix, YouTube, Amazon Video, etc.).

[0038] As used herein, the term "computing device" refers to an electronic device equipped with at least a processor, a memory, and a device for presenting output (such as the audio / video components of multimedia content). In some embodiments, the computing device may include a wireless communication device (such as a transceiver and an antenna) configured to communicate with a wireless communication network. The computing device may include an augmented / virtual reality device, a cellular phone, a smart phone, a portable computing device, a personal or mobile multimedia player, a laptop computer, a tablet computer, a 2-in-1 laptop / desktop computer, a smartbook, an ultrabook, a cellular phone supporting multimedia Internet, an entertainment device (e.g., a wireless game controller, a music and video player, a satellite radio component, etc.), a smart ring, a smart necklace, smart glasses, smart contact lenses, a non-contact sleep tracking device, smart furniture such as a smart bed or a smart sofa, smart exercise equipment, an Internet of Things (IoT) device, and any one or all of similar electronic devices including a memory, a wireless communication component, and a programmable processor. In some embodiments, the computing device may be a personal wearable device. As used herein, the term "smart" in connection with a device refers to a device including a processor for automated operation, for collecting and / or processing data, and / or programmable to perform all or part of the operations described with respect to the various embodiments.

[0039] The term "augmented reality device" or "XR device" may be used interchangeably herein to refer to one or more mobile computing devices configured to allow a user of the XR device to experience augmented reality (AR), virtual reality (VR), mixed reality (MR), and everything in between. The XR device may be a single integrated electronic device or a combination of separate electronic devices. For example, a single electronic device forming the XR device may include and combine functionality from a smart phone, a mobile VR headset, and AR glasses into a single wearable XR. Alternatively, one or more of a smart phone, a mobile VR headset, AR glasses, and / or other computing devices may work together as separate devices, which, according to various embodiments, may be collectively regarded as an XR device.

[0040] The term "multimedia content" is used herein to refer to communication content observable by a user of a VR device that combines different content forms (such as text, audio, images, animation, video, and / or other elements) into a single interactive presentation, as contrasted with traditional mass media (such as print materials or audio recordings), which are characterized by little or no interaction among users. For example, multimedia content can include videos that contain both audio and video components, audio slideshows, animated videos, and / or other audio and / or video presentations that may include tactile feedback, augmented reality (AR) elements, extended reality (ER) overlays, MR overlays, etc. Multimedia content can be recorded for replay (i.e., rendering) on demand or in real-time (streaming) on a computer, laptop, smartphone, and other computing or electronic devices.

[0041] The term "system-on-a-chip" (SOC) is used herein to refer to a single integrated circuit (IC) chip that contains multiple resources and / or processors integrated on a single substrate. A single SOC can contain circuitry for digital, analog, mixed-signal, and radio-frequency functions. A single SOC can also include any number of general-purpose and / or specialized processors (such as digital signal processors, modem processors, video processors, etc.), memory blocks (such as ROM, RAM, flash memory, etc.), and resources (such as timers, voltage regulators, oscillators, etc.). The SOC can also include software for controlling the integrated resources and processors and for controlling peripheral devices.

[0042] The term "system-in-package" (SIP) can be used herein to refer to a single module or package that contains multiple resources, computing units, cores and / or processors on two or more IC chips, a substrate, or an SOC. For example, a SIP can include a single substrate on which multiple IC chips or semiconductor die are stacked in a vertical configuration. Similarly, a SIP can include one or more multi-chip modules (MCMs) on which multiple ICs or semiconductor die are encapsulated into a unified substrate. A SIP can also include multiple independent SOCs that are coupled together via high-speed communication circuitry and packaged closely together (such as on a single motherboard or within a single computing device). The proximity of the SOCs enables high-speed communication as well as sharing of memory and resources.

[0043] Figure 1AIt is a component block diagram of a multi-source multimedia environment 100 suitable for implementing various embodiments. The multi-source multimedia environment 100 may include a computing device 120 in the form of a smart phone, which is configured to receive input from a user 5, particularly selections associated with multimedia content. The computing device 120 may alternatively be a different form of computing device (such as smart glasses, etc.), or include more than one computing device working together. In the environment 100, the user 5 is equipped with the computing device 120 and has arrived at a venue 10 including a remote media player 140 in the form of a television. For example, the venue 10 may be a bar, restaurant, stadium, airport, event venue, etc. The remote media player 140 is playing (i.e., rendering) a first multimedia content 145 (e.g., a live news stream) and may be configured to stream different multimedia content.

[0044] According to various embodiments, the remote media player 140 may play at least a portion of the first multimedia content 145 by rendering a video component of the first multimedia content 145 on a display. The remote media player 140 may also optionally render an audio component of the first multimedia content 145 from one or more speakers. However, the user 5 may desire to play one of the audio component, video component, or other perceivable media components of the first multimedia content 145 through the computing device 120. For example, although the user 5 may be able to see a display on the remote media player 140 having the corresponding video component of the first multimedia content 145 thereon, the user 5 may not be able to hear the audio component. Thus, desiring to hear the audio component of the first multimedia content 145, according to various embodiments, the user 5 may obtain the audio component and have it rendered by the computing device 120.

[0045] By providing a user input on the computing device 120, the user 5 may initiate a process of obtaining the audio component of the first multimedia content 145. For example, by aiming the camera on the computing device 120 at the remote media player 140 and focusing on the first multimedia content 145, the computing device 120 may receive a user input (e.g., in the form of a sampled image) indicating the content that the user 5 desires to render. The user input may be used to determine the media content selected by the user 5 for rendering that content. In the case where the desired content (e.g., the audio component) is rendered by the computing device 120 synchronously with the rendering of the video component on the remote media player 140, the user 5 may have a more enjoyable experience of observing and consuming the first multimedia content 145.

[0046] Alternatively, although user 5 may be able to hear the sound of the audio component of the first multimedia content 145 emitted by the speakers of the remote media player 140 or even other remote speakers in the venue 10, user 5 may not be able to see its video component (e.g., due to overcrowding in the venue 10 or the direction in which the user is sitting not facing the display). Therefore, desiring to see the video component of the first multimedia content 145, in accordance with various embodiments, user 5 may obtain the video component and have it rendered by the computing device 120. In the case where the video component is rendered by the computing device 120 synchronously with the rendering of the audio component from the remote media player 140, user 5 may have a more enjoyable experience of listening to and observing the first multimedia content 145.

[0047] In some embodiments, although user 5 may be able to see the video component rendered by the remote media player 140 and / or hear the audio component, user 5 may desire to perceive other media components, such as haptic feedback, subtitles, translations, sign language, or other overlays. For example, haptic seats / chairs, clothing, watches, speakers (e.g., subwoofers), etc. may be configured to provide additional perceivable components as part of the source multimedia content. In some embodiments, lights may be configured to flash, dim, or brighten in coordination with the source multimedia content. The user may desire to receive subtitles, translations, sign language, or other overlays locally on the user's computing device 120.

[0048] In some embodiments, user 5 may not wish to hear the audio component or may not wish to hear any louder audio, but may benefit from feeling haptic sensations associated with the content. For example, the user may wish to continue a conversation without adding the source media content to the ambient noise, but expects to feel things like explosions, collisions, roars of a crowd, engines, hitting a ball or kicking a ball with a bat, a tackle, or other similar events that can be expressed with haptic effects (e.g., vibrations or shakes).

[0049] In some embodiments, user 5 may wish to view multiple streams of a live game (i.e., separate video components) and selectively switch between audio streams (i.e., different audio components) to listen as needed.

[0050] The computing device 120 may be configured to receive communications from the local venue computing device 150, such as via the wireless link 132, which may be directly relayed to the local venue computing device 150 via a wireless router 130 having its own wired or wireless link 135. The wireless router 130 may provide wireless local area network (WLAN) capabilities, such as Wi-Fi networks or Bluetooth communications, such as to receive wireless signals from various wireless devices and provide access to the local venue computing device 150 and / or an external network (such as the Internet).

[0051] Alternatively, computing device 120 may be configured to communicate directly with remote media player 140 via wireless link 142 or communicate with local site computing device 150 via remote media player 140 acting in a similar wireless router capacity. As another alternative, computing device 120 may be configured to communicate via long-range wireless communication, such as using cellular communication via cellular network base station 160. In this manner, computing device 120 may also be configured to communicate with remote server 156 via wireless and / or wired connections 162, 164 to network 154, which may include a cellular wireless communication network.

[0052] Remote media player 140 may receive a stream of multimedia content, such as first multimedia content 145, via a wired or wireless link 144 to local site computing device 150. In this manner, local site computing device 150 may control how remote media player 140 renders content, what content to render, and when to render content. In various embodiments, local site computing device 150 may be located within or near site 10 or be remotely located like remote server 156 or a cloud-based system and be accessed via network 154, such as the Internet, via communication link 152.

[0053] Figure 1B is a block diagram of components of another multi-source multimedia environment 101 suitable for implementing various embodiments. Referring Figures 1A to 1B , the illustrated example multi-source multimedia environment 101 may include all of the elements, features, and functions described above with respect to Figure 1A the multi-source multimedia environment in (i.e., 100). Multi-source multimedia environment 101 illustrates an example in which user 5 is using a different computing device 122 in the form of smart glasses. Additionally, multi-source multimedia environment 101 includes a slightly different site 11, which may include a plurality of remote media players 140, 170, 180. Specifically, site 11 includes a first remote media player 140 that renders first multimedia content 145, a second remote media player 170 that renders second multimedia content 175, and a third remote media player 180 that renders third multimedia content 185.

[0054] According to various embodiments, user 5 may desire to play, via computing device 122, one of an audio component, a video component, or other perceivable media component of a first multimedia content 145, a second multimedia content 175, or a third multimedia content 185. For example, although user 5 may be able to see a display having a corresponding video component of the second multimedia content 175 thereon on the second remote media player 170, user 5 may not be able to hear the audio component. In fact, the first media player 140, the second media player 170, and the third media player 180 may not render audio to avoid generating too much noise and / or interfering with each other. Thus, desiring to hear the audio component of the second multimedia content 175, according to various embodiments, user 5 may obtain the audio component and have it rendered by computing device 122. In the case where the audio component is rendered by computing device 122 synchronously with the rendering of the video component on the second remote media player 170, user 5 may have a more enjoyable experience of viewing and consuming the second multimedia content 175.

[0055] By providing user input on computing device 122, user 5 may initiate the process of obtaining the audio component of the second multimedia content 175. For example, by pointing a finger at the second remote media player 170 and specifically in the direction 176 towards the second multimedia content 175, computing device 122 may recognize this gesture indicating the content that user 5 desires to render (e.g., using gesture recognition from the camera imaging 124 of the smart glasses). The user input may be used to determine the media content selected by user 5 for rendering the content. In the case where the desired content (e.g., the audio component) is rendered by computing device 122 synchronously with the rendering of the video component on the second remote media player 170, user 5 may have a more enjoyable experience of viewing and consuming the second multimedia content 175.

[0056] Alternatively, user 5 may be able to hear the sound of the audio component of the second multimedia content 175 (e.g., emitted by a nearby speaker), but user 5 may not be able to see its video component (e.g., due to crowding in venue 11 or the direction in which the user is sitting not facing the display). Thus, desiring to see the video component of the second multimedia content 175, according to various embodiments, user 5 may obtain the video component and have it rendered by computing device 122. In the case where the video component is rendered by computing device 122 synchronously with the rendering of the audio component from the second remote media player 170, user 5 may have a more enjoyable experience of listening to and viewing the second multimedia content 175.

[0057] Figure 1C is a component block diagram of another multi-source multimedia environment 102 suitable for implementing various embodiments. Referring to Figures 1A to 1C , the illustrated example multi-source multimedia environment 102 may include the above regardingFigure 1A and Figure 1B all of the components, features, and functions described in the multi-source multimedia environments (i.e., 100, 101). The multi-source multimedia environment 102 illustrates an example in which the user 5 uses again the computing device 120 in the form of a smart phone. Additionally, the multi-source multimedia environment 102 includes a slightly different venue 12, which may include a first remote media player 140, but now displays multiple multimedia contents 145, 175, 185, 195. Specifically, the first remote media player 140 is rendering the first multimedia content 145, the second multimedia content 175, and the third multimedia content 185 as well as the fourth multimedia content 195.

[0058] According to various embodiments, the user 5 may wish to play, via the computing device 120, one of the audio component, video component, or other perceivable media component of one of the first multimedia content 145, the second multimedia content 175, the third multimedia content 185, or the fourth multimedia content 195. Thus, desiring to hear the audio component of the fourth multimedia content 195, according to various embodiments, the user 5 may obtain the audio component and have it rendered by the computing device 120. In the case where the audio component is rendered by the computing device 120 synchronously with the rendering of the video component on the first remote media player 140, the user 5 may have a more enjoyable experience of observing and consuming the fourth multimedia content 195.

[0059] Alternatively, the user may be able to hear the sound of the audio component of the fourth multimedia content 195 (e.g., emitted by a nearby speaker), but the user 5 may not be able to see its video component. Thus, desiring to see the video component of the fourth multimedia content 195, according to various embodiments, the user 5 may obtain the video component and have it rendered by the computing device 120. In the case where the video component is rendered by the computing device 120 synchronously with the rendering of the audio component from the first remote media player 140, the user 5 may have a more enjoyable experience of listening to and observing the fourth multimedia content 195.

[0060] Figure 1D is a component block diagram of another multi-source multimedia environment 103 suitable for implementing various embodiments. Referring to Figures 1A to 1D , the illustrated example multi-source multimedia environment 103 may include the above regarding Figures 1A to 1CAll of the components, features, and functions described in the multi-source multimedia environment (i.e., 100, 101, 102). The multi-source multimedia environment 103 illustrates an example in which the user 5 is now using a computing device 121 in the form of a spreadsheet device. In addition, the multi-source multimedia environment 103 includes a slightly different venue 13, which may include a fourth remote media player 190 in the form of a speaker. In this way, the fourth remote media player 190 only renders the audio component of the fifth multimedia content 191.

[0061] According to various embodiments, the user 5 may wish to display the video component from the fifth multimedia content 191 via the computing device 121. Therefore, desiring to view the video component of the fifth multimedia content 191, the user 5 may obtain the video component and have it rendered by the computing device 121. In the case where the video component is rendered by the computing device 121 synchronously with the rendering of the audio component on the fourth remote media player 190, the user 5 may have a more enjoyable experience of observing and consuming the fifth multimedia content 191.

[0062] Figures 2A to 2C is a schematic diagram of an exemplary user input technique applicable to various embodiments. Referring to Figures 1A to 2C , the illustrated user input technique may enable the user to indicate that the user wishes to render a selected audio component, video component, or other perceivable media component on the computing device.

[0063] Figure 2A is a schematic diagram of a gesture-based user input technique 200 applicable to implementing various embodiments. Referring to Figures 1A to 2A , the illustrated exemplary gesture-based user input technique 200 may include using a gesture detection system of a computing device (e.g., 122), which can capture images and recognize gestures made by the user 5. Figure 2A is illustrated from a viewpoint perspective, which shows what the imaging system of the computing device can visually capture. Using computer vision and / or other computerized vision recognition systems, the computing device may be able to detect when the user performs the recognized gesture. For example, a gesture detection algorithm run by the processor of the computing device may be configured to detect a pointing gesture 25. The gesture detection algorithm may be configured to detect various different gestures that can trigger different operations. In this way, when the user 5 configures his or her hand in a specific manner (e.g., pointing a finger in a certain direction) or moves the hand and / or arm in a specific manner, such gestures may be recognizable when they meet certain predetermined characteristics.

[0064] In addition, the gesture detection algorithm can be configured to analyze and interpret the detected gesture to determine whether the characteristics of the gesture can provide additional information associated with the user input. For example, when a pointing gesture is detected, the detected gesture can be analyzed to determine the direction of the pointing and / or identify what the user is pointing at. In this example, interpreting the pointing gesture can determine that the user is pointing at the first multimedia content 145 being rendered on the first media player 140. Alternatively, instead of a pointing gesture (which can be rude or clumsy), the user can squint his / her eyes (which is sometimes a natural reaction when trying to see something better), pucker his / her lips (e.g., towards the source the user is interested in), quickly lift his / her head, keep his / her head up, turn his / her head to one side, cup his / her ears, etc. Larger gestures can indicate that the identified source is farther away.

[0065] Some embodiments can use a range sensor to determine how far an object is from the user / computing device to more easily determine what is being pointed at. A distance threshold can be used to exclude objects that are too far away as targets of the pointing gesture. In this way, the gesture detection system can exclude objects that are too far away in the background by establishing a threshold distance from the computing device. The threshold distance can be equal to or shorter than the distance at which the user can typically see and / or read the display. Accordingly, in response to determining that the identified object (and specifically, the identified multimedia content) is within the threshold distance of the computing device, the processor of the computing device interprets the pointing gesture as a selection of the identified multimedia content. The computing device can provide feedback to the user (such as visual, tactile, and / or audible output) to let the user know that the multimedia content has been identified.

[0066] In some embodiments, the user can provide supplementary gestures as further user input. Once a particular multimedia content has been identified, additional gestures of the user can indicate whether the user wishes to render the audio component, video component, or other perceivable media component of the identified multimedia content on the computing device. For example, swiping the pointing finger to the left can indicate that the user wants the audio component, while swiping the pointing finger to the right can indicate that the user wants the video component. Different additional gestures can mean a combination of those things or something completely different. Thus, interpreting the additional gesture can provide user input that enables the computing device to identify the audio component, video component, or one of the other perceivable media components of the identified multimedia content that the user wishes to render on the computing device.

[0067] Figure 2B is a schematic diagram of a gaze-based user input technique 201 suitable for implementing various embodiments. Refer to Figures 1A to 2B , the illustrated example of the gaze-based user input technique 201 can include an eye / gaze direction detection system of the computing device 122, which can perform eye tracking to determine what the user 5 is looking at.Figure 2B is illustrated from a viewpoint perspective, showing what the imaging system of a computing device can visually capture. In combination with eye tracking using computer vision and / or other computerized vision recognition systems, a computing device may be able to detect the object on which the user's focus is. For example, a gaze detection algorithm run by a processor of the computing device may be configured to detect the focus and / or direction of the user's gaze. The gaze detection algorithm may also be configured to detect the object or element being viewed. Thus, combining the identified objects or elements in the direction the user is viewing may allow the computing device to identify the multimedia content 145 being rendered on the display of the media player 140 in the direction of the user's gaze.

[0068] Figure 2C is a schematic diagram of a screen-based user input technique 202 applicable to implementing various embodiments. Refer to Figures 1A to 2C , the illustrated example of the screen-based user input technique 202 may use the characteristics of the display of the computing device 120 to determine what multimedia content the user (e.g., 5) wants. Figure 2C is illustrated from a viewpoint perspective, showing what the user sees. Specifically, the user views the computing device 120 in the foreground and the first media player 140 in the background. More specifically, the user points the camera of the computing device 120 at the first media player 140 such that the first media player 140 and its display are visible on the display of the computing device 120. An application running on the computing device 120 may determine the direction from which the user is perceiving the desired multimedia content based on the direction the camera is facing. Additionally, the desired multimedia content may be identified by the content that appears on the screen of the computing device 120 (or more specifically, in the target area 141 on the screen of the computing device 120). By keeping the desired multimedia content 185 in the target area 141 until the application registers the input, the user may provide a user input for selecting the desired multimedia content 185. Additional prompts may be provided for specifying which of the audio component or video component associated with the multimedia content the user wishes to have rendered on the computing device.

[0069] Figure 2D is a schematic diagram of another screen-based user input technique 203 applicable to implementing various embodiments. Refer to Figures 1A to 2D , the illustrated example of the screen-based user input technique 203 may use the characteristics of the touchscreen display of the computing device 120 to determine what multimedia content the user (e.g., 5) wants. Figure 2DIt is illustrated from the perspective of the viewpoint, showing what the user sees. Specifically, the user views the computing device 120 in the foreground and the first media player 140 in the background. More specifically, the user points the camera of the computing device 120 at the first media player 140 such that the first media player 140 and its display are visible on the display of the computing device 120. An application running on the computing device 120 can determine the direction from which the user is perceiving the desired multimedia content based on the direction the camera is facing. Additionally, the desired multimedia content can be identified by additional user input, such as a tap on the screen of the computing device 120 at a portion corresponding to the desired multimedia content 185. Additional cues can be provided for specifying which of the audio component or video component associated with the multimedia content the user wishes to render on the computing device 120.

[0070] Figure 2E is a schematic diagram of another XR overlay-based user input technique 204 applicable to implementing various embodiments. Refer to Figures 1A to 2E , the illustrated example of the XR overlay-based user input technique 204 can use the characteristics of the field-of-view overlay that can be rendered by a computing device (e.g., 122) in the form of smart glasses to determine what multimedia content the user 5 wants. Figure 2E It is illustrated from the perspective of the viewpoint, showing what the user sees. Specifically, the user 5 views through the computing device 122 and sees the first media player 140 in the background. According to various embodiments, the computing device 120 can project overlays 1, 2, 3, 4 onto the user's field of view to add labels to each of the first multimedia content 145, second multimedia content 175, third multimedia content 185, and fourth multimedia content 195. The projected overlays 1, 2, 3, 4 can appear to the user 5 as being on top of the first multimedia content 145, second multimedia content 175, third multimedia content 185, and fourth multimedia content 195. To select one of the projected overlays 1, 2, 3, 4, the user 5 can touch areas 245, 275, 285, 295 in the air, which appear to the user as corresponding to where the user would cover the corresponding overlays 1, 2, 3, 4 on the first media player 140.

[0071] An application running on the computing device (e.g., 122) can determine which of the overlays 1, 2, 3, 4 the user 5 has selected. In this way, the desired multimedia content can be identified by the user performing a virtual interaction with the screen of the first multimedia device 140. Additional user input, such as a swipe gesture in a specific direction, can specify which of the audio component or video component associated with the multimedia content the user wishes to render on the computing device 122.

[0072] Various embodiments may use alternative or additional user input techniques. For example, a computing device may be configured to receive verbal input from a user 5 (e.g., using speech recognition). As another example, a computing device may project a hologram or marker onto a selection within the user's field of view, which may be used to enter and / or confirm user input. Similarly, a computing device may present a list from which a user may select to provide user input. The list is obtained by the computing device from local and / or remote computing devices. Information used to populate such a list may be obtained actively, passively, or after a triggering event by the computing device, such as in response to a user request to do so.

[0073] Figure 3 FIG. 4 is a block diagram of components illustrating a non-limiting example of a computing and wireless modem system 300 in a computing device suitable for implementing any of the various embodiments, including within the computing device. The various embodiments may be implemented on several single-processor and multi-processor computer systems, including system-on-a-chip (SOC) or system-in-package (SIP).

[0074] Referring to FIGS. 1 through Figure 3 , the illustrated example computing system 300 (which may be a SIP in some embodiments) includes an SOC 302 coupled to a clock 306, a voltage regulator 308, a radio module 366 configured to transmit and receive wireless communications (including Bluetooth (BT) and Bluetooth Low Energy (BLE) messages) via an antenna (not shown) and an inertial measurement unit (IMU) 368. When the computing system 300 is used in a computing device, the radio module 366 may be configured to broadcast BLE and / or Wi-Fi beacons. In some implementations, the SOC 302 may operate as a central processing unit (CPU) of a user mobile device, which executes instructions by performing arithmetic, logical, control, and input / output (I / O) operations specified by the instructions of software applications.

[0075] The SOC 302 may include a digital signal processor (DSP) 310, a modem processor 312, a graphics processor 314, an application processor 316, one or more coprocessors 318 (such as a vector coprocessor) connected to one or more of these processors, a memory 320, custom circuitry 322, system components and resources 324, an interconnect / bus module 326, one or more temperature sensors 330, a thermal management unit 332, and a thermal power envelope (TPE) component 334. A second SOC (not shown) may include other elements, such as a 5G modem processor, a power management unit, an interconnect / bus module, multiple millimeter wave transceivers, additional memory, and various additional processors, such as an application processor, a packet processor, etc.

[0076] Each of the processors 310, 312, 314, 316, 318 may include one or more cores, and each processor / core may perform operations independently of other processors / cores. For example, the SOC 302 may include a processor that executes a first type of operating system (such as FreeBSD, LINUX, OS X, etc.) and a processor that executes a second type of operating system (such as MICROSOFT WINDOWS 10). In addition, any one or all of the processors 310, 312, 314, 316, 318 may be included as part of a processor cluster architecture (such as a synchronous processor cluster architecture, an asynchronous or heterogeneous processor cluster architecture, etc.).

[0077] The SOC 302 may include various system components, resources, and custom circuits for managing sensor data, analog-to-digital conversion, wireless data transmission, and for performing other specialized operations, such as decoding data packets and processing encoded audio and video signals for rendering in a web browser. For example, the system components and resources 324 of the SOC 302 may include power amplifiers, voltage regulators, oscillators, phase-locked loops, peripheral bridges, data controllers, memory controllers, system controllers, access ports, timers, and other similar components for supporting processors and software clients running on a user mobile device. The system components and resources 324 or the custom circuit 322 may also include circuits for docking with peripheral devices (such as cameras, electronic displays, wireless communication devices, external memory chips, etc.).

[0078] The SOC 302 may communicate via an interconnect / bus module 326. The various processors 310, 312, 314, 316, 318 may be interconnected via the interconnect / bus module 326 to one or more memory elements 320, system components and resources 324, and custom circuit 322, as well as a thermal management unit 332. The interconnect / bus module 326 may include an array of reconfigurable logic gates or implement a bus architecture (such as CoreConnect, AMBA, etc.). The communication may be provided by a high-performance interconnect, such as a high-performance on-chip network (NoC).

[0079] The SOC 302 may also include an input / output module (not illustrated) for communicating with resources external to the SOC (such as a clock 306 and a voltage regulator 308). Resources external to the SOC (such as the clock 306, voltage regulator 308) may be shared by two or more internal SOC processors / cores in the internal SOC processors / cores.

[0080] Figure 4A is a communication flowchart illustrating an example method 400 for multi-source multimedia output and synchronization in a computing device. Referring to FIGS. 1 to Figure 4A, Method 400 illustrates an example of a multi-source multimedia output and synchronization scenario involving gaze-based user input combined with metadata retrieval, according to various embodiments.

[0081] Method 400 may be initiated in response to a media player (e.g., remote media player 140) rendering multimedia content (such as first multimedia content 145). A content provider located outdoors may deliver multimedia content to the venue 11 via the Internet or other communication network 154. For example, a cable television provider or an Internet service provider may supply a stream 410 of multimedia content to a local computing device 150, such as a cable television box or a router, which may then deliver the received stream 412 to a media player that is remote from the user 5 (i.e., separated from the user). The received stream 412 may be rendered as first multimedia content 145 on the media player.

[0082] In various embodiments, user 5 wearing a wearable computing device 122 and looking at first multimedia content 145 may initiate multi-source multimedia output and synchronization via a user input to the computing device 122 that is designed to initiate the process. For example, using gesture-based commands, user 5 may initiate the process by performing a predetermined gesture, such as a pointing gesture.

[0083] In response to user 5 looking at first multimedia content 145, a camera or other sensor of computing device 122 may capture an image 414 that includes first multimedia content 145 from the media player. The processor of computing device 122 may scan the captured image 414 and register the detection of first multimedia content 145. Additionally, in response to user 5 pointing at something (i.e., performing a gesture), a camera or other sensor of computing device 122 may capture an additional image 416 of the user performing the gesture. The processor of computing device 122 may scan the additional captured image 416 and register that the user has performed a multimedia selection gesture. A combination of performing a predetermined gesture (e.g., a pointing gesture) in the direction of the registered multimedia content 145 may initiate the process of multi-source multimedia output and synchronization. The user's gesture may be considered a received user input that selects one of the audio component or the video component associated with the multimedia content being rendered by a remote media player within the perceivable distance of the user of computing device 122.

[0084] In response to receiving user input in the form of an additional image 416 that includes a launch gesture, a processor of the computing device 122 may attempt to identify the multimedia content 145 selected by the user. To do so, the processor of the computing device 122 may send a query to the computing device. For example, the processor may use a radio module to send a local query 420 to a local computing device 150 (such as a venue computer that houses a database). The local query 420 may be received by a router 130 in the venue 10. The router 130 may be configured to provide access to the local computing device 150 by passing the local query as a secure communication 422 to the local computing device 150.

[0085] In response to receiving the secure communication 422, the local computing device 150 may perform a database lookup to identify the selected multimedia content. For example, the local computing device 150 may identify a first media content 145 as the selected multimedia content. The local computing device 150 may thus send an intermediate response 430 to the router 130, and the router 130 may send the intermediate response as a query response 432 to the computing device 122. The query response 432 may include metadata that specifically identifies the first multimedia content 145. Alternatively, the query response 432 may include a link for obtaining identification information from a remote database (e.g., 156).

[0086] Using the obtained metadata, the computing device 122 may send a request to obtain the multimedia content from the source of the multimedia content. For example, the metadata may not only identify the first multimedia content 145, but may also indicate that the local computing device 150 may supply an audio component, a video component, or other perceivable media components from the source of the identified multimedia content. In this way, the computing device may send a local request 440 for the multimedia content to the local computing device 150 (via the router 130).

[0087] Alternatively, the metadata may not indicate how to obtain the multimedia content, or may indicate that it must be obtained from a remote server (e.g., 156), such as from a content service provider. In this case, the computing device may send a remote request 442 for the multimedia content to a remote computing device via the communication network 154.

[0088] Multimedia content displayed at some commercial venues (e.g., 10) may be restricted by a paywall, which can inhibit the ability of a computing device to locate or obtain the desired multimedia content on the display at that venue. Thus, according to various embodiments, a commercial venue may at least temporarily extend its subscription (i.e., license) to computing devices having access to the local area network of the venue, such that a user at the commercial venue can pass through the paywall and obtain more detailed information about the multimedia stream (e.g., via Wi-Fi or BTE access) and / or even obtain the multimedia content itself using the extended subscription. Some venues may offer this as an automatic guest pass or, optionally, offer the extended subscription service at the cost or in the manner of providing the subscription. Thus, the query response 432 from the router 130 may include subscription access to the multimedia content from the source of the media content. In this way, the computing device 122 may later receive the identified audio component, video component, or other perceivable media component of the multimedia content based on the subscription access.

[0089] If available, the local computing device 150 may respond to the request 440 by establishing a connection 450 with the computing device 122 (e.g., via the router 130), the connection being configured to provide a data stream to the computing device for delivering the requested first multimedia content 145. Although the user 5 may have selected only one of the audio or video components for rendering by the computing device 122, both the audio and video components and any XR enhancements of the requested first multimedia content 145 may be obtained by the computing device 122 (i.e., received from the local computing device 150).

[0090] Additional content components not selected by the user 5 for rendering on the computing device 122 may be used by the computing device 122 to effectively render the selected one of the audio component, video component, or other perceivable media component from the obtained multimedia content. For example, although the computing device 122 may only need to render the audio component of the multimedia content, the obtained video component may be used by the computing device 122 to properly synchronize the delivery of the audio component. Similarly, the computing device 122 may use any XR enhancements (such as overlays) as part of the rendering of the selected multimedia content. Thus, since the media player (e.g., 140) may continue to receive the first multimedia content 145, both the media player and the computing device 122 may receive both the audio component and the video component as part of the obtained multimedia content.

[0091] Alternatively, if the computing device must obtain the multimedia content from a remote server (e.g., 156) that received the request 442, the remote server may respond by establishing a connection 452 configured to provide a data stream to the computing device 122 for delivering the requested first multimedia content 145. The delivery of the requested first multimedia content 450 from the remote server to the computing device 122 may be similar to the delivery from the local computing device 150, although access through the router 130 may not be necessary (e.g., when using cellular communication services). As described above, both the media player and the computing device 122 may receive both the audio component and the video component as part of the obtained multimedia content.

[0092] Once the computing device 122 begins receiving the first multimedia content 145, the computing device 122 may begin to render 460 a selected one of the audio component, the video component, or other perceivable media components from the obtained multimedia content. When rendering the obtained multimedia content, the computing device 122 may synchronize the rendering with the rendering by the media player of the streams 470, 472 from the source.

[0093] To synchronize the rendering of a selected one of the audio component, the video component, or other perceivable media components from the obtained multimedia content, the computing device 122 may ensure that the delivery timing of the selected component matches the delivery timing of the other components rendered by the remote media player. For synchronization, the computing device 122 may need to accelerate or slow down the output of the requested audio or video component in order to match the timing of the multimedia stream from the remote media player.

[0094] The processing delay associated with video rendering may be used to synchronize the audio rendered by the computing device 122. In normal video rendering, the audio component and the video component may arrive together, but the video component may be delayed as part of the rendering process. In various embodiments, the computing device 122 may utilize the delay in the rendering of the video component to buffer the audio component and use the buffered audio to synchronize based on additional video image sampling completed as part of the synchronization process. In other words, since the processing time of the video stream is significantly longer than the processing time of the audio stream, the received audio stream may be buffered in order to synchronize the audio output from the remote media player with the video output.

[0095] As part of video sampling for a synchronization process, computing device 122 may capture an image observable on a media player and associated with a particular point in an associated audio stream at time X. From this point, as long as the collective processing delay T of the media player and / or computing device 122 is known, this collective processing delay can be used to determine the timing (X + T) of the output of the synchronized audio stream. This processing delay T can be 10 - 100 milliseconds, but can be sufficient time to compute and output the received audio stream in synchronization with multimedia streaming on the media player. Even live multimedia is not technically live; it is delayed by 10 - 100 milliseconds, and this delay can be used to time the synchronization. Various embodiments may employ other known synchronization techniques.

[0096] A synchronization system receives data regarding a multimedia stream (either shared locally with a streaming device or obtained from a remote source), which includes an audio component and a video component. Thereafter, the audio from the multimedia stream can be synchronized by matching the video observed on the media player with the video in the multimedia stream.

[0097] By continuing to observe the multimedia content stream to synchronize sound (such as by lip - reading), the synchronization of sound and imaging can be continuous. This can achieve fine - grained synchronization or refine an existing synchronization. For non - live broadcasts, audio events within the multimedia content can be known and used to synchronize the replay (e.g., knowing when a particular sound is expected to occur during multimedia replay). For example, a video sequence including applause or another obnoxious screech associated with a visual event can be used for synchronization.

[0098] Interruptions (e.g., commercials) or changes in a multimedia broadcast may result in a new multimedia streaming detection event, which restarts the process from the beginning. However, since commercials are typically not live streams, their content can be easily obtained in advance.

[0099] Figure 4B is a communication flow diagram illustrating an example method 401 for multi - source multimedia output and synchronization in a computing device. Referring to FIGS. 1 to Figure 4B , method 401 illustrates an example of a multi - source multimedia output and synchronization scenario involving gesture - based user input combined with remote network lookup to identify and obtain multimedia content according to various embodiments. Example method 401 can provide multi - source multimedia output and synchronization without support from a local venue or network. Without the need for support from a local venue, a user of a computing device can use the multi - source multimedia output and synchronization techniques of various embodiments in almost any venue.

[0100] Method 401 may be initiated in response to a media player (e.g., remote media player 140) rendering multimedia content (such as first multimedia content 145). A content provider located outdoors with a remote server 156 may deliver multimedia content to a venue 14 via the Internet or other communication network 154. For example, the content provider may supply a stream 411 of multimedia content via the communication network 154 to a media player (e.g., 140), which in turn may convey the received stream 413 to a media player that is remote from the user 5 (i.e., separated from the user). The received stream 413 may be rendered as first multimedia content 145 on the media player.

[0101] As described above with respect to method 400, user 500 may initiate multi-source multimedia output and synchronization. In response to receiving user input initiating the process, a processor of computing device 120 may attempt to identify the multimedia content 145 selected by the user. To do so, since there is no local database to query, the processor of computing device 122 may send a query to a remote computing device. For example, the processor may use a radio module to send a remote query 421 to a remote computing device 156 (such as a multimedia database).

[0102] The remote query 421 may be received by the communication network 154 and passed as a network query 423 to the remote computing device 156. The remote query 421 may request the identification of the selected multimedia content for its identification. Alternatively, if the identification of the multimedia content is known in some way, the remote query 421 may request information about the source of the multimedia content. The remote query 421 may include a sampling of the multimedia content, such as a short video or screenshot of the first multimedia content. Such sampling may be referred to as "fingerprinting" because the collected images are used to identify the content.

[0103] Multimedia content fingerprinting may take small samples of a multimedia stream to match against a database of multimedia, thereby identifying what specific multimedia content was captured and at what point within it is being viewed. The database being searched may be local to the venue presenting the multimedia in question, or a remote database or collection of databases. The remote database or collection of databases may form a repository of information about all or most multimedia, thus providing multimedia lookup. Even if not all multimedia content is available for lookup, such a service can be useful to users if there is enough multimedia content available for lookup. Fingerprinting may be continuous, at regular intervals, at intervals initiated by the media player (pushed) or the computing device (e.g., pulled from user input or a process on the computing device).

[0104] The search for multimedia content can also use additional sensor information. For example, sensor data from computing device 120 (e.g., location, date / time, orientation) can be used to assist in the identification of multimedia content. The remote server can determine the venue based on such sensor data, which can narrow down or precisely identify what multimedia content is being streamed there. Such location, orientation, and time information can identify the institution, the location / orientation within the institution (e.g., near one or more multimedia displays) to assist in identification selection.

[0105] The remote query 421 can be received by the communication network 154 and passed as a network query 423 to the remote computing device 156. If available, the remote server 156 can respond by establishing a connection with the computing device 120 through a series of communications 431, 433, 441, 443 between the remote server 421, the network 154, and the mobile device 120, where the connection is configured to provide a data stream to the computing device 120 for delivering the requested first multimedia content 145. In some embodiments, the remote server 156 can transmit requests 431, 433 for code or license information for the mobile device 120, which indicates that the user has paid or otherwise received permission or access to receive the multimedia content, and the mobile device 120 can reply with the requested code or license information in messages 441, 443. Once the connection is established, the remote server 154 can begin delivering the requested first multimedia content to the computing device 120 via streaming communications 451, 453, similar to delivery from a local computing device (e.g., 150), although access through a local router (e.g., 130) may not be necessary (e.g., when using cellular communication services). As described above, both the media player and the computing device 120 can receive both the audio component and the video component as part of the acquired multimedia content.

[0106] Once the computing device 120 begins receiving the first multimedia content 145, the computing device 120 can begin rendering 460 a selected one of the audio component, the video component, or other perceivable media components from the acquired multimedia content. When rendering the acquired multimedia content, the computing device 120 can synchronize the rendering with the rendering by the media player of the streams 471, 472 from the source.

[0107] Figure 5Ais a process flow diagram illustrating method 500 for multi-source multimedia output and synchronization in a computing device. Referring to FIGS. 1-5, the components for performing each operation of method 500 may be performed by a processor (e.g., 302, 310, 312, 314, 316, and / or 318) and / or transceiver (e.g., 366) of a computing device (e.g., 120, 122), etc. Alternatively, the components for performing each of the operations of method 500 may be a processor of a computing device, a computing device associated with a local site, or other computing devices working in combination (e.g., a remote computing device such as remote server 156).

[0108] In block 510, the computing device may receive a user input that selects one of an audio component or a video component associated with user-identified multimedia content being rendered by a remote media player within a perceivable distance of the user of the computing device. The user input may indicate that the user desires to render the selected audio component, video component, or other perceivable media component on the computing device.

[0109] In some embodiments, the user may select multimedia content from a plurality of multimedia content observable by the user, wherein the selection is communicated to the computing device via various methods.

[0110] In some embodiments, receiving a user input by the computing device may include detecting a gesture performed by the user. The detected gesture may be interpreted by the computing device to determine whether it identifies multimedia content being rendered within a threshold distance of the computing device. By interpreting the detected gesture, the computing device may identify one of an audio component, a video component, or other perceivable media component of the identified multimedia content that the user desires to render on the computing device.

[0111] In some embodiments, receiving a user input by the computing device may include receiving a camera image of multimedia when the user points a camera of the computing device or a connected mobile device at the remote media player. The computing device may use the received camera image to identify multimedia content being rendered within the visual range of the computing device.

[0112] In block 512, a processor of the computing device may identify the multimedia content. In some embodiments, identifying multimedia content being rendered on a display within a perceivable distance of the user of the computing device may include detecting the user's gaze direction. Additionally, the multimedia content being rendered on the display in the user's gaze direction may be identified. In some embodiments, identifying multimedia content being rendered on a display within a perceivable distance of the user of the computing device may include receiving a user input indicating the direction from which the user is perceiving the multimedia content. Additionally, the multimedia content may be identified based on the received user input.

[0113] In block 514, a processor of a computing device may obtain the identified multimedia content from a source of the multimedia content. The computing device receiving the multimedia content may mean that each of the remote media player and the computing device receives both an audio component and a video component as part of the obtained multimedia content.

[0114] In some embodiments, obtaining the identified audio component, video component, or other perceivable media component of the multimedia content from a source of the multimedia content may include obtaining metadata about the multimedia content. Moreover, the obtained metadata may be used to identify the source of the multimedia content. Additionally, the audio component, video component, or other perceivable media component may be obtained from the identified source of the multimedia content.

[0115] In some embodiments, obtaining the identified audio component, video component, or other perceivable media component of the multimedia content from a source of the multimedia content may include sending a query about the multimedia content to a remote computing device. Additionally, an identification of the source of the multimedia content may be requested. Additionally, the audio component, video component, or other perceivable media component may be obtained from the identified source of the multimedia content.

[0116] In some embodiments, obtaining the identified audio component, video component, or other perceivable media component of the multimedia content from a source of the multimedia content may include obtaining a subscription access to the multimedia content from the source of the media content. Additionally, the identified audio component, video component, or other perceivable media component of the multimedia content may be received based on the subscription access.

[0117] In block 516, a processor of the computing device may render a selected one of the audio component, video component, or other perceivable media component from the obtained multimedia content. Such rendering by the computing device may be synchronized with rendering performed by a remote media player within a perceivable distance of the user.

[0118] In some embodiments, rendering, by a computing device, a selected one of an audio component, a video component, or other perceivable media components synchronously with rendering by a remote media player within a perceivable distance of a user may include sampling one or more of an audio component, a video component, or other perceivable media components of multimedia content being rendered by the remote media player within the perceivable distance of the user. Moreover, a timing difference may be determined between samples of one or more of an audio component, a video component, or other perceivable media components of the multimedia content being rendered and an audio component, a video component, or other perceivable media components obtained from a source of the multimedia content. Additionally, a selected one of an audio component, a video component, or other perceivable media components may be rendered by the computing device such that the user will perceive the selected one of the audio component, the video component, or other perceivable media components so rendered as being synchronous with the perceivable multimedia content rendered by the remote media player.

[0119] Figure 5B Illustrated are additional operations 501 executable by a processor of a computing device for outputting and synchronizing multi-source multimedia content in addition to method 500. In block 518, the processor may sample one of an audio component or a video component being rendered by a display, where the query sent includes at least a portion of the sampled audio component or video component.

[0120] In some embodiments, at least one of an identification of multimedia content or synchronization with rendering by a remote media player may be based on information received in response to a query sent.

[0121] Figure 6 is a component block diagram of a user mobile device 120 adapted to function as a user mobile device or consumer user equipment (UE) when configured with processor-executable instructions to perform operations of various embodiments. Referring to FIGS. 1 through Figure 6 , the user mobile device 120 may include a SOC 302 (e.g., SOC-CPU) coupled to a second SOC 604 (e.g., a SOC with 5G capabilities). The first SOC 302 and the second SOC 604 may be coupled to internal memories 606, 616, a display 612, and coupled to a speaker 614. Additionally, the user mobile device 120 may include an antenna 624 for transmitting and receiving electromagnetic radiation, which may be connected to a radio module 366 configured to support a wireless local area network data link (e.g., BLE, Wi-Fi, etc.) and / or coupled to a wireless wide area network (e.g., a cellular telephone network) of one or more processors in the first SOC 302 and / or the second SOC 604. The user mobile device 120 generally also includes menu selection buttons 620 for receiving user input.

[0122] The exemplary user mobile device 120 may also include an inertial measurement unit (IMU) 368 that includes a plurality of microelectromechanical sensor (MEMS) elements configured to sense movement associated with acceleration and rotation of the device and provide this movement information to the SOC 302. Additionally, one or more of the processors in the first SOC 302 and the second SOC 604, and the radio module 366 may include digital signal processor (DSP) circuitry (not shown separately).

[0123] Various embodiments (including those discussed above with reference Figure 1A to FIGS. 11I) may be implemented on a variety of wearable devices, an example of which is illustrated Figure 7 in the form of smart glasses 700. Referring Figures 1A to 7 to, the smart glasses 700 may operate like conventional glasses but have enhanced computer features and sensors such as a built-in camera 735 and a heads-up display or AR features on or near the lenses 731. Like any glasses, the smart glasses may include a frame 702 coupled to temple arms 704 that fit alongside the wearer's head and behind the ears. When the nose pads 706 on the bridge 708 are placed on the wearer's nose, the frame 702 holds the lenses 731 in place in front of the wearer's eyes.

[0124] In some embodiments, the smart glasses 700 may include an image rendering device 714 (e.g., an image projector) that may be embedded in one or both of the temple arms 704 of the frame 702 and configured to project an image onto the optical lens 731. In some embodiments, the image rendering device 714 may include a light-emitting diode (LED) module, a light tunnel, a homogenizing lens, an optical display, a folding mirror, or other well-known projector or head-mounted display components. In some embodiments (e.g., those in which the image rendering device 714 is not included or used), the optical lens 731 may be or may include a transparent or partially transparent electronic display. In some embodiments, the optical lens 731 includes image generating elements such as a transparent organic light-emitting diode (OLED) display element or a liquid crystal on silicon (LCOS) display element. In some embodiments, the optical lens 731 may include separate left and right eye display elements. In some embodiments, the optical lens 731 may include or operate as an optical waveguide for delivering light from the display element to the wearer's eyes.

[0125] The smart glasses 710 may include a plurality of external sensors configured to obtain information regarding the wearer's movements and external conditions useful for sensing images, sounds, muscle movements, and other phenomena useful for detecting when the wearer interacts with the described virtual user interface. In some embodiments, the smart glasses 700 may include a camera 735 configured to image objects in front of the wearer as a still image or video stream, which may be sent to another computing device (e.g., mobile device 120) for analysis. Additionally, the smart glasses 700 may include a lidar sensor 740 or other ranging device. In some embodiments, the smart glasses 700 may include a microphone 710 positioned and configured to record sounds near the wearer. In some embodiments, multiple microphones may be positioned at different locations on the frame 702, such as at the distal end of the temple 704 near the jaw, to record sounds made when the user taps on a selection object on the hand, etc. In some embodiments, the smart glasses 700 may include a pressure sensor (such as on the nose pad 706) configured to sense facial movement to calibrate distance measurements. In some embodiments, the smart glasses 700 may include other sensors (e.g., thermometer, heart rate monitor, body temperature sensor, pulse oximeter, etc.) for collecting information related to the environment and / or user conditions useful for identifying the wearer's interaction with the virtual user interface.

[0126] The processing system 712 may include processing and communication SOCs 902, 904, which may include one or more processors 902, 904, and one or more of the processors may be configured with processor-executable instructions to perform the operations of the various embodiments. The processing and communication SOCs 902, 904 may be coupled to internal sensors 720, internal memory 722, and communication circuitry 724, which is coupled to one or more antennas 726 for establishing a wireless data link with an external computing device (e.g., mobile device 120), such as via a Bluetooth link. The processing and communication SOCs 902, 904 may also be coupled to a sensor interface circuit 728 configured to control and receive data from the camera 735, microphone 710, and other sensors positioned on the frame 702.

[0127] The internal sensors 720 may include an IMU, which includes an electronic gyroscope, accelerometer, and magnetic compass configured to measure the movement and orientation of the wearer's head. The internal sensors 720 may also include a magnetometer, altimeter, odometer, and atmospheric pressure sensor, as well as other sensors useful for determining the orientation and movement of the smart glasses 700. Such sensors may be used in various embodiments to detect head movements, which may be used to adjust distance measurements as described.

[0128] The processing system 712 may also include a power source such as a rechargeable battery 730, which is coupled to the SOCs 902, 904, and external sensors on the frame 702.

[0129] Figure 8 is a component block diagram of a local site computing device 150 suitable for various embodiments. Referring to FIGS. 1 to Figure 8 and, the local site computing device 150 may generally include a processor 801 coupled to a volatile memory 802 and a mass non-volatile memory such as a disk drive 803. The local site computing device 150 may also include a peripheral memory access device coupled to the processor 801, such as a floppy disk drive, a compact disc (CD) or digital video disc (DVD) drive 806. The local site computing device 150 may also include a network access port 804 (or interface) coupled to the processor 801 for establishing a data connection to a network (such as the Internet and / or a local area network coupled to other system computers and servers). The local site computing device 150 may be coupled to one or more antennas (not shown) for transmitting and receiving electromagnetic radiation, and the one or more antennas may be connected to a wireless communication link. The local site computing device 150 may include additional access ports for coupling to peripheral devices, external memories, or other devices, such as USB, Firewire, Thunderbolt, etc.

[0130] The processors of the user mobile device 120 and the local site computing device (e.g., 150) can be any programmable microprocessor, microcomputer, or one or more multiprocessor chips that can be configured by software instructions (applications) to perform various functions including those described in the various embodiments below. In some user mobile devices, multiple processors may be provided (such as one processor dedicated to wireless communication functions within an SOC (e.g., 604) and one processor dedicated to running other applications within an SOC (e.g., 302)). Generally, software applications may be stored in the memory 606, and then they are accessed and loaded into the processor. The processor may include internal memory sufficient to store the software application instructions.

[0131] The various embodiments illustrated and described are provided only as examples of various features of the claims. However, the features shown and described with respect to any given embodiment need not be limited to the associated embodiment, and may be used or combined with other embodiments shown and described. Additionally, the claims are not intended to be limited to any one exemplary embodiment. For example, one or more operations of methods 400, 401, 500, 501 may replace or be combined with one or more operations of methods 400, 401, 500, 501.

[0132] Specific implementation examples are described in the following paragraphs. Although some of the following specific implementation examples are described according to example methods, other example specific implementations may include: example methods implemented by a local site computing device (or other entity) and / or a user mobile device discussed in the following paragraphs, the local site computing device (or other entity) and / or the user mobile device including a processor configured to perform the operations of these example methods; example methods implemented by a local site computing device (or other entity) and / or a user mobile device discussed in the following paragraphs, the local site computing device (or other entity) and / or the user mobile device including components for performing the functions of these example methods; example methods implemented in a processor used in a local site computing device (or other entity) and / or a user mobile device, the processor being configured to perform the operations of these example methods; and example methods implemented as a non-transitory processor-readable storage medium storing processor-executable instructions thereon, the processor-executable instructions being configured to cause a processor or a modem processor to perform the operations of these example methods.

[0133] Example 1. A method for multi-source multimedia output and synchronization in a computing device, the method comprising: receiving a user input that selects one of an audio component or a video component associated with multimedia content being rendered by a remote media player within a perceivable distance of a user of the computing device, wherein the user input indicates that the user desires to render the selected audio component, video component, or other perceivable media component on the computing device; identifying, by a processor of the computing device, the multimedia content; obtaining the identified multimedia content from a source of the multimedia content; and rendering, by the computing device in synchronization with the rendering performed by the remote media player within the perceivable distance of the user, the selected one of the audio component, video component, or other perceivable media component from the obtained multimedia content.

[0134] Example 2. The method according to Example 1, wherein the multimedia content is selected by the user from a plurality of multimedia contents observable by the user.

[0135] Example 3. The method according to any one of Examples 1 or 2, wherein receiving a user input comprises: detecting a gesture performed by the user; interpreting the detected gesture to determine whether it identifies the multimedia content being rendered within a threshold distance of the computing device; and identifying one of the audio component, video component, or other perceivable media component of the identified multimedia content that the user desires to render on the computing device.

[0136] Example 4. The method according to any one of Examples 1 to 3, wherein identifying the multimedia content being rendered on a display within a perceivable distance of the user of the computing device includes: detecting the gaze direction of the user; and identifying the multimedia content being rendered on the display in the gaze direction of the user.

[0137] Example 5. The method according to Example 4, wherein identifying the multimedia content being rendered on the display within a perceivable distance of the user of the computing device includes: receiving user input indicating the direction from which the user is perceiving the multimedia content; and identifying the multimedia content based on the received user input.

[0138] Example 6. The method according to any one of Examples 1 to 5, wherein obtaining the identified audio component, video component, or other perceivable media component of the multimedia content from the source of the multimedia content includes: obtaining metadata about the multimedia content; using the obtained metadata to identify the source of the multimedia content; and obtaining the audio component, video component, or other perceivable media component from the identified source of the multimedia content.

[0139] Example 7. The method according to any one of Examples 1 to 6, wherein obtaining the identified audio component, video component, or other perceivable media component of the multimedia content from the source of the multimedia content includes: sending a query about the multimedia content to a remote computing device; requesting identification of the source of the multimedia content; and obtaining the audio component, video component, or other perceivable media component from the identified source of the multimedia content.

[0140] Example 8. The method according to Example 7, the method further includes sampling one of the audio component or video component being rendered by the remote media player, wherein the sent query includes at least a portion of one of the sampled audio component or video component.

[0141] Example 9. The method according to Example 7, wherein at least one of the identification of the source of the multimedia content or the synchronization with the rendering performed by the remote media player is based on information received in response to the sent query.

[0142] Example 10. The method according to any one of Examples 1 to 9, wherein obtaining the identified multimedia content from the source of the multimedia content includes: obtaining a subscription access to the multimedia content from the source of the multimedia content; and receiving the identified audio component, video component, or other perceivable media component of the multimedia content based on the obtained subscription access.

[0143] Example 11. The method according to any one of Examples 1 to 10, wherein rendering, by the computing device, a selected one of the audio component, video component, or other perceivable media component synchronously with the rendering performed by the remote media player within the perceivable distance of the user includes: sampling one or more of the audio component, video component, or other perceivable media component of the multimedia content being rendered by the remote media player within the perceivable distance of the user; determining a timing difference between the samples of one or more of the audio component, video component, or other perceivable media component of the multimedia content being rendered and the audio component, video component, or other perceivable media component obtained from the source of the multimedia content; and rendering, by the computing device, the selected one of the audio component, video component, or other perceivable media component such that the user will perceive the selected one of the audio component, video component, or other perceivable media component so rendered as being synchronous with the multimedia content rendered by the remote media player.

[0144] Example 12. The method according to any one of Examples 1 to 11, wherein the computing device is an augmented reality (XR) device.

[0145] Many different cellular and mobile communication services and standards will be available or are expected in the future, and all of these services and standards can implement and benefit from various aspects. Such services and standards may include, for example, the 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE) systems, 3rd Generation Wireless Mobile Communication Technology (3G), 4th Generation Wireless Mobile Communication Technology (4G), 5th Generation Wireless Mobile Communication Technology (5G), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), 3GSM, General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA) systems (e.g., cdmaOne, CDMA1020TM), EDGE, Advanced Mobile Phone System (AMPS), Digital AMPS (IS-136 / TDMA), Evolution-Data Optimized (EV-DO), Digital Enhanced Cordless Telecommunications (DECT), Worldwide Interoperability for Microwave Access (WiMAX), Wireless Local Area Network (WLAN), Wi-Fi Protected Access I and II (WPA, WPA2), Integrated Digital Enhanced Network (iDEN), C-V2X, V2V, V2P, V2I, and V2N, and so on. For example, each of these technologies involves the sending and receiving of voice, data, signaling, and / or content messages. It should be understood that any reference to terms and / or technical details related to individual telecommunication standards or technologies is for illustrative purposes only and is not intended to limit the scope of the claims to a particular communication system or technology, unless specifically recited in the claim language.

[0146] The foregoing method descriptions and process flow diagrams are provided only as illustrative examples and are not intended to require or imply that the operations of the various embodiments must be performed in the order given. As will be appreciated by those skilled in the art, the order of operations in the foregoing embodiments may be performed in any order. Words such as "thereafter," "then," "next," etc. are not intended to limit the order of operations; these words are used to guide the reader through the description of the method. Additionally, any reference to an element of a claim in the singular form (e.g., a reference using the articles "a," "an," or "the") should not be construed as limiting that element to the singular.

[0147] The various illustrative logical blocks, modules, components, circuits, and algorithmic operations described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and operations have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the claims.

[0148] The hardware for implementing the various illustrative logical components, logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented using a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. While a general-purpose processor may be a microprocessor, in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of receiver intelligent objects, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Alternatively, some operations or methods may be performed by circuitry specific to a given function.

[0149] In one or more embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. Operations of a method or algorithm disclosed herein may be implemented in a processor-executable software module or processor-executable instructions that may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium accessible by a computer or a processor. By way of example and not limitation, such non-transitory computer-readable or processor-readable storage media may include RAM, ROM, EPROM, flash memory, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, disk and optical disks include compact disk (CD), laser disk, optical disk, digital versatile disk (DVD), floppy disk, and Blu-ray disk, where disks typically reproduce data magnetically, while optical disks reproduce data optically with lasers. Combinations of the above are also included within the scope of non-transitory computer-readable and processor-readable media. Additionally, operations of a method or algorithm may reside as code and / or instructions in one or any combination or set on a non-transitory processor-readable storage medium and / or a computer-readable storage medium, which may be incorporated into a computer program product.

[0150] The foregoing description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the claims. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the scope of the claims. Thus, the present disclosure is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.

Claims

1. A method for multi-source multimedia output and synchronization in a computing device, the method comprising: Receiving a user input that selects one of an audio component, a video component, or other perceivable media component associated with multimedia content being rendered by a remote media player within a perceivable distance of a user of the computing device, wherein the user input indicates that the user desires to render the selected one of the audio component, the video component, or the other perceivable media component on the computing device; Identifying, by a processor of the computing device, the multimedia content; Obtaining the identified multimedia content from a source of the multimedia content; And Rendering, by the computing device, in synchronization with the rendering performed by the remote media player within the perceivable distance of the user, the selected one of the audio component, the video component, or the other perceivable media component from the obtained multimedia content.

2. The method according to claim 1, wherein the multimedia content is selected by the user from a plurality of multimedia contents observable by the user.

3. The method according to claim 1, wherein receiving a user input comprises: Detecting a gesture performed by the user; Interpreting the detected gesture to determine whether it identifies the multimedia content being rendered within a threshold distance of the computing device; And Identifying one of the audio component, the video component, or the other perceivable media component of the identified multimedia content that the user desires to render on the computing device.

4. The method according to claim 1, wherein identifying the multimedia content being rendered on a display within a perceivable distance of the user of the computing device comprises: Detecting a gaze direction of the user; And Identifying the multimedia content being rendered on the display in the gaze direction of the user.

5. The method according to claim 4, wherein identifying the multimedia content being rendered on the display within a perceivable distance of the user of the computing device comprises: Receiving a user input indicating a direction from which the user perceives the multimedia content; And Identifying the multimedia content based on the received user input.

6. The method according to claim 1, wherein obtaining the identified audio component, video component, or other perceivable media component of the multimedia content from a source of the multimedia content comprises: Obtaining metadata about the multimedia content; Using the obtained metadata to identify a source of the multimedia content; And Obtaining the audio component, video component, or other perceivable media component from the identified source of the multimedia content.

7. The method according to claim 1, wherein obtaining the identified audio component, video component, or other perceivable media component of the multimedia content from a source of the multimedia content comprises: Sending a query about the multimedia content to a remote computing device; Requesting an identification of a source of the multimedia content; And Obtaining the audio component, video component, or other perceivable media component from the identified source of the multimedia content.

8. The method according to claim 7, wherein the method further comprises: Sampling one of the audio component or the video component being rendered by the remote media player, wherein the query sent includes at least a portion of one of the sampled audio component or video component.

9. The method according to claim 7, wherein at least one of the identification of the source of the multimedia content or the synchronization with the rendering performed by the remote media player is based on information received in response to the query sent.

10. The method according to claim 1, wherein obtaining the identified multimedia content from the source of the multimedia content comprises: Obtaining a subscription access to the multimedia content from the source of the multimedia content; And Receiving the identified audio component, video component or other perceivable media component of the multimedia content based on the obtained subscription access.

11. The method according to claim 1, wherein rendering, by the computing device, one of the audio component, the video component or other perceivable media component synchronously with the rendering performed by the remote media player within the perceivable distance of the user of the computing device comprises: Sampling one or more of the audio component, video component or other perceivable media component of the multimedia content being rendered by the remote media player within the perceivable distance of the user; Determining a timing difference between a sample of one or more of the audio component, video component or other perceivable media component of the multimedia content being rendered and the audio component, video component or other perceivable media component obtained from the source of the multimedia content; And Rendering, by the computing device, one of the audio component, video component or other perceivable media component such that the user will perceive the one of the audio component, video component or other perceivable media component so rendered as being synchronized with the multimedia content rendered by the remote media player.

12. The method according to claim 1, wherein the computing device is an augmented reality (XR) device.

13. A computing device, the computing device comprising: A transceiver; And A processor, the processor being coupled to the transceiver and configured to: Receive a user input that selects one of the audio component, video component or other perceivable media component associated with the multimedia content being rendered by a remote media player within the perceivable distance of a user of the computing device, wherein the user input indicates that the user desires to render the one of the audio component, video component or other perceivable media component on the computing device; Identify the multimedia content; Obtain the identified multimedia content from the source of the multimedia content; And Render, by the computing device, in synchronization with the rendering by the remote media player within the perceivable distance of the user, a selected one of the audio component, video component, or other perceivable media components from the obtained multimedia content.

14. The computing device according to claim 13, wherein the processor is configured such that the multimedia content is selected by the user from a plurality of multimedia contents observable by the user.

15. The computing device according to claim 13, wherein the processor is further configured to receive user input by: detecting a gesture performed by the user; interpreting the detected gesture to determine whether it identifies the multimedia content being rendered within a threshold distance of the computing device; and identifying one of the audio component, video component, or other perceivable media components of the identified multimedia content that the user desires to render on the computing device.

16. The computing device according to claim 13, wherein the processor is further configured to identify the multimedia content being rendered on a display within the perceivable distance of the user of the computing device by: detecting the gaze direction of the user; and identifying the multimedia content being rendered on the display in the gaze direction of the user.

17. The computing device according to claim 16, wherein the processor is further configured to identify the multimedia content being rendered on the display within the perceivable distance of the user of the computing device by: receiving user input indicating the direction from which the user is perceiving the multimedia content; and identifying the multimedia content based on the received user input.

18. The computing device according to claim 13, wherein the processor is further configured to obtain the identified audio component, video component, or other perceivable media component of the multimedia content from the source of the multimedia content by: obtaining metadata about the multimedia content; using the obtained metadata to identify the source of the multimedia content; and obtaining the audio component, video component, or other perceivable media component from the identified source of the multimedia content.

19. The computing device according to claim 13, wherein the processor is further configured to obtain the identified audio component, video component, or other perceivable media component of the multimedia content from the source of the multimedia content by: sending a query about the multimedia content to a remote computing device; requesting an identification of the source of the multimedia content; and obtaining the audio component, video component, or other perceivable media component from the identified source of the multimedia content.

20. The computing device according to claim 19, wherein the processor is further configured to: sample one of the audio component or video component being rendered by the remote media player, wherein the sent query includes at least a portion of one of the sampled audio component or video component.

21. The computing device according to claim 19, wherein the processor is further configured to identify the source of the multimedia content or synchronize the rendering performed by the remote media player based on information received in response to the sent query.

22. The computing device according to claim 13, wherein the processor is further configured to obtain the identified multimedia content from the source of the multimedia content by: obtaining a subscription access to the multimedia content from the source of the multimedia content; and receiving the identified audio component, video component, or other perceivable media component of the multimedia content based on the obtained subscription access.

23. The computing device according to claim 13, wherein the processor is further configured to render, by the computing device, a selected one of the audio component, the video component, or other perceivable media components synchronously with the rendering performed by the remote media player within the perceivable distance of the user by: sampling one or more of the audio component, video component, or other perceivable media components of the multimedia content being rendered by the remote media player within the perceivable distance of the user; determining a timing difference between a sample of one or more of the audio component, video component, or other perceivable media components of the multimedia content being rendered and the audio component, video component, or other perceivable media components obtained from the source of the multimedia content; and rendering, by the computing device, a selected one of the audio component, video component, or other perceivable media components such that the user will perceive the selected one of the audio component, video component, or other perceivable media components so rendered as being synchronized with the multimedia content rendered by the remote media player.

24. The computing device according to claim 13, wherein the computing device is an augmented reality (XR) device.

25. A computing device, the computing device comprising: means for receiving a user input that selects one of an audio component, a video component, or other perceivable media components associated with multimedia content being rendered by a remote media player within a perceivable distance of a user of the computing device, wherein the user input indicates that the user desires to render the selected one of the audio component, the video component, or other perceivable media components on the computing device; means for identifying the multimedia content; means for obtaining the identified multimedia content from the source of the multimedia content; and means for rendering, by the computing device, a selected one of the audio component, the video component, or other perceivable media components from the obtained multimedia content synchronously with the rendering performed by the remote media player within the perceivable distance of the user.

26. The computing device according to claim 25, wherein the means for receiving a user input comprises: means for detecting a gesture performed by the user; A component for interpreting a detected gesture to determine whether it identifies the multimedia content being rendered within a threshold distance of the computing device; and A component for identifying one of the audio component, video component, or other perceivable media component of the identified multimedia content that the user wishes to render on the computing device.

27. The computing device according to claim 25, wherein the component for obtaining the identified audio component, video component, or other perceivable media component of the multimedia content from the source of the multimedia content includes: A component for obtaining metadata about the multimedia content; A component for using the obtained metadata to identify the source of the multimedia content; and A component for obtaining the audio component, video component, or other perceivable media component from the identified source of the multimedia content.

28. The computing device according to claim 25, wherein the component for obtaining the identified audio component, video component, or other perceivable media component of the multimedia content from the source of the multimedia content includes: A component for sending a query about the multimedia content to a remote computing device; A component for requesting an identification of the source of the multimedia content; and A component for obtaining the audio component, video component, or other perceivable media component from the identified source of the multimedia content.

29. The computing device according to claim 25, wherein the component for obtaining the identified multimedia content from the source of the multimedia content includes: A component for obtaining a subscription access to the multimedia content from the source of the multimedia content; and A component for receiving the identified audio component, video component, or other perceivable media component of the multimedia content based on the obtained subscription access.

30. A non-transitory processor-readable storage medium storing processor-executable instructions thereon, the processor-executable instructions being configured to cause a processor of a computing device to perform operations, the operations including: Receiving a user input that selects one of the audio component, video component, or other perceivable media component associated with multimedia content being rendered by a remote media player within a perceivable distance of a user of the computing device, wherein the user input indicates that the user wishes to render the selected one of the audio component, video component, or other perceivable media component on the computing device; Identifying the multimedia content; Obtaining the identified multimedia content from the source of the multimedia content; and Rendering, by the computing device in synchronization with the rendering by the remote media player within the perceivable distance of the user, the selected one of the audio component, video component, or other perceivable media component from the obtained multimedia content.