Interactions between encapsulation and synchronization of state between devices
By establishing sessions between playback devices, identifying and synchronizing the playback modes of target devices, the problem of low bandwidth efficiency in network device self-organizing groups is solved, achieving efficient media rendering and resource saving.
Patent Information
- Application Number
- CN202211429008.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-09-28
- Filing Date
- 2018-10-19
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2038-10-19
AI Technical Summary
Existing network devices in self-organizing groups have low bandwidth efficiency for media playback applications, especially under uplink bandwidth constraints or digital rights management (DRM) constraints, which causes the first device to retransmit content to other devices in the group, resulting in a waste of communication resources.
By establishing sessions between playback devices, identifying the playback mode of the target device, and synchronizing the rendering of media assets between devices, devices can directly transmit assets from the media source stream. The master device coordinates the rendering operations of other devices to achieve consistent playback.
It improves the media playback efficiency of network device self-organizing groups, saves communication resources, and ensures consistent media rendering even under bandwidth-limited or DRM-constrained conditions.
Smart Images

Figure CN115713949B_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with application number 201811220623.4 and titled "Interaction of packaging and synchronizing state between devices" and filed on October 19, 2018. TECHNICAL FIELD
[0002] The present disclosure relates to media playback operations, and in particular to the synchronization of playback devices grouped together for playback in an ad hoc manner. BACKGROUND
[0003] Media playback devices are well known in consumer applications. Typically, they involve playback of video or audio at a consumer device such as a media player device (television, audio system). Recently, playback applications have expanded to include playback over network devices, such as over a Bluetooth or WiFi network. In network applications, playback typically involves streaming playback content from a first player device, such as a smartphone, to connected rendering devices. The first player device can not store the playback content itself; it can stream the content from another source. SUMMARY
[0004] According to some embodiments of the present disclosure, a method is provided, comprising: in response to a command to change a playback session of networked rendering devices, identifying from the command a target device of the command; determining from a function of the target device a playback mode of the target device; and synchronizing playback of the target device with other devices that are members of the session, the synchronization including identifying at least one asset to be rendered according to the playback session and playback timing of the asset. BRIEF DESCRIPTION OF DRAWINGS
[0005] Figure 1 A system according to one aspect of the present disclosure is shown.
[0006] Figure 2 A block diagram of a rendering device according to one aspect of the present disclosure.
[0007] Figure 3 A state diagram showing the progression of a synchronization session according to embodiments of the present disclosure.
[0008] Figure 4 A signal flow diagram showing an exemplary process of adding a device to a session according to one aspect of the present disclosure.
[0009] Figure 5 A signal flow diagram showing an exemplary rendering of playback content according to one aspect of the present disclosure.
[0010] Figure 6A signal flow diagram to illustrate an example rendering of playback content according to another aspect of the disclosure.
[0011] Figure 7 A signal flow diagram to illustrate an example process of removing a rendering device from a session according to another aspect of the disclosure.
[0012] Figure 8 An example user interface that can be rendered in a displayable device according to one aspect of the disclosure is shown.
[0013] Figure 9 An example user interface for managing devices in a session is shown.
[0014] Figure 10 A signal flow diagram to illustrate an example group management of playback content according to one aspect of the disclosure.
[0015] Figure 11 An example user interface for managing devices in a session is shown.
[0016] Figure 12 A signal flow diagram to illustrate an example group management according to another aspect of the disclosure.
[0017] Figure 13 A method according to one aspect of the disclosure is shown.
[0018] Figure 14 And Figure 15 Two use cases for inputting commands to manage rendering devices according to aspects of the disclosure are shown.
[0019] Figure 16 A method according to another aspect of the disclosure is shown.
[0020] Figure 17 A method according to yet another aspect of the disclosure is shown.
[0021] Figure 18 A method according to another aspect of the disclosure is shown.
[0022] Figure 19 A network diagram for a system according to yet another aspect of the disclosure.
[0023] Figure 20A A block diagram of a virtual assistant system according to one aspect of the disclosure is shown.
[0024] Figure 20B A function of a virtual assistant system according to one aspect of the disclosure is shown.
[0025] Figure 21 A system according to another aspect of the disclosure is shown.
[0026] Figure 22 A communication flow for session management is shown in accordance with one aspect of the disclosure.
[0027] Figure 23 A method of session management is shown in accordance with one aspect of the disclosure. DETAILED DESCRIPTION
[0028] Aspects of the disclosure provide techniques for managing media playback among a group of self-organizing playback devices. Such techniques can involve establishing a session among the playback devices, where the playback devices communicate information about their playback capabilities. Based on the playback capabilities of the devices, a playback mode can be derived for the session. Playback operations can be synchronized among the member devices of the session, where the devices receive an identification of an asset to be rendered according to the playback operation and timing information for playback of the asset. When the devices are able to do so, the devices can stream the playback asset directly from a media source. This conserves communication resources.
[0029] While media playback over a self-organizing group of network devices is known, it is often inefficient. Such applications typically result in a first device in the group streaming playback content to each device in the group. This results in low bandwidth efficiency in such applications as the first device retrieves content for streaming that is then re-transmitted by the first device to be rendered by other devices in the group. However, the techniques of the disclosure allow playback to be performed in a consistent manner across all devices of a session, even when uplink bandwidth is constrained or when digital rights management ("DRM") constraints limit access to assets.
[0030] Figure 1 A system 100 is shown in accordance with one aspect of the disclosure. The system 100 can include a plurality of rendering devices 110.1-110.N provided in communication with one another via a network 120. The rendering devices 110.1-110.N can cooperate to render media items in a coordinated manner.
[0031] During operation, one of the rendering devices (e.g., device 110.1) can operate in the role of a "master" device. The master device 110.1 can store data, referred to for convenience as a "playback queue," representative of a playlist of media assets to be played by the group of rendering devices 110.1-110.N. The playback queue can identify, for example, a plurality of assets to be played in succession, a current playback mode (e.g., audio assets to be played according to a "random" or "repeat" mode), etc. In other aspects, the assets can be identified on the basis of an algorithm, e.g., an algorithmic station, or in the case of a third-party application, a data structure maintained by the application. The master device 110.1 can communicate with the other rendering devices 110.2-110.N to synchronize playback of the media assets.
[0032] The other rendering devices 110.2-110.N in the group, referred to as "secondary" devices, can play the asset in synchronization. In one aspect, each of the devices 110.1-110.N can stream the asset directly from the media source 130 on the network. The secondary devices 110.2-110.N typically store data representative of the asset currently being played and a timebase for synchronized rendering. In another aspect, a rendering device (e.g., device 110.3) is not capable of streaming the asset from the media source 130, and the primary device 110.1 can stream the data directly to the non-capable rendering device 110.3. In another aspect, the primary device 110.1 can decrypt and decode the asset and then stream the asset data (e.g., audio bytes) to the secondary device without regard to the capabilities of the secondary device.
[0033] The principles of the present discussion find application with a variety of different types of rendering devices 110.1-110.N. These can include a smartphone or tablet computer, as represented by the rendering device 110.1; a display device 110.2; and a speaker device 110.1-110.N. Although not shown in FIG. 1, the principles of the present disclosure can be applied to other types of devices, such as a laptop computer and / or personal computer, a personal media player, a set-top box, an optical disc-based media player, a server, a projector, a wearable computer, an embedded player (e.g., a car music player), and a personal gaming device. Figure 1
[0034] The rendering devices 110.1-110.N can be provided in one or more networks in communication with each other, collectively shown as the network 120. The network 120 can provide a communication architecture through which the various rendering devices 110.1-110.N discover each other. For example, in a residential application, the network 120 can provide a communication architecture through which rendering devices in a common household can discover each other. In an office application, the network 120 can provide a communication architecture through which rendering devices in a common building, common enterprise, and / or common campus can discover each other. Unless otherwise indicated herein, the architecture and / or topology of the network 120, including the number of networks 120, is immaterial to the present discussion.
[0035] Figure 2 A block diagram of a rendering device 200 in accordance with one aspect of the present disclosure. The rendering device 200 can include a processor 210, a memory 220, a user interface 230, a transceiver 240, and a display 250 and / or speakers 260 appropriate to the type of device.
[0036] The rendering device may include a processor 210 that executes program instructions stored in memory 220. These program instructions may define an operating system 212 for the rendering device 200; a synchronization manager 214 (discussed herein); and various application programs 216.1-216.N for the rendering device 200, relating to rendering device 200 and corresponding rendering devices 110.1-110.N. Figure 1 The playback operation is performed by the rendering device 200 and / or the corresponding rendering device 110.1-110.N. Figure 1 Rendering.
[0037] The rendering device 200 may also include other functional units, such as a user interface 230, a transceiver (“TX / RX”) 240, a display 250, and / or a speaker 260. The user interface 230 provides controls through which the rendering device 200 interacts with the user. User interface components may include various operator controls, such as buttons, switches, pointing devices, etc., and various output components, such as indicators, lights, speakers, and / or displays.
[0038] The TX / RX 240 can provide an interface through which the rendering device 200 communicates with the network 120. Figure 1 It communicates with other rendering devices via extensions and with media sources as needed.
[0039] Display 250 and / or speaker 260 represent rendering components through which device 200 renders playback content: video and / or audio, depending on the type of rendering device. The display and speaker are shown separately from user interface 230 merely to emphasize the playback operations that can be performed by rendering device 200. In practice, the same display and / or speaker involved in user interface operations will also be involved in playback operations.
[0040] As noted, the principles of this disclosure relate to the synchronization of playback operations between various rendering devices, and specifically, to multiple types of rendering devices. Therefore, a given rendering device does not need to have… Figure 2 All components shown. For example, a speaker device does not need to include a display. A set-top box device does not need to include either a speaker or a display. Therefore, when the principles of this disclosure are put into operation, it is expected that... Figure 2 The block diagram has some discrepancies.
[0041] Figure 3To illustrate the progression of a synchronization session 300 according to embodiments of the present disclosure, a state diagram is shown. The synchronization session 300 can proceed according to several main stages, including session revision stages (shown as stages 310, 320, respectively), a synchronization / playback stage 330, and optionally, a queue migration state 340. The session revision stages can be entered to add a rendering device to the synchronization session (stage 320) or to remove a rendering device from the session (stage 320). The synchronization / playback stage 330 represents the stage of operation in which devices that are currently members of the common synchronization group exchange information about playback. The queue migration stage 340 can be entered to transfer the state of the playback queue from one rendering device to another, effectively designating a new master device for a group of devices.
[0042] The synchronization session 300 can be initiated by the first two rendering devices (possibly more) to the session 300. Typically, the session is initiated from a first rendering device (e.g., device 110.1), which indicates that the rendering device 110.1 is to be included in the session. A second rendering device (e.g., device 110.2) can be identified. Figure 1 Figure 1
[0043] According to this identification, the session 300 can enter a revision state 310 to add the selected device to the session 300. The devices 110.1, 110.2 can communicate with each other to negotiate a playback mode for rendering. Once the negotiation is complete, the devices can enter a synchronization / playback state 330 during which the member devices of the session exchange information for rendering.
[0044] The session can progress to other states in response to other control inputs. For example, if an input is received to add another device (e.g., device 110.3) to the session 300, the session 300 can return to state 310 in which the master device (e.g., device 110.1) negotiates a playback mode for rendering with the new device 110.3.
[0045] If an input is received to remove a device from the session 300, the session 300 can progress to another revision state 320 in which the identified device (e.g., device 110.3) is removed from the session. In this state, the master device 110.1 can communicate with the device 110.3 to be removed to terminate the device’s membership in the session. The device 110.3 should terminate rendering.
[0046] Upon removal of a device from the session 300, the session 300 can progress to several other states. Removal of a device can isolate the last two devices of the session (e.g., devices 110.2 and 110.3) (i.e., no group of devices in the session), which can result in the end of the session 300.
[0047] In another use case, the device being deleted can be the "master" device of the session. In this case, deleting the master device can cause the session to enter a queue migration state 340, in which a new master device is selected from the devices remaining in the session, and playback queue information is transferred to the newly selected master device. Once the new master device is selected, the session 300 can return to the sync / playback state 330.
[0048] In another aspect, queue migration can be performed in response to operator control. For example, an operator can interact with a device that is not currently acting as a master device. The operator can engage in control to change the playback session, for example, by changing the playback mode or the playback asset (e.g., switching from one asset playlist to another). In response, the devices within the session can proceed to the queue migration state 340 to transfer the queue to the device with which the operator interacted, and then return to the sync / playback state 330 once the operator engages the new playback mode.
[0049] In yet another aspect, deletion of a device 320 can be performed in response to device failure, if necessary, deletion of queue migration 340 can be performed. The exchange of synchronization information can involve messaging between devices within the session. If a message is not received from a given device within a predetermined period of time, it can be interpreted by the other devices as a device failure. In response, the session can proceed to state 320 to delete the failed device from the session, and if the failed device was the master device, perform queue migration in state 340.
[0050] In general, the playback session can be supported by various user interface components at the devices. Some example user interfaces are shown in Figure 8 and Figure 9 In one aspect, state changes can be implemented as atomic transactions prior to updating the user interface. In this way, the user interface can be periodically refreshed during operation in the sync / playback state 330, which simplifies the presentation of the user interface.
[0051] If neither a termination event nor a queue migration event is triggered, when session modification is complete at state 320, the session 300 can return to the sync / playback state 330.
[0052] Figure 4 A signal flow 400 between devices for adding a device to the session 300 Figure 3 ) is shown in accordance with one aspect of the disclosure. In this example, two rendering devices 110.2, 110.3 are added to the session from a first rendering device 110.1 that manages the session.
[0053] As shown, a session 300 can be started by adding two devices (here, device 110.1 and device 110.2) to the session. In response to user input indicating that device 110.2 is to be added to the session, rendering device 110.1 can transmit a message to rendering device 110.2 requesting its playback capabilities (message 410). Rendering device 110.2 can respond by providing information about its playback capabilities (message 420).
[0054] Device 110.1 can establish a session object that identifies device 110.1 and device 110.2 as members of the session and various features of the playback rendering operation to be performed (block 430). Thereafter, rendering devices 110.1, 110.2 can exchange synchronization information about the playback rendering operation (message 440).
[0055] The synchronization information can include information describing both the type of device and the rendering modes supported by the device. For example, a device can identify itself as video-capable, audio-capable, or both. A device can identify supported playback applications (e.g., iTunes, Spotify, Pandora, Netflix, etc.) and, where applicable, account identifiers associated with such information.
[0056] From such information, the master device can establish a session object that identifies the asset to be rendered, timing information for the asset, and other playback mode information (e.g., the application to be used for rendering). State information from the session object can be distributed to other devices in the session throughout its lifetime.
[0057] The playback rendering operation can vary as the type of media to be played by rendering devices 110.1, 110.2 changes, and thus the type of synchronization information can also vary. Consider the following use cases:
[0058] Both rendering devices 110.1, 110.2 are playing from a music playlist. In this case, the master device can store data representing the playback queue, e.g., the audio assets that make up the audio playlist, any content services through which the audio assets are accessed (e.g., iTunes), the playback mode (e.g., random, play order), etc. In one aspect, the synchronization information can identify the assets being played (e.g., by URL and digital rights management token), timing information for rendering the music assets, and the roles of devices 110.1, 110.2 (e.g., whether the device is playing the entire asset or a designated channel of the asset— left, right). Rendering devices 110.1, 110.2 can use this information to fetch the music assets identified in the playlist and synchronize rendering of the music assets. Changes in the playback mode (e.g., skip to previous track, skip to next track, random) can result in new synchronization information being transmitted between the devices.
[0059] In another aspect, the synchronization information can identify the content service being used and identify a handle (possibly a user ID, password, session ID) that identifies the service parameters through which the service is being accessed, as well as any parameters of the respective roles of the devices in playback (e.g., video, audio, audio channel, etc.). In such cases, the rendering devices 110.1, 110.2 can authenticate themselves with the server 130 that supports the content service. The server 130 can maintain information about the assets being played, the order of playback, etc., and can provide the assets directly to the devices.
[0060] Both rendering devices 110.1, 110.2 play ordinary audiovisual assets. In this case, the master device can store data representing the playback queue, e.g., the video components and audio components from the assets to be rendered, any content services through which the assets are accessed, and playback settings. Many audiovisual assets can contain multiple video components (e.g., representations of a video at different screen resolutions and / or encoding formats, representations of a video at different viewing angles, etc.) and multiple audio components (e.g., audio tracks in different languages). The synchronization information can identify the asset components to be rendered by each device in the group and timing information for rendering the respective components (e.g., reference points between the timelines of the assets). The rendering devices 110.1, 110.2 can use this information to obtain the components of the audiovisual assets that are relevant to the respective types of devices. For example, a display device can obtain the video portions of the assets that are consistent with the configuration of the display, and a speaker device can obtain the audio portions of the assets that are consistent with the rendering settings (e.g., the English audio track or the French audio track as specified by the playback settings).
[0061] As shown, when a new master device is selected, a queue migration can be performed (state 340). The selection of the new master device can be performed after an exchange of messages between the devices, and can be performed in various ways. In a first aspect, various devices for the session can not be capable of functioning as a master device, e.g., because they are feature-limited devices. Thus, based on device capabilities, devices can be excluded from being candidates for operating as a master device. The master device can be selected among the candidate devices based on the features of the devices. For example, the devices can vary based on whether they are battery-powered or line-powered, based on their connectivity bandwidth to the network 120, based on their load (e.g., whether they are close to their heat-emission limits), based on features such as the frequency of interrupts to process other tasks.
[0062] Selection of a master device can be performed in response to one or more of these factors, preferably (optionally) factors that are indicative of higher reliability. For example, it can be determined that a line-powered device is more reliable than a battery-powered device. It can be determined that a device operating near its heat limit is less reliable than a device not operating near its heat limit. A device with higher bandwidth connectivity can be considered more reliable than a device with low bandwidth. A device that is not frequently interrupted can be considered more reliable than a device that is interrupted at a higher rate. Selection of a new master device can result from consideration of these factors, whether individually or collectively.
[0063] Figure 4 Operations for adding another device 110.3 to the session are also shown. Here, the session is shown being added to the session in response to user input at rendering device 110.1. Rendering device 110.1 can transmit a request for information of its playback capabilities to rendering device 110.3 (message 450). Rendering device 110.2 can respond by providing information about its playback capabilities (message 460).
[0064] Device 110.1 can modify the session object to add device 110.3 (block 470). In the illustrated example, the session object identifies devices 110.1, 110.2, and 110.3 as members of the session. It also identifies various features of the playback rendering operation to be performed. These features can change based on the overlap of capabilities of devices 110.1, 110.2, and 110.3 according to the features identified in block 430. Thereafter, rendering devices 110.1, 110.2, and 110.3 can exchange synchronization information about the playback rendering operation (message 480).
[0065] Figure 5 A signal flow diagram showing an example rendering of playback content according to session 300( Figure 3 ) in accordance with an aspect of the disclosure. In this example, three rendering devices 110.1-110.3 are members of a common playback session.
[0066] In this playback example, each rendering device has the capability to playback media content using a general purpose application. Thus, each rendering device 110.1-110.3 streams its respective playback content from media source 130 via a network and renders the playback content independently of one another, represented by blocks 520, 530, and 540. In this example, rendering devices 110.1-110.3 communicate with one another to synchronize playback (message 510), but they acquire the playback content to be rendered through independent communication processes with media source 130.
[0067] Figure 6 A signal flow diagram showing an example rendering of playback content according to session 300( Figure 3The following is an exemplary signal flow diagram for rendering playback content. In this example, three rendering devices 110.1 to 110.3 are members of a normal playback session.
[0068] In this playback example, the rendering device does not have the general capability to play back media content using a generic application. Devices 110.1 and 110.2 have the capability to obtain their respective playback content from media source 130, but device 110.3 does not.
[0069] In this example, rendering devices 110.1 and 110.2 stream their respective playback content from media source 130 via the network and render the playback content independently of each other, as indicated by boxes 620 and 630. One of the devices (device 110.1 in this example) can acquire content for rendering device 110.3 and transmit that playback content to rendering device 110.3 (box 640). Therefore, rendering device 110.3 acquires its content from another device 110.1 in the session rather than from the media source.
[0070] In some applications, rendering device 110.3 may not have the ability to render content in a given playback mode due to Digital Rights Management (DRM) issues. For example, if an attempt is made to add a device that does not have the rights to participate in the playback mode of the current activity (e.g., because the device to be added does not have an account logged in from the media source from which the content is being played) to an ongoing session, several results may occur. In one case, a non-supporting rendering device may not be added. In another case, an alternative playback mode may be derived to play the same assets in the current playback mode (e.g., by switching to an alternative application from which the new device does indeed have access rights).
[0071] Figure 7 To illustrate another aspect of this disclosure, from session 300 ( Figure 3 The signal flow diagram illustrates an exemplary process for removing rendering device 110.2. In this example, three rendering devices 110.1 to 110.3 are members of a normal playback session from the outset. In this example, these three devices 110.1 to 110.3 may participate in periodic synchronization communication (information 710) to synchronize playback between devices.
[0072] At a certain point, user input indicating that device 110.2 should be removed from the session can be received. In response, the master device (here, device 110.1) can modify the session object to recognize device 110.2 as an abandoned device. The master device 110.1 can transmit a message 730 to the abandoned device 110.2 indicating that the device will end playback. In response, the rendering device 110.2 can end its playback (box 740). Thereafter, synchronization messages 750 related to playback will be exchanged between devices 110.1 and 110.3 that remain in the session.
[0073] On the other hand, device removal can be initiated at the device being removed. In this event (not shown), the device being removed can terminate its playback and transmit a message to the master device notifying the master device of the device's removal.
[0074] Although the foregoing discussion has presented session management functions controlled by a general-purpose device (device 110.1), the principles of this discussion are not limited to this. Therefore, the principles of this discussion seek applications where different devices change session state. Figure 3 Therefore, adding a device to a session (state 310) can be initiated by any device (including the device to be added to the session). Similarly, removing a device from a session (state 320) can be initiated by any device (master or secondary). Queue migrations, as discussed, can be user-controlled or initiated by an event that removes a master device from a session.
[0075] As an illustrative example, consider an arrangement where an auxiliary device joins an existing session in which another device is already managing the device's activities. This could be, for example, in an arrangement using a smart speaker 110.3 ( Figure 1 —This occurs when a speaker device manages a playback session for a group of devices residing in the session. A user can add his / her smartphone 110.1 to the session. In this case, the smartphone 110.1 can initiate an operation to join the session itself. Figure 3 (State 310). Initially, it may assume the role of an auxiliary device and begin rendering media assets (e.g., video content, image content) associated with the media asset being played. Eventually, the user can participate in playback control operations via smartphone 110.1, which may trigger a queue migration event (state 340) that transmits playback queue information to smartphone 110.1. Furthermore, the user may remove smartphone 110.1 from the session (state 320), which may trigger a second queue migration event (state 340) that transmits playback queue information to another device in the session—possibly back to smart speaker 110.3.
[0076] Alternatively, after a queue migration event in which the smart phone 110.1 becomes the master device, a user (or another user) can engage in playback control operations through the speaker 110.3. In such cases, another queue migration event can occur in which playback queue information is transmitted to the smart speaker (state 340).
[0077] Queue migration need not be performed in all cases in which a user enters commands through a connected device. In other embodiments, devices that are members of a group can share remote control user interfaces that display information to users at these devices about an ongoing playback session. Such remote control displays can allow users to enter commands that, if not entered at the master device, can be communicated to the master device to change playback modes. Queue migration need not occur in these use cases.
[0078] Figure 8 An example user interface 800 for session 300( Figure 3 ) management is shown that can be rendered in a displayable device according to one aspect of the disclosure. The user interface 800 can include controls 810 for managing device membership in a session, controls 820, 825 to manage playback of assets in a playback session, and areas 830, 835 for rendering asset information.
[0079] The controls 810 can include controls for adding or removing devices from a session. In the example shown, the control area 810 includes a first control 812 for adding devices, a second control 814 that displays devices that are currently in the session, and a third control 816 for accessing other session management controls. Figure 8
[0080] The controls 820, 825 can control asset playback. In the example shown, they include play / pause controls, controls to make playback skip back to a previous asset or forward to a next asset, volume controls, and controls to make playback jump to a specified position along the playback timeline of an asset. Although not shown in this example, controls can also be invoked to change playback modes (e.g., normal play, random, repeat), change playlists, change services used to receive media, change user accounts used to receive media, etc. Further, user controls specific to the type of asset being rendered can be provided, with a different set of user controls for audio assets than for video assets. Figure 8
[0081] The areas 830, 835 can provide a display of asset content or metadata about that content. For audio information, it can include graphical images and / or textual information associated with the rendered audio (e.g., artist images, artwork, track names, etc.). For video, it can include video data of the asset itself.
[0082] Figure 9 An example user interface 900 for managing devices in a session 300( Figure 3 ) is shown. In this example, the user interface 900 can include a control area 910 for adding and / or deleting devices in the session, which can include other areas 920, 930, 940-945 for displaying other information. For example, area 920 can display information about the playback mode for the session. Area 930 can provide user controls for managing playback. Areas 940-945 can display status indicators for other devices that are not members of the session but can be added.
[0083] Figure 10 A signal flow diagram to illustrate example group management of playback content according to an aspect of the disclosure. Figure 10 The illustrated techniques find application in an environment in which a new device (here, rendering device 110.3) attempts to join an ongoing session 300( Figure 3 ), but is prevented from discovering the master device 110.1 in the group session, for example, due to user permissions or other constraints that prevent direct communication between devices. For ease of discussion, assume that the joining device 110.3 can discover and communicate with another member of the group. In the example of Figure 10 , the joining device 110.3 can communicate with device 110.2.
[0084] In this example, playback synchronization 1010 is performed between the two devices 110.1, 110.2 that are members of the group. The joining device 110.3 can send a request to join the group to the non-master device 110.2 (message 1020). In response, the non-master device 110.2 can relay the join request message to the master device 110.1 (message 1030). The master device can request the capability to join device 110.3 via a communication to the non-master device 110.2 (message 1040), and the non-master device 110.2 can relay the communication to the joining device 110.3 (message 1050). The joining device 110.3 can identify its capabilities in a response message (message 1060), which is sent to the non-master device 110.2 and relayed to the master device 110.1 (message 1070). The master device can modify the session object to identify rendering device 110.3 as a destination device (block 1080). Thereafter, playback synchronization 1090 can be performed among the three rendering devices 110.3.
[0085] Figure 10 Embodiments of the disclosure can find application using so-called "dumb devices," which refer to devices that do not have the functionality to act as a master device in group management. Thus, in Figure 10The illustrated aspects, even in cases where the rendering device 110.2 is not capable of performing session management operations itself, the rendering device 110.2 can relay group management messages between other devices.
[0086] In another aspect, in cases involving dumb devices, the master device can provide such devices with user interface controls that provide group controls as well as controls that apply only to the local device. Figure 11 An exemplary user interface 1100 for managing devices in a session 300 Figure 3 ) is shown. In this example, the user interface 1100 can include a control area 1110 for controlling the devices within a session. It can include a first control 1112, shown as a slider bar, for controlling the playback volume of all devices in the current session. It can include other controls 1114, 1116 for controlling the playback volume of session devices individually. One such control, for example control 1114, can control the local device.
[0087] In response to user interaction with the local control 1114, the device can perform the action directly indicated by the user input. Thus, in response to interaction with the volume control, the rendering device can change its volume output accordingly. In response to interaction with a control that represents the session group as a whole or as another device, the rendering device can report the user input to the master device, which will issue volume control commands to the other devices in the session.
[0088] In another aspect, the session devices can perform operations to simulate queue migration in cases where regular migration is not possible. Consider a case where a session remains in place, it includes both full-featured rendering devices (devices that can act as master devices) as well as other dumb devices (devices that cannot act as master devices). In such implementations, the master device can provide the dumb devices with a user interface that presents session management functionality. It can happen that a user controls the session to remove the current master device from the session, which will move the queue management responsibility to another dumb device. In such cases, the session devices can respond in a number of ways:
[0089] • In one aspect, the master device can search for other devices in the session that can act as master devices (excluding dumb devices). If such a device is found, a queue migration can be performed to move the queue management responsibility to another capable device.
[0090] • In another aspect, the master device can retain the queue management responsibility. It can mute itself to simulate being removed from the session but still perform queue management operations.
[0091] Figure 12 A communication flow 1200 for session 300 Figure 3 ) management according to another aspect of the disclosure is shown.Figure 12 In this scenario, the application discovers that queue management is performed by the first device (here, rendering device 110.2), and the user attempts to add a device (rendering device 110.3) that operates according to access control permissions.
[0092] In communication flow 1200, an add request is entered at rendering device 110.1, instructing rendering device 110.3 to be added to the session; the add request (message 1210) is transmitted to the queue manager, rendering device 110.2. In response, rendering device 110.2 can send a join request message (message 1220) to rendering device 110.3. Rendering device 110.3 can respond with a credential challenge (message 1230), which is sent to the queue manager, rendering device 110.2. In response, rendering device 110.2 can pass the credential challenge to rendering device 110.1, from which it receives the add request (message 1240). Rendering device 110.1 can provide credentials to the queue master device 110.2 (message 1250), and the queue master device 110.2 relays the credentials to the rendering device 110.3 to be added. Assuming the credentials are accepted (box 1270), the rendering device 110.3 to be joined can convey a response message granting the joining request (message 1280). Playback synchronization 1290 can be performed between rendering devices 110.2 and 110.3.
[0093] The principles of this disclosure enable application discovery in a network environment, where associations between devices can be created based on device location. In residential applications, a player device can be associated with a single room within a house (e.g., kitchen, living room, etc.). In commercial applications, a player device can be associated with a separate meeting room, office, etc. Portable devices often operate according to protocols that automatically discover nearby devices. In such applications, the identifier of such devices can be automatically populated in the control area 910 to allow operators to add and remove devices during a session.
[0094] In one aspect, devices can be configured to be automatically added to and removed from a session. For example, an audio player in a smartphone can automatically establish a session with an embedded player in a car (e.g., because the car is started). The playback session can then automatically render media via the car's audio system upon joining the session. If the smartphone detects that the connection with the embedded audio player has been lost due to, for example, the car being turned off, it may disassemble the session. Furthermore, if / when it detects other devices later in the playback, such as home audio components, it can add those devices. Thus, the principles of this disclosure can produce a playback experience that makes playback "follow" and the operator become someone who experiences it in their daily life.
[0095] Figure 13 An example of a session 300 according to one aspect of this disclosure is shown. Figure 3 Method 1300 for management. Method 1300 may be invoked in response to a command that identifies an action to be taken on a target device (such as "play jazz in the kitchen"). In response to such a command, method 1300 will classify the device that is the target of the command (e.g., a kitchen media player). The target device may be classified as a smart device or a dumb device.
[0096] When the target device is categorized as a "smart device", method 1300 determines whether the target device is a member of the playback group (box 1315). If not, the target device can retrieve the selected content and begin rendering (box 1320).
[0097] If the target device is determined to be a member of the playback group at box 1315, method 1300 may determine the target device's role as the group's master device, a group's secondary device, or a group's "silent master device" (box 1325). When the target device is classified as a secondary device, method 1300 may cause the target device to remove itself from the group (box 1330) and become its own master device in the new group. Thereafter, method 1300 may proceed to box 1320, and the target device may render the identified content.
[0098] If the target device is classified as the master device at box 1325, method 1300 may split the group (box 1335). Splitting the group removes the target device from the previously defined group and allows the previously defined group to continue its previous actions. Method 1300 may perform a queue migration, under which some other devices in the previously defined group may assume the role of master device and continue the group's management and its rendering operations. The target device may become its own master device in a new group (initially formed only by the target device). Method 1300 may proceed to box 1320, and the target device may render the identified content.
[0099] If the target device is classified as a silent master device at box 1325, method 1300 may cause the silent master device to interrupt the playback of the group (box 1345). Method 1300 may proceed to box 1320, and the target device may render the identified content.
[0100] If the target device is classified as a dumb device at box 1310, method 1300 determines whether the target device is currently playing content (box 1350). If so, method 1300 causes the target device to stop playback (box 1355). Thereafter, or if the target device is not playing content, method 1300 may designate another device within the target device's communication range as operating as a "silent master" on behalf of the target device (box 1360). The silent master device may retrieve the content identified in the command and stream the retrieved content to the target device for rendering (box 1365). The silent master device does not need to locally play the retrieved content via its own output; in fact, the silent master device may play content different from what will be played by the target device in response to the command.
[0101] exist Figure 13 During the operation of Method 1300, classifying a device as a "smart device" or a "dumb device" may be based on identification of the device's capabilities through its playback capabilities, its usage rights, or a combination thereof. A target device may be classified as a dumb device when it is unable to retrieve and play the identified content on its own. For example, the target device may not be an internet-enabled device and therefore will not be able to download content from internet-based media services. Alternatively, even if the target device is able to download content from such services, it may not have the account information or other authentication information required to obtain access to such services. In such scenarios, the target device may be classified as a dumb device. Alternatively, if the target device has the capability to download and render the identified content, the target device may be classified as a smart device.
[0102] Queue migration operations may not be available on all devices. In such use cases, method 1300 may interrupt playback of a group (not shown) in which the target device is a member. Alternatively, method 1300 may also begin playback of the identified stream on the target device (which becomes its own master device) and may also cause the target device to become the silent master device of other devices in its previous group (operation not shown).
[0103] Method 1300 ( Figure 14 Find applications that are suitable for various usage scenarios and have a variety of devices. Figure 14 and Figure 15 Two such use cases are shown. Figure 14 This illustrates a scenario where an operator presents verbal commands to a playback device 1400 that may or may not be the target device. Figure 15 This illustrates another use case where an operator presents commands to a control device 1500 (shown as a smartphone in this example) via touchscreen input; Figure 15 In the example, control device 1500 is not the target device.
[0104] In one aspect, a control method may be provided to identify a target device and direct commands to a device that can provide control over the target device (e.g., the target device itself or the target device's master device). One such control method is shown in... Figure 13 In this process, method 1300 may determine whether a command is received at the target device (box 1365). If yes, method 1300 may proceed to box 1315 and participate in the operations described above. If no, method 1300 may determine the state of the target device (box 1370), for example, whether it participates in a group with a master device. Method 1300 may relay the command to the target device or the target device's master device (box 1375). For example, if the target device participates in a group, method 1300 may relay the command to the target device's master device. If the target device can act as its own master device, method 1300 may relay the command to the target device itself.
[0105] Figure 16 An alternative method for session 300 according to another aspect of this disclosure is shown. Figure 3 Method 1600 manages applications where a command is directed to a location that may have several target devices (e.g., "Play jazz in the kitchen" when several devices are present in the kitchen). Method 1600 identifies the target device to which the command is directed and then determines whether the target device has been configured to belong to a stereo pair (box 1610). If so, any device paired with the target device emphasized by the command can be evaluated under Method 1600.
[0106] Method 1600 identifies whether the target device is a member of any currently operating rendering group (box 1620). It determines whether any target device is a member of such a group (box 1625). If not, method 1600 determines whether the target device is available to be used as its own master device (box 1630). If the target device is available to be used as a master device, method 1600 makes the target device the master device of the new group (box 1635), and the target device retrieves the content identified by the command and begins rendering that content (box 1640). In cases where the target device is paired with another device, the devices can negotiate together to designate one as the master device and the other as the auxiliary device.
[0107] At box 1630, if the target device cannot be used as the master device, method 1600 may find another device to act as the silent master device for the target device (box 1645). When a silent master device is available, method 1600 may cause the silent master device to retrieve the identified content and push the rendered data of the content to the target device (box 1650). Although in Figure 16If no silent master device is found for the identified target device, as not shown in the diagram, method 1600 will return an error in response to the command.
[0108] At box 1625, if the target device is a member of the group, the method determines whether the group contains devices other than the target device (box 1655). If not, the target device group can begin rendering the identified content (box 1640).
[0109] If the target device is a member of a group containing the device pointed to by the command (box 1655), method 1600 may split the target device from the old group (box 1660). Method 1600 may determine whether the target device is the master device of the existing group (box 1665). If the target device is the master device of the existing group, method 1600 may perform queue migration for devices that are members of the existing group from which the target device was split (box 1670). If successful, the previously rendered events can continue using other devices that are members of the existing group.
[0110] Splitting the group (box 1660) will result in the formation of a new group using the target device. Method 1600 can proceed to box 1630 and perform the operations of boxes 1630-1650 to begin rendering operations for the new group.
[0111] As discussed above, a device may operate as the master device for a given group based on its operating parameters, such as having appropriate account information to render the identified content, being online powered (rather than battery powered), the quality of the network connection, and the device type (e.g., when rendering audio, a speaker device may take precedence over other types of media players). When identifying whether a target device is suitable to be used as a master device (box 1630), each target device within the group may be evaluated based on a comparison of the device's capabilities required to render the content. When no target device within the group is suitable to be used as a master device, other devices that are not members of that group may also be evaluated as silent master devices based on a comparison of the device's capabilities required to render the identified content, and the capabilities of candidate master devices to communicate with target devices in a new group may also be evaluated. In some cases, when no target device is found to be suitable to operate as a master device, and when no other device is found to be suitable to operate as a silent master device, method 1600 may terminate with an error message (not shown).
[0112] In some use cases, queue migration may fail (box 1670). In this case, method 1600 will cause devices in the existing group—any devices that will not join the target device in the new group—to stop playing (operation not shown). In another alternative, also not shown, method 1600 may use devices that were previously part of an existing group and will not join the target device in the new group to perform the operations of boxes 1630-1650 in parallel processing (operation not shown).
[0113] Figure 17 This illustrates a session 300 that can be initiated in response to a command to stop playback at a target device (e.g., "stop playback in the kitchen"), as shown in one aspect of this disclosure. Figure 3 Method 1700 manages the device. Similar to existing methods, this command does not require input at the target device. Method 1700 also identifies the target device for the command and then determines whether the target device is paired with other devices (box 1710). If paired, all paired devices are considered target devices for the purposes of method 1700. Method 1700 may determine the group to which the target device belongs (box 1730) and may classify the target devices (box 1740).
[0114] If the target device is a master device without auxiliary devices, method 1700 will stop playback on the target device (box 1750).
[0115] If the target device is a master device with auxiliary devices, method 1700 performs a queue migration (box 1760) on the playback content to establish another device as the master device. Thereafter, method 1700 may stop playback on the target device, as shown in box 1750. As in the previous embodiments, if the queue migration fails for any reason, method 1700 may optionally stop playback on all devices in the current group (operation not shown).
[0116] If the target device is an auxiliary device, method 1700 will remove the target device from its current group (box 1770) and stop playback on the target device, as shown in box 1750.
[0117] If the target device is a silent master device, method 1700 can keep the target device silent (box 1780).
[0118] Figure 18 An example of a session 300 according to one aspect of this disclosure is shown. Figure 3 Another method of management 1800. Method 1800 can discover the application when a command is entered, which identifies the desired content in relative terms (e.g., "add this music to the kitchen", where the music is not directly identified). Figure 18Method 1800 can be used to identify the content referenced by the command.
[0119] Method 1800 can begin by determining whether the device to which the command was input (referred to as the "command device" for convenience) is playing content (box 1810). If so, method 1800 can designate the command device as a host group for the purposes of this method. This "host group" is the group that the target device will ultimately join for rendering purposes.
[0120] If, at box 1810, method 1800 determines that the command device is not playing, then method 1800 will identify the host group in an alternative manner (box 1830). In one aspect, method 1800 may determine how many groups are currently involved in the playback operation (box 1832). If only one group is identified, that group can be designated as a host group (box 1834). If multiple groups are identified, method 1800 may designate one of those groups as a host group based on ranking criteria such as audio proximity to the command device (and therefore the user), Bluetooth proximity to the command device, data representing the physical layout from the device to the command device, and / or heuristics such as the group that most recently started playback or the group that most recently received user interaction.
[0121] Once the host group is specified, method 1800 determines the target device to be added to the host group (box 1840). Similarly, the target device (the kitchen appliance in this example) can be identified. Method 1800 can determine if the target device is paired with any other device, such as via a stereo pair (box 1842). If so, the paired device is also designated as the target device (box 1844). Method 1800 then adds the target device to the host group (box 1850).
[0122] The principles of this disclosure extend to other use cases. For example, a command such as "move music to the kitchen" can be executed as a removal operation to remove a target device from one playback group and an addition operation to add the target device to another playback group. Therefore, the techniques disclosed in the above embodiments can be cascaded to provide more complex device management features.
[0123] This disclosure utilizes a computer-based virtual assistant service that responds to voice command input. In one aspect, the virtual assistant can associate user commands in a manner that develops command contexts, allowing the virtual assistant to associate, for example, device- or media-independent commands with target device or media content.
[0124] Figure 19 This is a network diagram of an exemplary system 1900, in which a virtual assistant service can be used to manage session 300 ( Figure 3System 1900 includes multiple user devices 1910.1 to 1910.3 that communicate with server 1920 via communication network 1930. In this example, virtual assistant service 1940 is shown to be provided by server 1920. In an alternative implementation, virtual assistant service 1940 may be provided by one of user devices 1910.1 to 1910.3, or it may be distributed across multiple devices 1910.1 to 1910.3. The implementation details regarding the placement of virtual assistant 1940 within system 1900 are expected to be tailored to the needs of individual applications.
[0125] The virtual assistant 1940 of the present invention may include a voice process 1942, a natural language process 1944, and a flow control process 1946. The voice process 1942 receives audio data representing a spoken command input from one of the devices 1910.1 to 1910.3, and may generate a text representation of the speech. The natural language process 1944 determines the intent of the text representation from the audio. The flow control process 1946 generates conversational commands from the intended data generated by the natural language process 1944. These commands may be output to devices 1910.1 to 1910.3 to achieve conversational variations consistent with the spoken commands.
[0126] During operation, user commands can be entered on any device that accepts verbal input, and these commands are integrated into a system for session control. Therefore, commands can be entered on smartphone devices 1910.1, smart speaker devices 1910.2, media players 1910.3, etc. These devices can capture audio representing the verbal commands and relay the audio to device 1920, which operates as a virtual assistant 1940. The virtual assistant can parse the audio into text and further parse it into session control commands, which can be output to devices 1910.1 to 1910.3 to manage playback sessions.
[0127] It is anticipated that users will input verbal commands targeting devices different from those they are interacting with. Thus, a user could input a command to smartphone 1910.1 to change playback at speaker 1910.2 or digital media player 1910.3. At another time, a user could input a command at speaker 1910.2 in one room (not shown) to change playback at another speaker (not shown) in another room. Finally, it is anticipated that users can input commands that do not explicitly provide all the information needed to achieve a change in playback. For example, a user could input a command that is not specific to the media to be presented or the device targeted by the command.
[0128] This disclosure develops "context" data that identifies the device and / or media as the subject of user commands. The context can change over time in response to operator commands. When a new command with altered playback is received, the virtual assistant 1940 will refer to the currently developed context to determine the subject of the command.
[0129] The following example illustrates how to develop a context. In this example, the user can enter the following command:
[0130] Command 1: "Play jazz music in the kitchen."
[0131] Command 2: Pause
[0132] The first command identifies the media item to be played (a jazz playlist) and the target device (a kitchen speaker). Therefore, the virtual assistant 1940 can generate session commands to make the kitchen device display the identified media item. The virtual assistant 1940 can also store data identifying the context, including the kitchen device.
[0133] Upon receiving command 2, the virtual assistant 1940 can determine that the command is neither media-specific nor target device-specific. The virtual assistant 1940 can refer to the context currently developed for the user and identify the target device based on that context. Thus, the virtual assistant 1940 can generate a session command to pause playback on the kitchen device.
[0134] User commands can lead to the expansion of currently developed groups. For example, consider the following command sequence:
[0135] Command 1: "Play jazz music in the kitchen."
[0136] Command 2: "Play this song in the living room."
[0137] As discussed above, the first command can develop a context group that includes the kitchen playback device. The second command (especially if entered at a kitchen device) can expand the context group to include target devices from other locations. Virtual Assistant 1940 can recognize the context of the current playback, where "this song" refers to the jazz playlist presented in the kitchen, and can add living room devices to the group that is playing the jazz playlist.
[0138] Furthermore, user commands can cause the context group to collapse. Consider the following command sequence:
[0139] Command 1: "Play classical music in all spaces"
[0140] Command 2: "Stop playing in the kitchen"
[0141] The first command can create a context group that includes all devices adjacent to the device in which the command is entered (e.g., all devices in a house). The second command can cause the target device to be removed from the group. In this example, the virtual assistant 1940 will remove the kitchen device from the session, but allow the other devices in the group to continue playing. The virtual assistant 1940 can also cause the kitchen device to be removed from the context. In this respect, if another command (e.g., "play jazz") is received, the virtual assistant 1940 can then designate the device as the target of a later received command within the context (all devices except the kitchen device).
[0142] On the other hand, a command that causes the session group to shrink can result in the context group being cleared. Therefore, in the aforementioned example, upon receiving command 2, the virtual assistant could remove the kitchen device from the session and clear the context. In this respect, commands received later (e.g., "play jazz") can be processed even when the virtual assistant has no available activity context. For example, the command could be interpreted as being directed to the local device where the command was entered. Therefore, virtual assistant 1940 would cause a jazz playlist to be presented on the local device.
[0143] Continuing this example, the virtual assistant can also check the playback status of the local device that will start playing jazz. If the local device is part of a session group (e.g., a collection of remaining devices playing classical music), the virtual assistant can switch all devices in the session group to the new media item. If the local device is not part of a session group (e.g., the kitchen device removed from the session group in command 2), the virtual assistant can create a new session group using the kitchen device. In either case, the set of devices identified as the target device by the virtual assistant can become the new context for processing new user commands.
[0144] Similarly, a command that causes playback to stop in a large area (e.g., "Stop in all spaces") may result in the context being cleared. In this case, a later received "Play Jazz" command will not reference the context "all spaces" but will be resolved differently, for example, by playing jazz on the local device where the command was entered.
[0145] In one aspect, devices identified as members of a stereo pair (e.g., left speaker and right speaker) can be added to and removed from the context as units rather than individual devices.
[0146] The above discussion addressed contexts, including commands targeting devices whose target devices are not explicitly identified as such. Alternatively, contexts can be developed for media items.
[0147] Consider the following example:
[0148] Command 1: "Play jazz music in the kitchen."
[0149] Command 2: "Play here".
[0150] For command 2, the target device is identified by its position relative to the user—the target device is the device on which the user inputs the verbal command. However, the media item was not identified.
[0151] In one aspect, the virtual assistant can also develop context for recognizing media (such as media items or media playlists) based on previous user commands. In the aforementioned example, command 1 identifies a jazz playlist as media to be presented on the kitchen device. When processing command 2, the context (jazz) can provide recognition of media the user expects to play on the new device.
[0152] Consider the following example:
[0153] Command 1: "Play jazz music".
[0154] Command 2: "Play it in the kitchen."
[0155] Here, similarly, command 1 identifies the media playlist to be played on the local device. Command 1 also provides context for subsequent commands. Upon receiving command 2, virtual assistant 1940 can identify the media playlist of the command based on the context (jazz). In this example, virtual assistant 1940 will issue a session command to add the kitchen device to the group playing the jazz playlist.
[0156] Contexts can be developed for different users interacting with the virtual assistant. Therefore, in residential applications, the virtual assistant can develop contexts for individual family members. When a user enters a new command, the virtual assistant can refer to the context developed for that user to identify the target device and / or media item to apply the command to.
[0157] Figure 20A A block diagram of a virtual assistant system 2000 according to various examples is shown. In some examples, the virtual assistant system 2000 may be implemented on a standalone computer system. In some examples, the virtual assistant system 2000 may be distributed across multiple computers. In some examples, some modules and functions of the virtual assistant may be divided into a server part and a client part, wherein the client part resides on one or more user devices 1910.1-1910.3 and communicates with the server part (e.g., server 1920) via one or more networks, for example, as... Figure 19 As shown. In some examples, Virtual Assistant System 2000 can be... Figure 19The specific implementation of the server system 1920 shown is illustrated. It should be noted that the virtual assistant system 2000 is merely an example of a virtual assistant system, and the virtual assistant system 2000 may have more or fewer components than shown, may combine two or more components, or may have different configurations or arrangements of components. Figure 20A The various components shown may be implemented in hardware, software instructions for execution by one or more processors, firmware (including one or more signal processing integrated circuits and / or application-specific integrated circuits), or a combination thereof.
[0158] The virtual assistant system 2000 may include a memory 2002, one or more processors 2004, an input / output (I / O) interface 2006, and a network communication interface 2008. These components may communicate with each other via one or more communication buses or signal lines 2010.
[0159] In some examples, memory 2002 may include non-transitory computer-readable media, such as high-speed random access memory and / or non-volatile computer-readable storage media (e.g., one or more disk storage devices, flash memory devices or other non-volatile solid-state memory devices).
[0160] In some examples, I / O interface 2006 can couple input / output devices 2016 of the virtual assistant system 2000, such as a display, keyboard, touchscreen, and microphone, to user interface module 2022. I / O interface 2006, combined with user interface module 2022, can receive user input (e.g., voice input, keyboard input, touch input, etc.) and process this input accordingly. In some examples, such as when the virtual assistant is implemented on a standalone user device, virtual assistant system 2000 may include interfaces 1910.1-1910.3 that facilitate connection to devices 1910.1-1910.3. Figure 19 The virtual assistant system 2000 may refer to either the communication component or the I / O communication interface. In some examples, the virtual assistant system 2000 may represent the server portion 1920 of the virtual assistant implementation. Figure 19 And can be located in user equipment (e.g., equipment 1910.1-1910.3). Figure 19 The client-side portion of the interface interacts with the user.
[0161] In some examples, the network communication interface 2008 may include one or more wired communication ports 2012 and / or wireless transmission and reception circuitry 2014. The one or more wired communication ports may receive and transmit communication signals via one or more wired interfaces such as Ethernet, Universal Serial Bus (USB), FireWire, etc. The wireless circuitry 2014 may receive RF signals and / or optical signals from the communication network and other communication devices, and transmit RF signals and / or optical signals to the communication network and other communication devices. Wireless communication may use any of a variety of communication standards, protocols, and technologies, such as GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol. The network communication interface 2008 enables the virtual assistant system 2000 to communicate with other devices via networks such as the Internet, intranets, and / or wireless networks (such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs)).
[0162] In some examples, memory 2002 or its computer-readable storage medium may store programs, modules, instructions, and data structures, including all or a subset of the following: operating system 2018, communication module 2020, user interface module 2022, one or more applications 2024, and virtual assistant module 2026. Specifically, memory 2002 or its computer-readable storage medium may store instructions for executing procedures. One or more processors 2004 may execute these programs, modules, and instructions, and read data from or write data to data structures.
[0163] Operating system 2018 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, WINDOWS, or embedded operating systems such as VxWorks) may include various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitate communication between various hardware, firmware, and software components.
[0164] The communication module 2020 facilitates communication between the virtual assistant system 2000 and other devices via the network communication interface 2008. For example, the communication module 2020 can communicate with electronic devices (such as devices 1910.1-1910.3). Figure 19 The communication module 2020 may also include various components for processing data received by the wireless circuit 2014 and / or the wired communication port 2012.
[0165] The user interface module 2022 can receive commands and / or input from a user (e.g., from a keyboard, touchscreen, pointing device, controller, and / or microphone) via the I / O interface 2006 and generate user interface objects on the display. The user interface module 2022 can also prepare output (e.g., voice, sound, animation, text, icons, vibration, haptic feedback, lighting, etc.) and deliver the output to the user via the I / O interface 2006 (e.g., through a display, audio channel, speaker, touchpad, etc.).
[0166] Application 2024 may include programs and / or modules configured to be executed by one or more processors 2004. For example, if the virtual assistant system is implemented on a standalone user device, application 2024 may include user applications such as games, calendar applications, navigation applications, or email applications. If the virtual assistant system 2000 is on server 1920 ( Figure 19 If implemented on ), then Application 2024 may include, for example, resource management applications, diagnostic applications, or scheduling applications.
[0167] The memory 2002 may also store the virtual assistant module 2026 (or the server portion of the virtual assistant). In some examples, the virtual assistant module 2026 may include the following sub-modules or subsets or supersets thereof: input / output processing module 2028, speech-to-text (STT) processing module 2030, natural language processing module 2032, dialogue flow processing module 2034, task flow processing module 2036, service processing module 2038, and speech synthesis module 2040. Each of these modules may have access to one or more, or subsets or supersets of, the following systems or data and models of the virtual assistant module 2026: knowledge ontology 2060, vocabulary index 2044, user data 2048, task flow model 2054, service model 2056, and ASR system.
[0168] In some examples, using the processing modules, data, and models implemented in the Virtual Assistant Module 2026, the Virtual Assistant may perform at least some of the following: converting voice input into text; recognizing user intent expressed in natural language input received from the user; proactively eliciting and obtaining the information needed to fully infer the user intent (e.g., by deambiguity of words, names, intents, etc.); determining a task flow to satisfy the inferred intent; and executing the task flow to satisfy the inferred intent.
[0169] In some examples, such as Figure 20B As shown, the I / O processing module 2028 can... Figure 20A The I / O devices in 2016 interact with the user or through Figure 20A The network communication interface 2008 in the system is connected to user equipment (e.g., Figure 19The I / O processing module 2028 interacts with devices 1910.1-1910.3 to receive user input (e.g., voice input) and provides a response to the user input (e.g., as voice output). The I / O processing module 2028 may optionally acquire contextual information associated with the user input from the user device, either along with or shortly after receiving the user input. Contextual information may include user-specific data, vocabulary, and / or preferences associated with the user input. In some examples, the contextual information may also include the software and hardware state of the user device at the time the user request is received, and / or information related to the user's surrounding environment at the time the user request is received. In some examples, the I / O processing module 2028 may also send follow-up questions related to the user request to the user and receive responses from the user. When a user request is received by the I / O processing module 2028 and the user request may include voice input, the I / O processing module 2028 may forward the voice input to the STT processing module 2030 (or a speech recognizer) for speech-to-text conversion.
[0170] STT processing module 2030 may include one or more ASR systems. The one or more ASR systems can process speech input received through I / O processing module 2028 to produce recognition results. Each ASR system may include a front-end speech preprocessor. The front-end speech preprocessor can extract representational features from the speech input. For example, the front-end speech preprocessor can perform a Fourier transform on the speech input to extract spectral features characterizing the speech input as a sequence of representational multidimensional vectors. Furthermore, each ASR system may include one or more speech recognition models (e.g., sound models and / or language models) and can implement one or more speech recognition engines. Examples of speech recognition models may include hidden Markov models, Gaussian mixture models, deep neural network models, n-gram language models, and other statistical models. Examples of speech recognition engines may include engines based on dynamic time warping and engines based on weighted finite-state transducers (WFST). One or more speech recognition models and one or more speech recognition engines can be used to process the extracted representational features from a front-end speech preprocessor to produce intermediate recognition results (e.g., phonemes, phoneme strings, and subwords) and ultimately produce text recognition results (e.g., words, word strings, or sequences of symbols). In some examples, the speech input may be processed at least partially by a third-party service or on the user's device (e.g., Figure 19 The STT processing module 2030 processes the text into recognition results on devices 1910.1-1910.3. Once the STT processing module 2030 generates recognition results containing text strings (e.g., words, or sequences of words, or sequences of symbols), the recognition results can be transmitted to the natural language processing module 2032 for intent inference.
[0171] In some examples, the STT processing module 2030 may include a vocabulary of recognizable words, and / or be accessible via the speech-to-letter conversion module 2031. Each vocabulary word may be associated with one or more candidate pronunciations of a word represented in a speech recognition alphabet. Specifically, the vocabulary of recognizable words may include words associated with multiple candidate pronunciations. For example, the vocabulary may include words related to... and The candidate pronunciations are associated with the word "tomato". Furthermore, vocabulary words can be associated with custom candidate pronunciations based on previous speech input from the user. Such custom candidate pronunciations can be stored in the STT processing module 2030 and can be associated with a specific user via a user profile on the device. In some examples, candidate pronunciations of words can be determined based on the spelling of the word and one or more linguistic and / or phonetic rules. In some examples, candidate pronunciations can be generated manually, for example, based on known canonical pronunciations.
[0172] In some examples, candidate pronunciations can be ranked based on their prevalence. For example, candidate speech... The ranking can be higher than Because the former is a more commonly used pronunciation (e.g., among all users, for users in a specific geographic region, or for any other suitable subset of users). In some examples, candidate pronunciations can be ranked based on whether they are custom candidate pronunciations associated with a user. For example, custom candidate pronunciations can rank higher than standard candidate pronunciations. This can be used to identify proper nouns with unique pronunciations that deviate from the canonical pronunciation. In some examples, candidate pronunciations can be associated with one or more phonological features, such as geographic origin, country, or ethnicity. For example, candidate pronunciations... Possibly associated with the United States, and the candidate pronunciation It may be associated with the United Kingdom. Furthermore, the ranking of candidate pronunciations can be based on one or more characteristics of the user (e.g., geographic origin, country, ethnicity, etc.) stored in the user profile on the device. For example, it can be determined from the user profile that the user is associated with the United States. Based on the user's association with the United States, candidate pronunciations can be ranked... (Related to the United States) ranked higher than candidate pronunciations (Related to the UK) Higher. In some examples, one of the ranked candidate pronunciations can be selected as the predicted pronunciation (e.g., the most likely pronunciation).
[0173] When a voice input is received, the STT processing module 2030 can be used (e.g., using a sound model) to determine the phonemes corresponding to the voice input, and then attempt (e.g., using a language model) to determine the words that match those phonemes. For example, if the STT processing module 2030 can first identify a sequence of phonemes corresponding to a portion of the voice input... Then it can then determine that the sequence corresponds to the word "tomato" based on the vocabulary index 2044.
[0174] In some examples, the STT processing module 2030 can use fuzzy matching techniques to determine words in a utterance. Therefore, for example, the STT processing module 2030 can determine phoneme sequences. This corresponds to the word "tomato," even if the specific phoneme sequence is not a candidate phoneme sequence for that word.
[0175] In some examples, the Natural Language Processing (NLP) module 2032 may be configured to receive metadata associated with the voice input. The metadata may indicate whether NLP should be performed on the voice input (or a sequence of words or symbols corresponding to that voice input). If the metadata indicates that NLP will be performed, the NLP module may receive the sequence of words or symbols from the STT processing module to perform NLP. However, if the metadata indicates that NLP will not be performed, the NLP module may be disabled, and the sequence of words or symbols (e.g., a text string) from the STT processing module may be output from the virtual assistant. In some examples, the metadata may further identify one or more domains corresponding to a user request. Based on these one or more domains, the NLP processor may disable domains in the knowledge ontology 2060 other than these one or more domains. Thus, the NLP is constrained to these one or more domains in the knowledge ontology 2060. Specifically, these one or more domains in the knowledge ontology, instead of other domains, may be used to generate structured queries (described below).
[0176] The virtual assistant's natural language processing module 2032 ("natural language processor") can acquire a sequence of words or symbols ("symbol sequence") generated by the STT processing module 2030 and attempt to associate the symbol sequence with one or more "executable intentions" recognized by the virtual assistant. An "executable intention" can represent a task that can be performed by the virtual assistant and may have an associated task flow implemented in the task flow model 2054. An associated task flow can be a series of programmed actions and steps taken by the virtual assistant to perform the task. The capabilities of the virtual assistant may depend on the number and type of task flows implemented and stored in the task flow model 2054, or in other words, on the number and type of "executable intentions" recognized by the virtual assistant. However, the effectiveness of the virtual assistant may also depend on its ability to infer the correct "one or more executable intentions" from a user request expressed in natural language.
[0177] In some examples, in addition to the sequence of words or symbols obtained from the STT processing module 2030, the natural language processing module 2032 may also receive (e.g., from the I / O processing module 2028) contextual information associated with the user request. The natural language processing module 2032 may optionally use the contextual information to clarify, supplement, and / or further define the information included in the sequence of symbols received from the STT processing module 2030. Contextual information may include, for example, user preferences, the hardware and / or software state of the user's device, sensor information collected before, during, or shortly after the user request, previous interactions (e.g., conversations) between the virtual assistant and the user, etc. As described herein, the contextual information can be dynamic and may vary with the time, location, content, and other factors of the conversation.
[0178] In some examples, natural language processing may be based on, for example, a knowledge ontology 2060. A knowledge ontology 2060 may be a hierarchical structure containing many nodes, each node representing an "executable intent" or an "attribute" related to one or more of the "executable intent" or other "attributes." As mentioned above, an "executable intent" may represent a task that the virtual assistant can perform, i.e., the task is "executable" or can be done. An "attribute" may represent a parameter associated with an executable intent or a sub-aspect of another attribute. The links between executable intent nodes and attribute nodes in knowledge ontology 2060 can define how the parameters represented by the attribute nodes are related to the task represented by the executable intent nodes.
[0179] In some examples, the knowledge ontology 2060 may consist of executable intent nodes and attribute nodes. Within the knowledge ontology 2060, each executable intent node may be directly linked to or linked to one or more attribute nodes through one or more intermediate attribute nodes. Similarly, each attribute node may be directly linked to or indirectly linked to one or more executable intent nodes through one or more intermediate attribute nodes.
[0180] An executable intent node, along with its linked conceptual nodes, can be described as a "domain". In this discussion, each domain may be associated with a corresponding executable intent and involves a set of nodes associated with a particular executable intent (and the relationships between these nodes). Each domain may share one or more attribute nodes with one or more other domains.
[0181] In some examples, knowledge ontology 2060 may include all domains (and therefore executable intents) that the virtual assistant can understand and act upon. In some examples, knowledge ontology 2060 may be modified, such as by adding or removing entire domains or nodes, or by modifying the relationships between nodes within knowledge ontology 2060.
[0182] Furthermore, in some examples, nodes associated with multiple related executable intents can be clustered under a “superdomain” in Knowledge Ontology 2060.
[0183] In some examples, each node in knowledge ontology 2060 may be associated with a set of words and / or phrases related to the attributes or executable intentions represented by the node. The corresponding set of words and / or phrases associated with each node may be referred to as the "vocabulary" associated with the node. The corresponding set of words and / or phrases associated with each node may be stored in a vocabulary index 2044 associated with the attributes or executable intentions represented by the node. The vocabulary index 2044 may optionally include words and phrases from different languages.
[0184] Natural Language Processing (NLP) module 2032 may receive a sequence of symbols (e.g., a text string) from STT processing module 2030 and determine which nodes the words in the sequence of symbols relate to. In some examples, if a word or phrase in the sequence of symbols (via lexical index 2044) is found to be associated with one or more nodes in knowledge ontology 2060, that word or phrase may “trigger” or “activate” those nodes. Based on the number and / or relative importance of the activated nodes, NLP module 2032 may select one executable intent from the executable intents as the task the user intends the virtual assistant to perform. In some examples, the domain with the most “triggered” nodes may be selected. In some examples, the domain with the highest confidence value (e.g., based on the relative importance of its various triggered nodes) may be selected. In some examples, the domain may be selected based on a combination of the number and importance of the triggered nodes. In some examples, additional factors, such as whether the virtual assistant has previously correctly interpreted similar requests from the user, are also considered in the node selection process.
[0185] User data 2048 may include user-specific information such as user-specific vocabulary, user preferences, user address, user's default and second languages, user's contact list, and other short- or long-term information for each user. In some examples, the natural language processing module 2032 may use user-specific information to supplement the information contained in the user input to further define the user's intent.
[0186] In some examples, once the natural language processing module 2032 identifies an executable intent (or domain) based on a user request, it can generate a structured query representing the identified executable intent. In some examples, the structured query may include parameters for one or more nodes within the domain of the executable intent, and at least some of these parameters may be populated with specific information and requirements specified in the user request. Based on a knowledge ontology, the structured query for a domain may include predetermined parameters. In some examples, based on voice input and text derived from the voice input using the STT processing module 2030, the natural language processing module 2032 can generate a partial structured query for a domain, where the partial structured query includes domain-related parameters. In some examples, the natural language processing module 2032 may utilize received contextual information to populate some parameters of the structured query, as discussed above.
[0187] In some examples, the natural language processing module 2032 may pass the generated structured query (including any completed parameters) to the task flow processing module 2036 (“task flow processor”). The task flow processing module 2036 may be configured to receive the structured query from the natural language processing module 2032, complete the structured query (if necessary), and perform the actions required to “complete” the user’s final request. In some examples, the various processes necessary to complete these tasks may be provided in the task flow model 2054. In some examples, the task flow model 2054 may include processes for obtaining additional information from the user, and task flows for performing actions associated with the executable intent.
[0188] In some use cases, to complete a structured query, the task flow processing module 2036 may need to initiate additional dialogue with the user to obtain additional information and / or clarify potentially ambiguous statements. When such interaction is necessary, the task flow processing module 2036 may invoke the dialogue flow processing module 2034 to participate in the dialogue with the user. In some examples, the dialogue flow processing module 2034 may determine how (and / or when) to request additional information from the user and receive and process the user's response. This question may be presented to the user via the I / O processing module 2028 and a response may be received from the user. In some examples, the dialogue flow processing module 2034 may present dialogue output to the user via audio and / or visual output and receive input from the user via verbal or physical (e.g., click) responses. Continuing with the above examples, when the task flow processing module 2036 invokes the dialogue flow processing module 2034 to determine parameter information for the structured query associated with the selected domain, the dialogue flow processing module 2034 may generate a question to be presented to the user. Once a response is received from the user, the dialogue flow processing module 2034 can either fill the structured query with the missing information or pass the information to the task flow processing module 2036 to complete the missing information based on the structured query.
[0189] Once the task flow processing module 2036 has completed a structured query for the executable intent, it can proceed to execute the final task associated with that intent. Therefore, the task flow processing module 2036 can execute the steps and instructions in the task flow model based on the specific parameters contained in the structured query.
[0190] In some examples, task flow processing module 2036 may, with the assistance of service processing module 2038 (“service processing module”), complete the task requested in the user input or provide the informational answer requested in the user input. In some examples, the protocols and application programming interfaces (APIs) required for each service may be specified through the corresponding service model in service model 2056. Service processing module 2038 may access the appropriate service model for a service and, based on the service model, generate a request for that service according to the protocols and APIs required by that service.
[0191] In some examples, the natural language processing module 2032, the dialogue processing module 2034, and the task flow processing module 2036 may be used jointly and repeatedly to infer and define the user's intent, obtain information to further clarify and refine the user's intent, and ultimately generate a response (i.e., output to the user or completion of a task) to satisfy the user's intent. The generated response may be a dialogue response to the voice input that at least partially satisfies the user's intent. Furthermore, in some examples, the generated response may be output as voice output. In these examples, the generated response may be sent to the speech synthesis module 2040 (e.g., a speech synthesizer), where the generated response may be processed to synthesize the dialogue response in speech form. In other examples, the generated response may be data content related to satisfying a user request in the voice input.
[0192] The speech synthesis module 2040 can be configured to synthesize speech output for presentation to a user. The speech synthesis module 2040 synthesizes speech output based on text provided by a virtual assistant. For example, the generated dialogue response may be in the form of a text string. The speech synthesis module 2040 can convert the text string into audible speech output. The speech synthesis module 2040 can use any suitable speech synthesis technique to generate speech output from text, including but not limited to: concatenation synthesis, unit selection synthesis, diphone synthesis, domain-specific synthesis, formant synthesis, articulation synthesis, Hidden Markov Model (HMM) based synthesis, and sine wave synthesis. In some examples, the speech synthesis module 2040 can be configured to synthesize individual words based on phoneme strings corresponding to these words. For example, phoneme strings may be associated with words in the generated dialogue response. Phoneme strings may be stored in metadata associated with the words. The speech synthesis module 2040 can be configured to directly process the phoneme strings in the metadata to synthesize words in speech form.
[0193] In some examples, instead of using the speech synthesis module 2040 (or otherwise), it can be used on remote devices (e.g., server system 1920). Figure 19Speech synthesis is performed on a server-side device, and the synthesized speech can be sent to the user device for output to the user. For example, this can occur in some implementations where the output of a virtual assistant is generated at a server system. And because server systems typically have greater processing power or more resources than user devices, it is possible to obtain higher quality speech output than would be achieved through client-side synthesis.
[0194] While the invention has been described in detail above with reference to some embodiments, variations in the scope and spirit of the invention will be apparent to those skilled in the art. Therefore, the invention should be considered limited only by the scope of the appended claims.
[0195] Figure 21 A system 2100 according to another aspect of this disclosure is illustrated. In this embodiment, rendering devices 2110.1-2110.n are provided, which communicate with server 2120 via communication network 2130. In one embodiment, server 2120 may be provided in the local area network where rendering devices 2110.1-2110.n reside, which may occur when server 2120 acts as a proxy device for the local network. In another embodiment, server 2120 may be provided at an Internet location, which may occur when server 2120 is integrated into an online service.
[0196] exist Figure 21 In the aspects shown, server 2120 can perform the session management operations described above. Therefore, server 2120 can store playback queue data and manage session data for rendering devices 2110.1-2110.n. Although server 2120 does not render media data as part of the playback, server 2120 can operate as a master device on behalf of the rendering devices 2110.1-2110.n participating in the playback. In this respect, server 2120 can act as a master device on behalf of many simultaneously active groups. Server 2120 can store data of the rendering devices 2110.1-2110.n registered to it in registry 2125.
[0197] Figure 22 This illustration shows a session 300, according to one aspect of the present disclosure, that can occur between rendering device 2210 and server 2220. Figure 3The server 2220 manages communication flow 2200 to register rendering device 2210 with the server 2200. Rendering device 2210 may send a registration message (msg. 2230) to the server 2200 to identify the device. The server 2220 will determine whether the device is a new device for which the server 2220 has not yet stored information (box 2232). If so, the server 2220 will send a request message (msg. 2234) to the rendering device 2210 to request information such as its capabilities, location, and account information. The rendering device 2210 may provide the requested information in a response message (msg. 2236), and the server 2220 will also store the rendering device in its registry (box 2238). If, at box 2232, the server 2220 determines that rendering device 2210 is not a new device, the server 2220 will mark rendering device 2210 as active in its registry (still at box 2238).
[0198] The principles of this disclosure can be found to have applications in self-organizing network environments, where individual rendering devices can turn on and off at indeterminate times and can gain and lose network connectivity at indeterminate times. Therefore, Figure 22 Method 2200 can be initiated by the rendering device as part of the standard power-on procedure or when it gains a network connection.
[0199] Figure 23 A session management method 2300 according to one aspect of this disclosure is illustrated. Method 2300 can be initiated when a user command (box 2310) is received at an input device (e.g., “Play jazz in the kitchen”). After receiving the command, the input device can report the command to server 2120. Figure 21 (Box 2315). At the server, method 2300 identifies the target device of the command from its registry data (Box 2320). For each target device thus identified, method 2300 takes appropriate steps from the operations shown in Boxes 2325 through 2385.
[0200] This method can determine the type of the target device (box 2325). If method 2300 classifies the target device as a smart device, then method 2300 will determine whether the target device is already a member of an active group (box 2330). If not, then method 2300 will cause the server to issue a command to the target device to start playback, thereby identifying the target device's playlist and its type in the group (box 2335).
[0201] If, at box 2330, method 2300 determines that the target device is already a member of the group, then method 2300 can determine the target device's role within the group (box 2340). If the target device is the master device, then method 2300 can determine to split the group (box 2345). Method 2300 can initiate a queue migration process for playing back the content of the already active group (box 2350), which may cause the server to issue commands (not shown) to other members of the already active group, designating one of these members as the new master device. Method 2300 may cause the server to issue commands to the target device to begin playback, identifying the target device's playlist and its role in the group (box 2335).
[0202] If, at box 2340, the method determines that the target device is acting as the mute master, then method 2300 can issue a command to the target device to interrupt playback on behalf of its already active group (box 2355). Method 2300 can also issue a command to another target device in an already active group, designating it as the new mute master (box 2360). Method 2300 can also cause the server to issue a command to the target device to start playback, identifying the target device's playlist and its role in the group (box 2335). In one aspect, the commands issued to the target device in boxes 2355 and 2335 can be merged into a public message or message group.
[0203] If, at box 2340, the method determines that the target device acts as an auxiliary device, then method 2300 can issue a command to the target device that assigns the target device to a new group (box 2365). Method 2300 can also cause the server to issue a command to the target device to start playback, identifying the target device's playlist and its role in the group (box 2335). Similarly, the commands issued to the target device in boxes 2365 and 2335 can be merged into a public message or message group.
[0204] If the method determines at box 2325 that the target device is a dumb device, then method 2300 can determine whether the target device is currently participating in playback (box 2370). If so, method 2300 can issue a command to the target device to stop playback (box 2375). Method 2300 can define a mute master for the target device 2380 (box 2380) and issue a command to the target device to identify its new master (box 2385). The commands issued to the target device in boxes 2375 and 2385 can be merged into a common message or message group. Method 2300 can issue a command to the designated master device 2385 to transmit the media stream to the target device (box 2385).
[0205] As described above, this disclosure proposes techniques for developing self-organizing rendering networks among multiple loosely connected rendering devices with various designs and functions. Devices can autonomously discover each other, and they can exchange communications to synchronize their rendering operations. These devices may have various functions. Some high-performance devices may have the ability to communicate with media sources, authenticate themselves as having access to media items rendered as part of a playback session, and autonomously download media items for playback. Other low-performance devices may not have such capabilities; these devices may operate in a "secondary" mode in a primary-secondary relationship, where these devices receive media content sent to them from another "primary" device without having, for example, the ability to authenticate themselves with media sources or autonomously download media items. Furthermore, rendering devices can control the rendering operations performed by the rendering network by accepting user input—via a touchscreen on a smartphone or computer, audio input captured by a microphone at a smartphone, computer, or smart speaker, etc.—through any of many different devices. It is expected that the various session management techniques disclosed herein can be adaptively applied in response to changing environmental guarantees between rendering sessions. For example, it is possible that a given session 300, when initiated, will be operated by the first rendering device (which operates as the primary device). Figure 1 ) is managed, however, over time, perhaps due to the removal of the first rendering device from the session, subsequent sessions will be managed by a server on the Internet. Figure 21 What may happen is that the session will be interrupted at a point during operation by communication stream 400 (…). Figure 4 ) management, which can be achieved at another point in its lifecycle through method 1300 ( Figure 13 It can be managed through method 2300 at another point in the lifecycle. Figure 23 This disclosure is designed to manage [the system]. The principles of this disclosure are adapted to this change.
[0206] The foregoing discussion identifies functional blocks that can be used in playback devices and servers constructed according to various embodiments of the invention. In some applications, the functional blocks described above can be provided as elements of an integrated software system, wherein the blocks can be provided as separate elements of a computer program, stored in memory and executed by a processing device. A non-transitory computer-readable medium may have program instructions for causing a computer to execute the functional blocks. In other applications, the functional blocks can be provided as discrete circuit components of a processing system, such as functional units within a digital signal processor or application-specific integrated circuit. Other applications of the invention can be embodied as hybrid systems of dedicated hardware and software components. Furthermore, it is not necessary to provide the functional blocks described herein as separate units. Unless otherwise stated, these implementation details are not essential to the operation of the invention.
[0207] As described above, one aspect of the present invention involves collecting and using data from various sources to improve the delivery of inspirational or other content that may be of interest to users. This disclosure contemplates that, in some instances, such collected data may include personal information that uniquely identifies or can be used to contact or locate specific individuals. Such personal information may include demographic data, location-based data, telephone numbers, email addresses, Twitter IDs, home addresses, data or records related to a user's health or health level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying information or personal information.
[0208] This disclosure recognizes that the use of such personal information data in the present invention can benefit users. For example, the personal information data can be used to deliver targeted content that is of interest to the user. Therefore, the use of such personal information data enables users to have planned control over the delivered content.
[0209] This disclosure assumes that entities responsible for collecting, analyzing, disclosing, transmitting, storing, or otherwise using such personal information data will comply with established privacy policies and / or privacy practices. Specifically, such entities should implement and adhere to privacy policies and practices recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy and security of personal information data. These policies should be readily accessible to users and should be updated as data collection and / or use change. Users' personal information should be collected for the legitimate and reasonable use of the entity and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should be conducted only after obtaining informed consent from users. In addition, such entities should consider taking any necessary steps to protect and safeguard access to such personal information data and ensure that other persons authorized to access personal information data comply with their privacy policies and processes. Additionally, such entities may be subject to third-party assessments to demonstrate their compliance with widely accepted privacy policies and practices. Furthermore, policies and practices should be adapted to the specific types of personal information data collected and / or accessed, and to applicable laws and standards, including specific considerations regarding jurisdiction. Therefore, different privacy practices should be maintained for different types of personal information data in each country.
[0210] Regardless of the foregoing, this disclosure also envisions implementation schemes for users to selectively prevent the use or access to personal information data. That is, this disclosure anticipates providing hardware and / or software components to prevent or block access to such personal information data. For example, with regard to advertising delivery services, the technology of this invention can be configured to allow users to opt-in or opt-out at any time during or after registration for the service to participate in the collection of personal information data. As another example, users may choose not to provide emotion-related data for a targeted content delivery service. In another example, users may choose to limit the length of time emotion-related data is retained, or completely prohibit the development of underlying emotional states. In addition to providing opt-in and opt-out options, this disclosure envisions providing notifications related to access to or use of personal information. For example, users may be notified when downloading an application that their personal information data will be accessed, and then reminded again before the personal information data is accessed by the application.
[0211] Furthermore, the purpose of this disclosure is to manage and process personal information data to minimize the risk of unintentional or unauthorized access or use. Once data is no longer needed, this risk can be minimized by restricting data collection and deleting data. Additionally, and where applicable, including in certain health-related applications, data deidentification can be used to protect user privacy. Where appropriate, deidentification can be facilitated by removing specific identifiers (e.g., date of birth, etc.), controlling the amount or characteristics of stored data (e.g., collecting location data at the city level rather than address level), controlling how data is stored (e.g., aggregating data among users), and / or other methods.
[0212] Therefore, while this disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, it also contemplates that various embodiments can be implemented without access to such personal information data. That is, various embodiments of the present invention will not be rendered inoperable due to the absence of all or part of such personal information data. For example, preferences can be inferred based on non-personal information data or an absolute minimum of personal information (e.g., content requested by a device associated with a user, other non-personal information available to the content delivery service, or publicly available information), thereby selecting content and delivering it to the user.
[0213] In order to help the Patent Office and any reader of any patent published in this application interpret the appended claims, the applicants wish to note that they do not intend any appended claim or claim element to reference 35 U.SC112(f) unless “means for…” or “steps for…” is expressly used in a particular claim.
[0214] This document specifically illustrates and / or describes several embodiments of the invention. However, it should be understood that modifications and variations of the invention are covered by the foregoing teachings and are within the scope of the appended claims without departing from the spirit and intended scope of the invention.
Claims
1. A method for multiple control groups, comprising: In response to a playback session control command received from the user at the first device in the current playback group of networked rendering devices participating in the playback session: The media session of the current playback group is identified based on the context of the previous command in the playback session, and Determine whether the received playback session control command references a playback device that is not the same as one or more devices in the current playback group participating in the playback session; When one or more devices in the current playback group are not the same as the referenced playback device: Identify the target device associated with the referenced playback device, and Based on the rendering capabilities of the identified target device, identify the asset component among multiple asset components in the media asset that will be rendered by the target device, wherein the asset component is a first representation of the media asset currently being played in the playback session. as well as The playback of the media asset on the target device is synchronized with playback on one or more other devices that are members of the playback session. The synchronization includes identifying one or more other asset components that will be rendered by the one or more other devices according to the playback session, and synchronizing the timing of the playback of the asset components with the timing of the one or more other asset components, which represent one or more other representations of the media asset.
2. The method according to claim 1, wherein, The command to play back media assets is a voice command.
3. The method of claim 1, wherein the media asset comprises representations of video from different perspectives, and a first representation is used for a first perspective, and the one or more other representations are used for one or more other perspectives.
4. The method according to claim 1, further comprising: Identify one or more additional target devices associated with the identified group; Add the target device and the one or more additional target devices to the playback session; as well as Remove devices that are not associated with the identified group from the playback session; Synchronizing playback includes synchronizing the target device with the one or more additional target devices.
5. The method of claim 2, wherein the playback queue for the playback session is stored on a first master device in the networked rendering device, and the operation of removing devices not associated with the identified group from the playback session removes the first master device, the method further comprising: The playback queue is migrated from the first master device to the second master device associated with the identified group.
6. The method according to claim 2, further comprising: Clear the context of the preceding command after removing devices that are not associated with the identified group.
7. The method according to claim 1, further comprising: When the target device is determined to be a dumb device: Identify the silent master device for the target device. The silent master device retrieves the asset components of the media assets for the identified target device, and The retrieved asset components are sent from the silent master device to the target device.
8. The method according to claim 1, further comprising: The ability of the target device to be used as a master device is determined based on the target device's Digital Rights Management (DRM) capabilities. When the target device cannot be used as the master device Identify the silent master device for the target device. The silent master device retrieves the asset component of the media asset identified for the target device, and The retrieved asset components are sent from the silent master device to the target device.
9. The method according to claim 1, further comprising: When the target device is a member of a stereo pair, playback synchronization includes both the target device and the other member of the stereo pair.
10. The method according to claim 1, wherein: The rendering capabilities of the identified target device include playback applications supported on the identified target device; as well as The identification of the asset components of the media asset is based on the playback application.
11. A playback device, comprising: Processing equipment, transceiver A memory system that stores program instructions that, when executed, cause the processing device to perform: In response to a playback session control command received from the user at the first device in the current playback group of networked rendering devices participating in the playback session: The media session of the current playback group is identified based on the context of the previous command in the playback session. Determine whether the received playback session control command references a playback device that is not the same as one or more devices in the current playback group participating in the playback session; When one or more devices in the current playback group are not the same as the referenced playback device: Identify the target device associated with the referenced playback device, and Based on the rendering capabilities of the identified target device, identify the asset component among multiple asset components in the media asset that will be rendered by the target device, wherein the asset component is a first representation of the media asset currently being played in the playback session. as well as The playback of the media asset on the target device is synchronized with playback on one or more other devices that are members of the playback session. The synchronization includes identifying one or more other asset components that will be rendered by the one or more other devices according to the playback session, and synchronizing the timing of the playback of the asset components with the timing of the one or more other asset components, which represent one or more other representations of the media asset.
12. The playback device of claim 11, wherein the command to play back the media asset is a voice command.
13. The playback device of claim 11, wherein the media asset includes representations of video from different viewpoints, and a first representation is used for a first viewpoint, and the one or more other representations are used for one or more other viewpoints.
14. The playback device of claim 11, wherein the instructions further cause the processing device to perform the following operations: Identify one or more additional target devices associated with the identified group; Add the target device and the one or more additional target devices to the playback session; and Remove devices that are not associated with the identified group from the playback session; Synchronizing playback includes synchronizing the target device with the one or more additional target devices.
15. The playback device of claim 12, wherein the playback queue for the playback session is stored on a first master device in the networked rendering device, and the operation of removing a device not associated with the identified group from the playback session removes the first master device, wherein the instructions further cause the processing device to perform the following operations: The playback queue is migrated from the first master device to the second master device associated with the identified group.
16. The playback device of claim 12, wherein the instructions further cause the processing device to perform the following operations: Clear the context of the preceding command after removing devices that are not associated with the identified group.
17. The playback device of claim 11, wherein the instructions further cause the processing device to perform the following operations: When the target device is determined to be a dumb device: Identify the silent master device for the target device. The silent master device retrieves the asset components of the media assets for the identified target device, and The retrieved asset components are sent from the silent master device to the target device.
18. The playback device of claim 11, wherein the instructions further cause the processing device to perform the following operations: The ability of the target device to be used as a master device is determined based on the target device's Digital Rights Management (DRM) capabilities. When the target device cannot be used as the master device Identify the silent master device for the target device. The silent master device retrieves the asset component of the media asset identified for the target device, and The retrieved asset components are sent from the silent master device to the target device.
19. The playback device of claim 11, wherein the instructions further cause the processing device to perform the following operations: When the target device is a member of a stereo pair, playback synchronization includes both the target device and the other member of the stereo pair.
20. The playback device according to claim 11, wherein: The rendering capabilities of the identified target device include playback applications supported on the identified target device; as well as The identification of the asset components of the media asset is based on the playback application.
21. A non-transitory computer-readable medium storing program instructions that, when executed by a processing device, cause the device to perform the following operations: In response to a playback session control command received from the user at the first device in the current playback group of networked rendering devices participating in the playback session: The media session of the current playback group is identified based on the context of the previous command in the playback session, and Determine whether the received playback session control command references a playback device that is not the same as one or more devices in the current playback group participating in the playback session; When one or more devices in the current playback group are not the same as the referenced playback device: Identify the target device associated with the referenced playback device, and Based on the rendering capabilities of the identified target device, identify the asset component among multiple asset components in the media asset that will be rendered by the target device, wherein the asset component is a first representation of the media asset currently being played in the playback session. as well as The playback of the media asset on the target device is synchronized with playback on one or more other devices that are members of the playback session. The synchronization includes identifying one or more other asset components that will be rendered by the one or more other devices according to the playback session, and synchronizing the timing of the playback of the asset components with the timing of the one or more other asset components, which represent one or more other representations of the media asset.
22. The computer-readable medium of claim 21, wherein the media asset comprises representations of video from different viewpoints, and a first representation is used for a first viewpoint, and the one or more other representations are used for one or more other viewpoints.
23. The computer-readable medium of claim 21, wherein the instructions further cause the device to perform the following operations: Identify one or more additional target devices associated with the identified group; Add the target device and the one or more additional target devices to the playback session; and Remove devices that are not associated with the identified group from the playback session; Synchronizing playback includes synchronizing the target device with the one or more additional target devices.
24. The computer-readable medium of claim 23, wherein the playback queue for the playback session is stored on a first master device in the networked rendering device, and the operation of removing a device not associated with the identified group from the playback session removes the first master device, the instructions further causing the device to perform the following operations: The playback queue is migrated from the first master device to the second master device associated with the identified group.
25. The computer-readable medium of claim 23, wherein the instructions further cause the device to perform the following operations: Clear the context of the preceding command after removing devices that are not associated with the identified group.
26. The computer-readable medium of claim 21, wherein the instructions further cause the device to perform the following operations: When the target device is determined to be a dumb device: Identify the silent master device for the target device. The silent master device retrieves the asset components of the media assets for the identified target device, and The retrieved asset components are sent from the silent master device to the target device.
27. The computer-readable medium of claim 21, wherein the instructions further cause the device to perform the following operations: The ability of the target device to be used as a master device is determined based on the target device's Digital Rights Management (DRM) capabilities. When the target device cannot be used as the master device Identify the silent master device for the target device. The silent master device retrieves the asset component of the media asset identified for the target device, and The retrieved asset components are sent from the silent master device to the target device.
28. The computer-readable medium of claim 21, wherein the instructions further cause the device to perform the following operations: When the target device is a member of a stereo pair, playback synchronization includes both the target device and the other member of the stereo pair.
29. The computer-readable medium of claim 21, wherein: The rendering capabilities of the identified target device include playback applications supported on the identified target device; as well as The identification of the asset components of the media asset is based on the playback application.