Access control of audio and video streams and control of representation of communication session
By providing side panel commands in the user interface, automatically controlling avatar movement and audio stream access, the problem of low switching efficiency and high security risks of collaborative systems in the prior art between multiple private packet sessions is solved, and more efficient and secure access management is achieved.
Patent Information
- Application Number
- CN202380086454.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-30
- Filing Date
- 2023-10-26
- Publication Date
- 2025-07-22
Smart Images

Figure CN120359726A_ABST
Abstract
Description
Background Art
[0001] There are various different types of collaboration systems that allow users to communicate. For example, some systems allow people to share content for collaboration by using video and audio streams, sharing files, chat messages, and the like. Some systems provide a user interface format that allows users to share content with an audience. Such systems can provide a specific set of permissions that allow users to assume specific roles, such as presenters, audience members, and the like.
[0002] Although some collaboration systems can provide a platform for multiple users to share live video and audio streams using a specific set of permissions for users to assume certain roles, such systems have multiple disadvantages. For example, when a communication session involves multiple different private breakout sessions, the user must take multiple different manual steps to search for the breakout session of interest and then take multiple manual steps to enter the individual breakout session. In an illustrative example, a meeting involving twenty people can have four breakout sessions: a first group of five people can participate in a private chat session discussing renovations, a second group of five people can participate in a private chat session discussing a new home, a third group of four people can participate in a private chat session discussing an office building, and a fourth group of six people can participate in a private chat session discussing a lease contract. To move an individual out of one group and into another private discussion, the person would have to take multiple manual steps to leave the group, search for another group of interest, and then enter that private discussion. These manual steps not only result in a lot of inefficiencies, but this manual process can also lead to security issues.
[0003] In some existing systems, security issues are created when a user joins a breakout discussion or breakout group. To join the group, the user may have to send a request to an administrator or group leader. The administrator may have to take multiple manual steps to change the permissions for the requesting user. Then, when the person leaves the group, those permissions may have to be changed back to the original state. This type of process involving manual entry to control access to files and control audio and video streams can lead to security issues because a person may make a mistake due to reversed input, or the permissions may be inadvertently left in an undesired state. Summary of the Invention
[0004] The technology disclosed herein provides features for managing a conference user interface for an event subgroup. Movement of an avatar or user representation and selective audio streaming in the user interface can be implemented in response to selecting a command corresponding to a particular subgroup (e.g., a "listen" command) from a list on a side panel. The disclosed technology includes various types of commands for controlling the movement of the avatar and controlling access to multiple selected audio streams for the user's computer. Controlling the visual representation of a user in a "listen" state and selective transmission of the corresponding audio signal in response to a particular command provided by the user, which includes but is not limited to voice instructions or pointer input selection of a discussion of a subgroup of a conference. The command can control the avatar and access to the signal when the command identifies: people in a group, a topic being discussed, a subgroup of people in a conference in a list, a reference to shared content for the user's subgroup, etc.
[0005] In response to the command, the avatar position can be moved from an original position to a second position near or within a graphical representation of the subgroup. In response to the command, the system also grants access to an audio stream generated by a computer of a subgroup member. When the user provides a second command (e.g., a leave discussion command), the avatar position can be moved back to the original position or out of the graphical representation of the subgroup. Additionally, in response to the second command, the system also revokes access to the audio stream generated by a computer of a subgroup member. Access to the stream can also control access to shared content (e.g., a file shared among people in the subgroup). Other commands cause the system to change access permissions for a video stream and target control of the audio stream. The operation state can be changed from a listen-only operation mode in which the audio stream can be unidirectionally transmitted to a full-join mode in which audio and video streams can be bidirectionally transmitted.
[0006] These features provide increased security by automatically controlling access permissions to shared content and streams. This eliminates the need for the user to provide requests or manual inputs to change access permissions and change the access permissions back to the original state. This can avoid situations where access permissions are inadvertently kept in an undesired state and also eliminates inadvertent inputs and incorrect permissions, which can lead to exposure of many different attack vectors.
[0007] Automatic graphics adjustment can also provide many technical benefits to a computing system. For example, by providing an adaptive adjustment of the graphical representation, each user of a communication session can benefit from group activities by obtaining a better context of the current situation. By providing this more detailed information and user stimuli, the system can promote user engagement to help the system reduce user fatigue. By reducing user fatigue, especially in a communication system, users can exchange information more effectively. This helps to mitigate the occurrence of lost or overlooked shared content. This can reduce the situations where users need to resend information. More effective communication of shared content can also help avoid the need for external systems (such as mobile phones for sending text messages and other messaging platforms). The systems and features described herein can also help reduce the reuse of network, processor, memory, or other computing resources.
[0008] By reading the following detailed description and reading the associated drawings, features and technical benefits other than those explicitly described above will be apparent. The present invention content is provided to introduce a selection of concepts further described below in the detailed description in a simplified form. The present invention content is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. For example, the term "technology" can refer to systems, methods, computer-readable instructions, modules, algorithms, hardware logic, and / or operations as permitted by the above context and the entire document. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The detailed description is described with reference to the accompanying drawings, in which the leftmost digit of the reference numeral identifies the drawing in which the reference numeral first appears. The same reference numerals in different drawings indicate similar or identical items. References to individual items among multiple items can use reference numerals with letters in an alphabetical sequence to refer to each individual item. A general reference to an item can use a specific reference numeral without an alphabetical sequence.
[0010] Figure 1 is a diagram of a user interface and system for providing access control to audio and video streams and control of the representation of a communication session.
[0011] Figure 2A Shows a user interface in the first stage of a process for providing access control to audio and video streams and control of the representation of a communication session.
[0012] Figure 2B Shows a user interface in the second stage of a process for providing access control to audio and video streams and control of the representation of a communication session, where the user interface shows a form of input command.
[0013] Figure 2CShows the user interface in the third stage of the process for providing access control to audio and video streams and control of the representation of a communication session, where the user interface shows the state of movement and representation changes in response to an input command and the state of change in access to the audio stream.
[0014] Figure 2D Shows the user interface in the fourth stage of the process for providing access control to audio and video streams and control of the representation of a communication session, where the user interface shows a second command for controlling the video stream and additional access to the video stream in response to the second command.
[0015] Figure 2E Shows the user interface in the fifth stage of the process for providing access control to audio and video streams and control of the representation of a communication session, where the user interface shows the result of the second command for controlling the video stream and additional access to the video stream.
[0016] Figure 3A Shows the user interface in the first stage of the process for providing access control to audio and video streams in response to movement of the representation of a communication session.
[0017] Figure 3B Shows the user interface in the second stage of the process for providing access control to audio and video streams in response to movement of the representation of a communication session, where a representation of a person is moving towards the representation of the discussion.
[0018] Figure 3C Shows the user interface in the second stage of the process for providing access control to audio and video streams in response to movement of the representation of a communication session, where a representation of a person is moving towards the representation of the discussion.
[0019] Figure 3D Shows the user interface in the second stage of the process for providing access control to audio and video streams in response to movement of the representation of a communication session, where a representation of a person is moving towards the representation of the discussion.
[0020] Figure 3E Shows the user interface in the second stage of the process for providing access control to audio and video streams in response to movement of the representation of a communication session, where a representation of a person is moving towards the representation of the discussion.
[0021] Figure 3F Shows the user interface in the second stage of the process for providing access control to audio and video streams in response to movement of the representation of a communication session, where a representation of a person is moving towards the representation of the discussion.
[0022] Figure 3G Shows a user interface in the second stage of a process for providing access control to audio and video streams in response to movement of a representation of a communication session, where a representation of a person is moving towards a representation of a discussion.
[0023] Figure 3H Shows a user interface in the second stage of a process for providing access control to audio and video streams in response to movement of a representation of a communication session, where a representation of a person is moving towards a representation of a discussion.
[0024] Figure 3I Shows a user interface in the second stage of a process for providing access control to audio and video streams in response to movement of a representation of a communication session, where a representation of a person is moving towards a representation of a discussion.
[0025] Figure 3J Shows a user interface in the second stage of a process for providing access control to audio and video streams in response to movement of a representation of a communication session, where a representation of a person is moving towards a representation of a discussion.
[0026] Figure 3K Shows a user interface in the second stage of a process for providing access control to audio and video streams in response to movement of a representation of a communication session, where a representation of a person is moving towards a representation of a discussion.
[0027] Figure 3L Shows a user interface in the second stage of a process for providing access control to audio and video streams in response to movement of a representation of a communication session, where a representation of a person is moving towards a representation of a discussion.
[0028] Figure 3M Shows a user interface in the second stage of a process for providing access control to audio and video streams in response to movement of a representation of a communication session, where a representation of a person is moving towards a representation of a discussion.
[0029] Figure 4A Shows aspects of a first output of a controlled spatial audio signal based on a first set of positions of representations on a user interface.
[0030] Figure 4B Shows aspects of a second output of other controlled spatial audio signals based on a second set of positions of representations on a user interface.
[0031] Figure 5 Shows aspects of an output of a controlled spatial audio signal based on a set of positions of representations in a 3D environment.
[0032] Figure 6Shows aspects of an embodiment in which subgroups can be identified and selected based on a choice of a topic of a related discussion or a person participating in the related discussion.
[0033] Figure 7 Is a flowchart showing aspects of routines of the disclosed technology.
[0034] Figure 8 Is a computer architecture diagram showing an illustrative computer hardware and software architecture of a computing system capable of implementing aspects of the technologies and techniques presented herein.
[0035] Figure 9 Is a computer architecture diagram showing a computing device architecture of a computing device capable of implementing aspects of the technologies and technologies presented herein. Detailed Description
[0036] Figure 1 Shows system 100 that controls access to audio and video streams and controls a representation indicating a modification of access to the audio and video streams. The communication session can be in the form of an online meeting, where the video stream and the audio stream are shared among users 10 of system 100. System 100 can include multiple computers 11, each corresponding to an individual user 10. For illustrative purposes, first user 10A is associated with first computer 11A, second user 10B is associated with second computer 11B, and other users are associated with other individual computers, up to user 10O being associated with computer 11O. These users can also be referred to as "user A", "user B", etc. respectively. This example is provided for illustrative purposes and should not be construed as limiting. It can be understood that the system can include any number of users and any number of devices.
[0037] Each user can be displayed as a two-dimensional 2D image in the user interface, or each user can be displayed as a three-dimensional representation, such as an avatar. The 3D representation can be a static model or a dynamic model that is animated in real time in response to user input. Although this example shows a user interface where the users are displayed as 2D images, some of which can include live video rendering, it can be understood that the technologies disclosed herein can be applied to other forms of representation, video, or other types of rendering. Computer 11 can be in the form of a desktop computer, a head-mounted display unit, a tablet computer, a mobile phone, etc. The system can generate a user interface showing aspects of the communication session to each user. In Figure 1 In the example of, the first user interface layout 101A can include multiple renderings 102 of one or more users 10.
[0038] Rendering 102 may include the rendering of two-dimensional (2D) images, which may include pictures or live video feeds of users 10. The user interface arrangement 101A includes a plurality of renderings 102, each rendering 102 being associated with an individual user 10 of a communication session. An individual cluster of users (e.g., cluster 1) represents an individual discussion group in which a subset of users is participating in a private session, where each person shares two-way video and audio signals. This means that permissions for each user in a subgroup (such as group 1 discussing AI and education topics) can hear and see each other's video streams. They can also share content such as files and other information. Permissions restrict others from receiving or sending video streams to this subgroup. For example, the first cluster 1 of renderings represents a discussion 103A among the subset of users 10L - 10O represented by respective renderings 102L - 102O located in association with cluster 1. This communication session also includes two other subgroups involving a second discussion 103B and a third discussion 103C. Individual users having a rendering 102 that is not part of a particular cluster have permissions that restrict them from receiving or sending audio or video streams or from sending audio or video streams to subgroup members.
[0039] The user interface arrangement 101A may also include a list 104 of individual discussions 103, each individual discussion 103 being associated with a graphical element depicting the individual cluster 1 - 3 representing the individual discussion 103. Each discussion 103 on the list 104 may include a first button 107 that allows a user (in this example, user A 10A) to receive an audio stream from the corresponding subgroup participating in the discussion 103. A second button 108 allows the user (e.g., user A 10A) to send and receive audio and video streams with the corresponding subgroup participating in the discussion 103. After selecting the second button, the system changes the user's permissions to allow them to send and receive audio and video streams with the corresponding subgroup, while also changing the position of the user's representation to move to the graphical element representing the subgroup.
[0040] Figures 2A - 2C An illustration of how a user such as the first user 10A can use a listen-only mode to change their association with a group (e.g., the first group 103A) is shown. For illustrative purposes, the user interface shown in Figure 2A is displayed to the first user 10A via an associated device 11A, and visual indicators are provided in the lower left of the drawing to show the volume of each discussion for the first user 10A for purposes of illustrating aspects of the present disclosure. In the Figure 2A operational state shown, where the rendering 102A of user A is not visually associated with the subgroup 103, the system restricts the device 11A of the first user 10A from communicating, receiving, or sending audio and video streams with other computers 11 of other users 10.
[0041] A user can receive an audio stream of a discussion subgroup in "listen-only mode" by selecting the "Listen" button for a specific subgroup participating in the discussion. In this example, as Figure 2B shown, the first user 10A provides an input to select the "Listen" button 107 for the first discussion (also referred to herein as Discussion 1, the first group 103A, or the representation of the first group 103A). As shown, the system 100 can receive an input indicating the selection of the first discussion 103A through interaction with the graphical element 107 associated with the first discussion 103A, where the graphical element 107 is positioned on the list 104 in association with the description of the discussion 103A.
[0042] As described herein, other types of inputs can be utilized to allow a user to listen to the audio stream of a specific discussion. For example, a user can select a button related to a topic, such as an AI button in the upper right corner of the user interface. If the AI button is selected and assuming the first discussion subgroup is discussing the AI topic, the system can highlight the representation of Discussion 103A and allow the user to listen to the audio stream of the user group 10L - 10O participating in that subgroup. A person can also listen to the group discussion by using a filter that allows searching by name or keywords of the conversation. For example, if the first user selects the "People" button and provides the name or identifier of a user in the subgroup, the system will allow the first user to listen to or join the subgroup. In addition to the listen-only access permission to the group's stream, these types of selections also cause the rendering of the user to be moved to the graphical representation of the group or within the graphical representation of the group.
[0043] In response to an input (e.g., a command) indicating the selection of the first discussion 103A from the list 104, as Figure 2C shown, the rendering 102A of the first user 10A associated with the input is moved to the discussion cluster of the first discussion 103A. As shown, the system moves the rendering 102A of the user 10A associated with the input to a position indicating the association between the rendering 102A and the rendering cluster 1 representing the selected discussion 103A. Thus, the first user 10A can now hear the discussion between the subset of users 10L - 10O in "listen-only" mode. Thus, in response to this input, the system modifies the access rights of the computing device 11A associated with the first user 10A to receive audio signals from the computing devices 11L - 11O of the subset of users 10L - 10O having renderings 102L - 102O positioned in association with the cluster 1.
[0044] Figures 2D - 2E Shows how a person can join a group using a Full Join that provides a two-way audio and / or two-way video exchange between the members of the selected subgroup and the requesting user. As Figure 2DAs shown, a user such as the first user 10A may provide input to select a discussion. For example, the system may receive a selection of a second graphical element 108 (e.g., a join button) associated with the first discussion 103A. The second graphical element 108 (join button) is positioned in association with a description of the first discussion 103A on the list 104. The description includes the names and / or topics of the people in the subgroup of the first discussion.
[0045] In response to selecting the second graphical element 108 of discussion 103A from list 104 or other form of input indicating joining, the system may cause a list of discussion elements 103A to be added to the discussion. Figure 2A The user interface arrangement 101A of the plurality of renderings 102 is Figure 2E 103A from list 104 or other form of input indicating joining, the system may modify access permissions of computing device 11A associated with first user 10A to allow computing device 11A to send audio signals to computing devices 11L-110 of subset of users 10L-100 having renderings 102L-102O positioned in association with cluster 1. In some embodiments, in response to selecting the second graphical element 108 of the discussion 103A from the list 104 or other form of input indicating joining, the system may modify the access permissions of the computing device 11A to allow two-way video and audio communications with the computing devices 11L-11O of the subset of users 10L-10O having renderings 102L-102O positioned in association with cluster 1 of the first discussion 103A.
[0046] Figures 3A - 3M Is the process for providing access control to audio and video streams in response to movement of representations of a communication session. Figure 3A A user interface in a first stage of a process for providing access control to audio and video streams in response to movement of a representation of a communication session is shown. Figure 1 The rendering 102A) shown as user 10A is not associated with any particular subgroup of discussion 103. Thus, the first user's computer is restricted from exchanging audio and video signals with the other users' devices.
[0047] Figures 3B - 3DThe rendering of the first user is shown to move to a position where the rendering of the first user has a graphical association with the third group. The movement of each user can be based on a voice command, a pointer input command, or any other suitable input from each user. In response to this graphical relationship, the permission of the first computer is modified to allow the first computer of the first user to receive an audio stream from the computer of a user whose representation is depicted in association with the third subgroup 3, or a user within the boundaries of the third subgroup 3. At the same time, User B and User C have moved within a predetermined distance of each other. When two people move within a predetermined distance of each other, the system changes their permissions so that they can participate in a conversation with two-way audio exchange (e.g., a voice call). This activity forms a fourth discussion with the new subgroup. They can upgrade to a video call with the approval of both parties.
[0048] Figure 3E A scenario is shown where User B and User C move their representations away from each other. When the distance between these representations exceeds a threshold, the system restricts their devices from being able to exchange audio or video streams. Any data defining the group object for the fourth discussion 4 is also removed or deactivated. Also shown in Figure 3D User A has moved but is still only listening to the third discussion 3 because their representation still shows a graphical association with the group, e.g., touching the boundary, within the boundary, having a threshold distance from a point within the graphical item (e.g., a circle) representing the group, etc.
[0049] Figure 3F and 3G shows that User B and User D (shown in Figure 1 ) have moved within a threshold distance of each other to create a new discussion group, the fifth discussion 5. This allows User B and User D to exchange audio and video signals. Once the group is formed, the users can move to the boundary of the group to only listen. But full joining can only occur after receiving a second command. Also shown, User C has also moved to a position showing a graphical association with the second group. This allows User C to listen to the audio stream of the people associated with this second discussion group.
[0050] As Figure 3H shown, User A has moved their representation to a position where their representation has a graphical association with the fifth discussion group. This allows User A to listen to the discussion between User B and User D. Then, as Figures 3I to 3J shown, User A moves their representation to a position where their representation has a graphical association with the first discussion group. This allows User A to listen to the discussion between the users (Users 10L - 10O) of that group.
[0051] Figures 3J to 3MThe transformed display of the user interface shown in shows how user A can transform to full participation in the first group. This transformation can be in response to a command or input indicating that user A desires to fully participate in the discussion for two-way audio and video communication. As shown throughout the transition, the user interface can focus on the group by magnifying a particular discussion group (e.g., the first discussion group), showing a representation of the group (rendering the surrounding oval), and showing the live video stream of each user on the right side of the user interface. Figure 3M The second user interface layout 101B of can have a self-view on the lower right part of the UI, and the left part can show a representation of the group and the live video stream shown on the right side of the UI. Other smaller representations of other groups can also be shown, e.g., together with group identifiers or topic headings.
[0052] In some configurations, when a user is not participating in a subgroup discussion (e.g., a breakout session or a private communication session with a user subgroup), the user can receive audio signals from multiple subgroups so that they can hear different conversations to help them choose which group to join. In some configurations, the audio streams from different conversations can be based on spatial audio technology so that the sounds seem to come from a specific direction.
[0053] Figure 4A Aspects of the first output of a controlled spatial audio signal based on the position of the first group represented on the user interface are shown. In this example, the representation of user A is located between two clusters, each cluster representing a different subgroup conversation. Assuming that the first discussion group 1 is to the left of the representation of user A, the system can play the audio stream from this first discussion group 1 in the left speaker of the speakers of user A's computing device. Similarly, assuming that the second discussion group 2 is to the right of the representation of user A, the system can play the audio stream from this second discussion group 2 in the right speaker of the speakers of user A's computing device.
[0054] Figure 4B Aspects of the second output of other controlled spatial audio signals based on the position of the second group represented on the user interface are shown. In the case where the representation of user A is now moved to a new position, as shown by path 151, the system modifies the source of the stream. In this new position, the system can lower the volume of the audio stream from the second discussion 2, and the system can play the audio stream from the first discussion 1 in the rear speaker of the speakers of user A's computing device.
[0055] In such a configuration, the system can cause an audio signal to be transmitted from the computing devices 11P-11T of user 10P-10T participating in other discussion subgroups to the computing device 11A associated with user 10A. The audio signal from the computing devices 11P-1IT is transmitted to the computing device 11A associated with user IDA, and the position of the rendering 102A of user 10A does not have a visual association with Cluster 1. For example, the user is not part of the subgroup. The first component of the audio signal (e.g., the left channel of a stereo signal) can include the audio signal from the first user cluster 10P-10Q participating in the first discussion 103N. The second component of the audio signal (e.g., the right channel of a stereo signal) can include the audio signal from the second user cluster 10R-10T participating in the second discussion 103M. The volume of the first component is based on the distance between the rendering 102A of user 10A and the representation of the first cluster, and the volume of the second component is based on the distance between the rendering 102A of user 10A and the representation of the second cluster.
[0056] Figure 5 Aspects of an output of a controlled spatial audio signal based on a set of positions represented in a 3D environment are shown. In such an embodiment, the position of the source of the first component is based on the position of the rendering 102A of user 10A relative to the position of the representation of the first cluster. The position of the source of the second component is based on the position of the rendering 102A of user 10A relative to the position of the representation of the second cluster.
[0057] In Figure 5 embodiments, the spatial audio signal can be generated by HRTF or Dolby Atmos technology. This allows the speakers to generate sounds that appear to come from a specific position relative to the user in the real world. Thus, if the computing device is to generate sounds that appear to come from the left side of the user for the user, the system can generate those signals. In this case, if the position 111 of the avatar is at the center of the sphere and the avatar is facing a specific direction, the system can generate a signal component that gives the user the experience that the sound comes from the left side when the representation of the discussion subgroup 103A is on the left side of the avatar. Similarly, if the position 111 of the avatar is at the center of the sphere and the avatar is facing a specific direction, the system can generate a second signal component that gives the user the experience that the sound comes from the right side when the representation of the discussion subgroup 103B is on the right side of the avatar.
[0058] Figure 6 Aspects of an embodiment are shown in which subgroups can be identified and selected based on a relevant discussion or the selection of a person participating in the relevant discussion. As Figure 6As shown, other types of inputs can be utilized to allow a user to listen to an audio stream of a specific discussion session. For example, the user can select the AI button in the upper right of the user interface. Assuming that the first discussion subgroup is discussing the AI topic, the system can highlight the representation of discussion 103A and allow the user to listen to the audio stream of the group of users 10L - 10O in the participating subgroup. A person can also listen to the group discussion by using a filter that allows searching by person's name or keywords of the conversation. For example, if the first user selects the "person" button and provides the item or identifier of the user in the subgroup, the system will allow the first user to listen to or join the subgroup.
[0059] In an embodiment related to Figure 6 the input indicating the selection of discussion 103A can include an operation for receiving an input indicating a topic. Then, the system can determine that the topic is relevant to the discussion. If the topic is relevant to the discussion, in response to determining that the topic is relevant to the discussion, the system can move the rendering 102A representing user 10A to the group and modify the access permission of the computing device 11A associated with user 10A to receive audio signals from the people in the subgroup.
[0060] The input indicating the selection of discussion 103A can include receiving an input indicating an identifier of a discussion participant. Then, the system can determine that the discussion participant identified in the input is participating in the discussion of the subgroup of users 10L - 10O. Then, in response to determining that the discussion participant identified in the input is participating in the discussion of the user subgroup, the system can move the rendering 102A representing user 10A to the subgroup and modify the access permission of the computing device 11A associated with user 10A to receive audio signals from the subgroup.
[0061] Figure 7 FIG. [FIGURE NUMBER] is a diagram showing aspects of routine 900 for computationally efficient management of access permissions and user interface layout. Those of ordinary skill in the art should understand that the operations of the methods disclosed herein are not necessarily presented in any particular order, and it is possible and expected that some or all of the operations may be performed in an alternative order. For ease of description and illustration, the operations have been presented in the order shown. Operations can be added, omitted, performed together, and / or performed simultaneously without departing from the scope of the appended claims.
[0062] It should also be understood that the methods shown may end at any time and need not be executed in their entirety. Some or all of the operations in the method and / or substantially equivalent operations may be performed by executing computer-readable instructions included on a computer storage medium, as defined herein. As used in the specification and claims, the term "computer-readable instructions" and variations thereof are used herein broadly to include routines, applications, application modules, program modules, programs, components, data structures, algorithms, etc. Computer-readable instructions may be implemented on various system configurations, including single-processor or multiprocessor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based programmable consumer electronics, combinations thereof, and the like. Although the exemplary routines described below operate on a system (e.g., one or more computing devices), it should be understood that the routines may be executed on any computing system that may include any number of computers working together to perform the operations disclosed herein.
[0063] Accordingly, it should be understood that the logical operations described herein are implemented as (1) a sequence of computer-implemented acts or program modules running on a computing system such as described herein, and / or (2) interconnected machine logic circuits or circuit modules within the computing system. The implementation is a matter of choice depending on the performance and other requirements of the computing system. Thus, the logical operations may be implemented in software, firmware, special purpose digital logic, and any combination thereof.
[0064] In addition, Figure 7 the operations shown in other figures may be implemented in association with the exemplary presentation user interface (UI) described above. For example, the various devices and / or modules described herein may generate, send, receive, and / or display data associated with the content of a communication session (e.g., live content, broadcast events, recorded content, etc.) and / or a presentation UI that includes a presentation of one or more participants of a remote computing device, avatar, channel, chat session, video stream, image, virtual object, and / or application associated with the communication session.
[0065] Routine 900 begins at operation 902, where the system displays a first UI arrangement showing the original computer state of the user's location and audio access permission. As Figure 1 shown, when a subgroup of users is in a private audio and video discussion, the user interface may show user renderings 102 arranged in a user cluster.
[0066] At operation 904, the system may receive a command to change the computer state. This may include a voice command or a device input, such as a pointer device, that indicates a subgroup of users in the discussion, the topic of the discussion, the people in the discussion, or any other information that identifies the discussion and / or the discussion participants.
[0067] At operation 906, the system moves the rendering of the user associated with the input to a position relative to the representation of the subgroup associated with the identified discussion. An example of this movement is shown in Figure 2C .
[0068] At operation 908, the system modifies the access rights of the computing device associated with the user to receive audio signals. In the example of Figure 2C , the new access rights allow the computer of User A to receive audio signals from the computers of users in the selected discussion group.
[0069] At operation 910, the system may receive a second command. This can be any form of input indicating that the user (e.g., User A) wants to fully join the discussion. An example of this input is shown in Figure 2D .
[0070] At operation 912, in response to the second command, the system changes the user interface format to display the video of the subgroup to User A.
[0071] At operation 914, the system modifies the access rights, where the new access rights allow the computer of User A to send audio and video signals to and receive audio and video signals from the computers of users in the selected discussion group.
[0072] Figure 8 is a diagram showing an example environment 1100 in which the system 1102 (which can be the system 100 of Figure 1 ) can implement the techniques disclosed herein. In some embodiments, the system 1102 can be used to collect, analyze, and share data defining one or more objects presented to the users of the communication session 1104.
[0073] As shown, the communication session 1104 can be implemented among a plurality of client computing devices 1106(1) to 1106(N) (where N is a number having a value of 2 or greater) associated with or part of the system 1102. The client computing devices 1106(1) to 1106(N) enable users (also referred to as individuals) to participate in the communication session 1104.
[0074] In this example, communication session 1104 is hosted by system 1102 over one or more networks 1108. That is, system 1102 can provide services that enable users of client computing devices 1106(1) through 1106(N) to participate in communication session 1104 (e.g., via live viewing and / or recorded viewing). Thus, the "participants" in communication session 1104 can include users and / or client computing devices (e.g., multiple users can be in a room participating in a communication session via use of a single client computing device), and each user can communicate with other participants. Alternatively, communication session 1104 can be hosted by one of client computing devices 1106(1) through 1106(N) using peer-to-peer technology. System 1102 can also host chat conversations and other team collaboration functions (e.g., as part of an application suite).
[0075] In some embodiments, such chat conversations and other team collaboration functions are considered to be external communication sessions distinct from communication session 1104. A computerized agent configured to collect participant data in communication session 1104 can be able to link to such external communication sessions. Thus, the computerized agent can receive information such as date, time, session details, etc., that enables connection to such external communication sessions. In one example, a chat conversation can be conducted in accordance with communication session 1104. Additionally, system 1102 can host communication session 1104 that includes at least multiple participants co-located at a meeting location (such as a conference room or auditorium) or located at different locations.
[0076] In the examples described herein, client computing devices 1106(1) through 1106(N) participating in communication session 1104 are configured to receive and render communication data for display on a user interface of a display screen. The communication data can include a collection of various instances or streams of live content and / or recorded content. The various instances or streams of live content and / or recorded content can be provided by one or more cameras (such as video cameras). For example, an individual live or recorded content stream can include media data associated with a video feed provided by a camera (e.g., audio and visual data capturing the appearance and speech of users participating in the communication session). In some embodiments, the video feed can include such audio and visual data, one or more still images, and / or one or more avatars. The one or more still images can also include one or more avatars.
[0077] Another example of a separate live and / or recorded content stream can include media data that includes an avatar of a user participating in a communication session and audio data that captures the user's voice. Another example of a separate live or recorded content stream can include media data that includes a file displayed on a display screen and audio data that captures the user's voice. Thus, the various live and / or recorded content streams within the communication data enable remote conferencing among a group of people and sharing of content within the group. In some embodiments, the various live and / or recorded content streams within the communication data can originate from multiple co-located cameras located in a space such as a room to record or live stream a presentation that includes one or more individuals presenting and one or more individuals consuming the presented content.
[0078] Participants or attendees can view the content of the communication session 1104 in real time as the event occurs, or alternatively, via a recording at a later time after the event (examples described herein), the client computing devices 1106(1) to 1106(N) participating in the communication session 1104 are configured to receive and render the communication data for display on a user interface of a display screen. The communication data can include a collection of various instances or streams of live and / or recorded content. For example, a single content stream can include media data associated with a video feed (e.g., audio and visual data that captures the appearance and voice of a user participating in a communication session). Another example of a single content stream can include media data that includes an avatar of a user participating in a meeting session and audio data that captures the user's voice. Another example of a single content stream can include media data and / or audio data, where the media data includes a content item displayed on a display screen and the audio data captures the user's voice. Thus, the various content streams within the communication data enable conferencing or broadcast presentation among a group of people dispersed at remote locations.
[0079] A participant or attendee of a communication session is a person within the range of a camera or other image and / or audio capture device such that the person's actions and / or sounds generated while the person is viewing and / or listening to the content shared via the communication session can be captured (e.g., recorded). For example, a participant may be sitting in a crowd watching a live broadcast of shared content at a broadcast location where a stage presentation is occurring. Or a participant can be sitting in an office conference room watching the shared content of a communication session with other colleagues via a display screen. Further still, a participant can be sitting or standing in front of a personal device (e.g., a tablet computer, a smart phone, a computer, etc.) watching the shared content of a communication session alone in their office or at home.
[0080] System 1102 includes device 1110. Device 1110 and / or other components of system 1102 may include distributed computing resources that communicate with each other and / or with client computing devices 1106(1) to 1106(N) via one or more networks 1108. In some examples, system 1102 may be an independent system responsible for managing aspects of one or more communication sessions such as communication session 1104. As an example, system 1102 may be managed by entities such as SLACK, WEBEX, GOTOMEETING, GOOGLE HANGOUTS, etc.
[0081] Network 1108 may include, for example, a public network such as the Internet, a private network such as an institutional and / or personal intranet, or some combination of private and public networks. Network 1108 may also include any type of wired and / or wireless network, including but not limited to local area networks (“LANs”), wide area networks (“WANs”), satellite networks, wired networks, Wi-Fi networks, WiMax networks, mobile communication networks (e.g., 3G, 4G, etc.), or any combination thereof. Network 1108 may utilize communication protocols, including packet-based and / or datagram-based protocols such as Internet Protocol (“IP”), Transmission Control Protocol (“TCP”), User Datagram Protocol (“UDP”), or other types of protocols. In addition, network 1108 may also include multiple devices that facilitate network communication and / or form the hardware foundation of the network, such as switches, routers, gateways, access points, firewalls, base stations, repeaters, backbone devices, etc.
[0082] In some examples, network 1108 may also include devices that enable connection to wireless networks, such as wireless access points (“WAPs”). Example WAPs support connections that send and receive data at various electromagnetic frequencies (e.g., radio frequency), including WAPs that support Institute of Electrical and Electronics Engineers (“IEEE”) 802.11 standards (e.g., 802.11g, 802.11n, 802.11ac, etc.) and other standards.
[0083] In various examples, device 1110 may include one or more computing devices operating in a cluster or other grouped configuration to share resources, balance loads, improve performance, provide failover support or redundancy, or for other purposes. For example, device 1110 may belong to various categories of devices, such as traditional server-type devices, desktop computer-type devices, and / or mobile-type devices. Thus, although shown as a single type of device or server-type device, device 1110 may include various device types and is not limited to a particular type of device. Device 1110 may represent, but is not limited to, a server computer, a desktop computer, a web server computer, a personal computer, a mobile computer, a laptop computer, a tablet computer, or any other kind of computing device.
[0084] A client computing device (e.g., one of client computing devices 1106(1) through 1106(N)) may belong to various device categories, which may be the same as or different from those of device 1110, such as traditional client-type devices, desktop computer-type devices, mobile-type devices, specialized-type devices, embedded-type devices, and / or wearable-type devices. Thus, client computing devices may include, but are not limited to, desktop computers, game consoles and / or gaming devices, tablet computers, personal data assistants (“PDAs”), mobile phone / tablet hybrid devices, laptop computers, telecommunications devices, computer navigation-type client computing devices, such as satellite-based navigation systems, including global positioning system (“GPS”) devices, wearable devices, virtual reality (“VR”) devices, augmented reality (“AR”) devices, implantable computing devices, automotive computers, network-enabled televisions, thin clients, terminals, Internet of Things (“IoT”) devices, workstations, media players, personal video recorders (“PVRs”), set-top boxes, cameras, integrated components for inclusion in a computing device (e.g., peripherals), appliances, or any other kind of computing device. Additionally, client computing devices may include combinations of the previously listed examples of client computing devices, such as, for example, a desktop computer-type device or a mobile-type device combined with a wearable device, etc.
[0085] Client computing devices 1106(1) through 1106(N) of various categories and device types may represent any type of computing device having one or more data processing units 1192 operably connected to a computer-readable medium 1194, such as via a bus 1116. In some instances, bus 1116 may include a system bus, a data bus, an address bus, a PCI bus, a Mini-PCI bus, and one or more of any various local, peripheral, and / or independent buses.
[0086] The executable instructions stored on the computer-readable medium 1194 may include, for example, an operating system 1119, a client module 1120, a profile module 1122, and other modules, programs, or applications that can be loaded and executed by the data processing unit 1192.
[0087] The client computing devices 1106(1) to 1106(N) may also include one or more interfaces 1124 to enable communication between the client computing devices 1106(1) to 1106(N) and other networked devices (such as device 1110) via the network 1108. Such network interfaces 1124 may include one or more network interface controllers (NICs) or other types of transceiver devices to send and receive communications and / or data over the network. Additionally, the client computing devices 1106(1) to 1106(N) may include input / output (“I / O”) interfaces (devices) 1126 that enable communication with input / output devices, including user input devices such as peripheral input devices (e.g., game controllers, keyboards, mice, pens, voice input devices such as microphones, cameras for obtaining and providing video feeds and / or still images, touch input devices, gesture input devices, etc.) and / or output devices such as peripheral output devices (e.g., displays, printers, audio speakers, haptic output devices, etc.). Figure 8 It is shown that the client computing device 1106(1) is connected to a display device (e.g., display screen 1129(1)) in a certain manner, which can display a UI according to the techniques described herein.
[0088] In Figure 8 the example environment 1100, the client computing devices 1106(1) to 1106(N) can use their respective client modules 1120 to connect to each other and / or to other external devices in order to participate in a communication session 1104 or to contribute activities to the collaborative environment. For example, a first user can use the client computing device 1106(1) to communicate with a second user of another client computing device 1106(2), and when the client module 1120 is executed, the users can share data, which can cause the client computing device 1106(1) to connect to the system 1102 and / or other client computing devices 1106(2) to 1106(N) via the network 1108.
[0089] The client computing devices 1106(1) to 1106(N) can use their respective profile modules 1122 to generate participant profiles ( Figure 8(not shown in the figure), and provide the participant profile to other client computing devices and / or the device 1110 of the system 1102. The participant profile may include one or more of the identity of the user or user group (e.g., name, unique identifier (“ID”), etc.), user data such as personal data, machine data such as location (e.g., IP address, room in a building, etc.), and technical capabilities. The participant profile can be used to register the participants of the communication session.
[0090] As Figure 8 shown, the device 1110 of the system 1102 includes a server module 1130 and an output module 1132. In this example, the server module 1130 is configured to receive media streams 1134(1) to 1134(N) from various client computing devices such as client computing devices 1106(1) to 1106(N). As described above, the media stream may include a video feed (e.g., audio and visual data associated with the user), audio data to be output together with the rendering of the user's avatar (e.g., an audio-only experience without sending the user's video data), text data (e.g., text messages), file data, and / or screen sharing data (e.g., documents, slides, images, videos displayed on a display screen, etc.). Thus, the server module 1130 is configured to receive a collection of various media streams 1134(1) to 1134(N) (this collection is referred to herein as “media data 1134”) during the live viewing of the communication session 1104. In some scenarios, not all client computing devices participating in the communication session 1104 provide media streams. For example, a client computing device can be a consumption or “listening” device only, such that it only receives the content associated with the communication session 1104 but does not provide any content to the communication session 1104. The communication session 1104 can have a start time and an end time, or the communication session 1104 can be ongoing. The communication session 1104 can also be classified as an event and have phases, where each phase causes the computer to change the role of an individual user as the event transitions through each phase.
[0091] In various examples, the server module 1130 may select aspects of the media stream 1134 to share with each of the participating client computing devices 1106(1) through 1106(N). Thus, the server module 1130 may be configured to generate session data 1136 based on the stream 1134 and / or pass the session data 1136 to the output module 1132. The output module 1132 may then transmit communication data 1139 to the client computing devices (e.g., the client computing devices 1106(1) through 1106(N) participating in the live viewing of the communication session). The communication data 1139 may include video, audio, and / or other content data provided by the output module 1132 based on the content 1150 associated with the output module 1132 and based on the received session data 1136. The devices 1110 of the system 1102 may also access the queue data 101 described above in connection with Figure 1 the standard data 1191 for defining the standards and / or thresholds described herein. The standard data 1191 may also include machine learning data accessible by a machine learning service or machine learning module, which may be part of the server module 1130 or part of a remote machine learning service, such as those accessible by a public API at a site run by IBM, Google, or Microsoft.
[0092] As shown, the output module 1132 sends the communication data 1139(1) to the client computing device 1106(1), and sends the communication data 1139(2) to the client computing device 1106(2), and sends the communication data 1139(3) to the client computing device 1106(3), and so on. The communication data 1139 transmitted to the client computing devices may be the same or may be different (e.g., the positioning of the content stream within the user interface may vary with the device).
[0093] In various embodiments, the device 1110 and / or the client module 1120 of the system 1102 may include a GUI rendering module 1140. The GUI rendering module 1140 may be configured to analyze communication data 1139 for delivery to one or more of the client computing devices 1106. Specifically, the UI rendering module 1140 at the device 1110 and / or the client computing device 1106 may analyze the communication data 1139 to determine an appropriate manner for displaying video, images, and / or content on the display screen 1129 of the associated client computing device 1106. In some embodiments, the GUI rendering module 1140 may provide video, images, and / or content to a rendered GUI 1146 presented on the display screen 1129 of the associated client computing device 1106. The rendered GUI 1146 may be caused to be presented on the display screen 1129 by the GUI rendering module 1140. The rendered GUI 1146 may include video, images, and / or content analyzed by the GUI rendering module 1140.
[0094] In some embodiments, the rendered GUI 1146 may include multiple sections or grids that may render or include video, images, and / or content for display on the display screen 1129. For example, a first section of the rendered GUI 1146 may include a video feed of a presenter or individual, and a second section of the rendered GUI 1146 may include a video feed of an individual consuming meeting information provided by the presenter or individual. The GUI rendering module 1140 may populate the first and second sections of the rendered GUI 1146 in a manner that appropriately mimics the environmental experience that the presenter and individual may be sharing.
[0095] In some embodiments, the GUI rendering module 1140 may zoom in or provide a zoomed view of an individual represented by a video feed to highlight the individual's reaction to the presenter, such as facial features. In some embodiments, the rendered GUI 1146 may include video feeds of multiple participants associated with a meeting, such as a general communication session. In other embodiments, the rendered GUI 1146 may be associated with a channel, such as a chat channel, an enterprise team channel, etc. Thus, the rendered GUI 1146 may be associated with an external communication session different from a general communication session.
[0096] Figure 9FIG. shows an example block diagram of an example device 1200 (also referred to herein as a "computing device") configured to generate and process some of the user interfaces disclosed herein. Device 1200 may generate data that may include one or more portions that may render or include video, images, and / or content for display on a display screen 1129. Device 1200 may represent one of the devices described herein. Additionally or alternatively, device 1200 may represent one of the client computing devices 1106.
[0097] As shown, device 1200 includes one or more data processing units 1202, a computer-readable medium 1204 (also referred to herein as computer storage medium 1204), and a communication interface 1206. The components of device 1200 are operably connected, for example, via a bus 1209, which may include a system bus, a data bus, an address bus, a PCI bus, a Mini-PCI bus, and one or more of any of a variety of local, peripheral, and / or independent buses.
[0098] As used herein, a data processing unit (such as data processing unit 1202 and / or data processing unit l192) may represent, for example, a CPU-type data processing unit, a GPU-type data processing unit, a field programmable gate array ("FPGA"), another class of digital signal processor ("DSP"), or other hardware logic components that may be driven by a CPU in some cases. Illustrative types of hardware logic components that may be utilized, for example but not limited to, include application specific integrated circuits ("ASICs"), application specific standard products ("ASSPs"), systems on a chip ("SOCs"), complex programmable logic devices ("CPLDs"), and the like.
[0099] As used herein, computer-readable media such as computer-readable medium 1204 and computer-readable medium 1194 may store instructions executable by a data processing unit. The computer-readable medium may also store instructions executable by an external data processing unit such as an external CPU, an external GPU, and / or executable by an external accelerator such as an FPGA-type accelerator, a DSP-type accelerator, or any other internal or external accelerator. In various examples, at least one CPU, GPU, and / or accelerator is incorporated into the computing device, while in some examples, one or more of the CPU, GPU, and / or accelerator are external to the computing device.
[0100] A computer-readable medium (which may also be referred to herein as a computer-readable medium) can include computer storage media and / or communication media. "Computer storage media", "non-transitory computer storage media", or "non-transitory computer-readable media" can include volatile memory, non-volatile memory, and / or any other permanent and / or auxiliary computer storage media, one or more of the removable and non-removable computer storage media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Thus, computer storage media or variants of the above terms include media in tangible and / or physical form that are included in a device or hardware component as part of a device or external to the device, including but not limited to random access memory ("RAM"), static random access memory ("SRAM"), dynamic random access memory ("DRAM"), phase change memory ("PCM"), read-only memory ("ROM"), erasable programmable read-only memory ("EPROM"), electrically erasable programmable read-only memory ("EEPROM"), flash memory, compact disc read-only memory ("CD-ROM"), digital versatile disc ("DVD"), optical card, or other optical storage media, cassette tapes, magnetic tapes, magnetic disk storage, magnetic cards, or other magnetic storage devices or media, solid-state memory devices, storage arrays, network-attached storage, storage area networks, hosted computer storage, or any other storage memory, storage device, and / or any storage media that can be used for local storage and maintenance of information for access at a computing device.
[0101] In contrast to computer storage media, communication media can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, or other transmission mechanism. As defined herein, computer storage media does not include communication media. That is, computer storage media does not include communication media that consists solely of a modulated data signal, a carrier wave, or a propagated signal itself.
[0102] Communication interface 1206 can represent, for example, a network interface controller ("NIC") or other types of transceiver devices that send and receive communications over a network. Additionally, communication interface 1206 can include one or more cameras and / or audio devices 1222 to enable generation of video feeds and / or still images, etc.
[0103] In the example shown, the computer-readable medium 1204 includes a data store 1208. In some examples, the data store 1208 includes a data repository, such as a database, a data warehouse, or other types of structured or unstructured data repositories. In some examples, the data store 1208 includes a corpus and / or a relational database having one or more tables, indexes, stored procedures, etc. to enable access to data including, for example, one or more of Hypertext Markup Language (“HTML”) tables, Resource Description Framework (“RDF”) tables, Web Ontology Language (“OWL”) tables, and / or Extensible Markup Language (“XML”) tables.
[0104] The data store 1208 can store data for the operation of processes, applications, components, and / or modules to be stored in the computer-readable medium 1204 and / or executed by the data processing unit 1202 and / or the accelerator. For example, in some examples, the data store 1208 can store meeting objects 1210, permission data 1212, and / or other data. The meeting object 1210 can include the total number of participants (e.g., users and / or client computing devices) in a communication session, the activities that occur in the communication session, a list of invitees to the communication session, and / or other data related to when and how the communication session is conducted or hosted. The object can also define subgroups, members of the subgroups, and other user information. The permission data 1212 stores all access rights for each user, e.g., whether a user's computer can receive an audio or video stream from a particular computer or send a video stream to a particular computer. The data store 1208 can also include context data 1214, which can include any information that defines the activities, criteria, or thresholds of the users disclosed herein.
[0105] Alternatively, some or all of the above data can be stored on a separate memory 1216 on one or more data processing units 1202, such as a memory on a CPU-type processor, a GPU-type processor, an FPGA-type accelerator, a DSP-type accelerator, and / or another accelerator. In this example, the computer-readable medium 1204 also includes an operating system 1218 and an application programming interface 1211 (API) configured to expose the functions and data of the device 1200 to other devices. Additionally, the computer-readable medium 1204 includes one or more modules, such as a server module 1230, an output module 1232, and a GUI rendering module 1240, although the number of modules shown is merely an example and the number can vary higher or lower. That is, the functions associated with the modules shown herein can be performed by a smaller number of modules or a larger number of modules on one device, or distributed across multiple devices.
[0106] It should be understood that, unless otherwise expressly stated, conditional language used herein such as "can", "could", or "may" is understood in context to present certain examples including certain features, elements, and / or steps, while other examples do not include certain features, elements, and / or steps. Thus, such conditional language is generally not intended to imply that certain features, elements, and / or steps are required in any way for one or more examples, or that one or more examples must include logic for deciding whether certain features, elements, and / or steps are included or will be performed in any particular example, with or without user input or prompting. Unless otherwise expressly stated, conjunctive language such as the phrase "at least one of X, Y, or Z" should be understood to present items, terms, etc. that can be X, Y, or Z or combinations thereof. Additionally, the words "the" or "if" may be used interchangeably. Thus, a phrase such as "determine that a criterion is met" may also be interpreted as "determine whether a criterion is met", and vice versa.
[0107] It should also be understood that many variations and modifications can be made to the above examples, and their elements should be understood to be other acceptable examples. All such modifications and variations are intended to be included within the scope of the present disclosure and protected by the appended claims.
[0108] Finally, although various configurations have been described in language specific to structural features and / or method acts, it should be understood that the subject matter defined in the appended representations need not be limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claimed subject matter.
Claims
1. A method for controlling access to audio and video streams of a communication session, performed by a computing system, the method comprising: Causing a display to include a plurality of rendered user interface arrangements, each rendering associated with an individual user of the communication session, the user interface arrangement further including clusters, each cluster representing an individual discussion of a subgroup of users, wherein a rendered cluster represents a discussion among a subset of users represented by individual renders located in association with the cluster, the user interface arrangement further including a list of individual discussions, each individual discussion associated with a graphical element depicting the individual cluster 1-3 representing the individual discussion; Receiving an input indicating selection of the discussion by interaction with the graphical element associated with the discussion, wherein the graphical element is located in association with the description of the discussion on the list; In response to an input selecting the discussion from the list: Moving the render representing the user associated with the input to a position indicating the association between the render representing the user and the rendered cluster representing the discussion among the subset of users, and Modifying the access rights of the computing device associated with the user to receive audio signals from the computing devices of the subset of users having renders located in association with the cluster.
2. The method according to claim 1, further comprising: Receiving a selection of a second graphical element associated with the discussion, wherein the second graphical element is located in association with the description of the discussion on the list; and In response to the selection of the second graphical element of the discussion from the list, modifying the access rights of the computing device associated with the user to send audio signals to the computing devices of the subset of users having renders located in association with the cluster.
3. The method according to claim 1, further comprising: Receiving a selection of a second graphical element associated with the discussion, wherein the second graphical element is located in association with the description of the discussion on the list; and In response to the selection of the second graphical element of the discussion from the list: Causing a transition from the user interface arrangement including the plurality of renders to a second user interface arrangement including individual renders of the subset of users, having graphical elements depicting the cluster and graphical elements representing other clusters, and Modifying the access rights of the computing device associated with the user to send audio signals to the computing devices of the subset of users having renders located in association with the cluster.
4. The method according to claim 1, wherein The input indicating selection of the discussion includes receiving an input indicating a topic, wherein the method further comprises: Determining that the topic is relevant to the discussion, wherein in response to determining that the topic is relevant to the discussion, moving the render representing the user and modifying the access rights of the computing device associated with the user to receive audio signals from the devices of the subgroup of users.
5. The method according to claim 1, wherein, The input indicating selection of the discussion includes receiving an input indicating an identifier of a discussion participant, wherein the method further comprises: Determine that the discussed participant identified in the input is participating in a discussion with a subgroup of users, wherein, in response to determining that the discussed participant identified in the input is participating in the discussion with the subgroup of users, move the rendering representing the user and modify the access rights of the computing device associated with the user to receive an audio signal from the devices of the subgroup of users.
6. The method according to claim 1, further comprising transmitting an audio signal from a computing device of a user participating in another discussion subgroup to the computing device associated with the user, wherein, When there is no visual association with the cluster at the location of the rendering of the user, simultaneously transmit the audio signal from the computing device to the computing device associated with the user, wherein a first component of the audio signal is from a first cluster of users participating in a first discussion, and a second component of the audio signal is from a second cluster of users participating in a second discussion, wherein the volume of the first component is based on the distance between the rendering of the user and the representation of the first cluster, and wherein the volume of the second component is based on the distance between the rendering of the user and the representation of the second cluster.
7. The method according to claim 1, wherein, The position of the source of the first component is based on the position of the rendering of the user relative to the position of the representation of the first cluster, and the position of the source of the second component is based on the position of the rendering of the user relative to the position of the representation of the second cluster.
8. The method according to claim 1, wherein Moving the rendering representing the user associated with the input includes: moving the rendering from an original position to the position, wherein the original position indicates that the rendering representing the user has no visual association with the rendering cluster representing the discussion, and wherein the position indicates a visual association between the rendering representing the user and the rendering cluster representing the discussion.
9. A computing device for controlling access to audio and video streams of a communication session, the computing device comprising: one or more processing units; and a computer-readable storage medium encoded with computer-executable instructions that cause the one or more processing units to perform the following operations: Cause a user interface layout including a plurality of renderings to be displayed, each rendering associated with an individual user of a communication session, the user interface layout further including clusters, each cluster representing an individual discussion of a subgroup of users, wherein a rendering cluster represents a discussion among a subset of users represented by individual renderings positioned in association with the cluster, the user interface layout further including a list of individual discussions, each individual discussion associated with a graphical element depicting the individual cluster representing the individual discussion; Receive an input indicating selection of the discussion by interaction with the graphical element associated with the discussion, wherein the graphical element is positioned in association with the description of the discussion on the list; In response to an input selecting the discussion from the list: Move the rendering representing the user associated with the input to a position indicating an association between the rendering representing the user and the rendering cluster representing the discussion among the subset of users, and Modify the access rights of the computing device associated with the user to receive an audio signal from the computing devices of the subset of users having renderings positioned in association with the cluster.
10. The computing device according to claim 9, wherein, The instruction further causes the one or more processing units to perform the following operations: Receive a selection of a second graphical element associated with the discussion, wherein the second graphical element is positioned in association with the description of the discussion on the list; and In response to the selection of the second graphical element of the discussion from the list, modify the access rights of the computing device associated with the user to send an audio signal to computing devices of a subset of the users having a rendering positioned in association with the cluster.
11. The computing device according to claim 9, wherein, The instruction further causes the one or more processing units to perform the following operations: Receive a selection of a second graphical element associated with the discussion, wherein the second graphical element is positioned in association with the description of the discussion on the list; and In response to the selection of the second graphical element of the discussion from the list: Cause a transition from the user interface layout including the plurality of renderings to a second user interface layout including individual renderings of a subset of the users, which has graphical elements depicting the cluster and graphical elements representing other clusters, and Modify the access rights of the computing device associated with the user to send an audio signal to computing devices of a subset of the users having a rendering positioned in association with the cluster.
12. The computing device according to claim 9, wherein, Indicating that the input selecting the discussion includes receiving an input indicating a topic, wherein the method further includes: Determine that the topic is relevant to the discussion, wherein in response to determining that the topic is relevant to the discussion, move the rendering representing the user and modify the access rights of the computing device associated with the user to receive an audio signal from devices of a subgroup of the users.
13. The computing device according to claim 9, wherein, Indicating that the input selecting the discussion includes receiving an input indicating an identifier of a discussion participant, wherein the method further includes: Determine that the discussion participant identified in the input is participating in a discussion with a subgroup of the users, wherein in response to determining that the discussion participant identified in the input is participating in the discussion with the subgroup of the users, move the rendering representing the user and modify the access rights of the computing device associated with the user to receive an audio signal from devices of a subgroup of the users.
14. The computing device according to claim 9, wherein, The instruction further causes the one or more processing units to transmit audio signals from computing devices of users participating in other discussion subgroups to the computing device associated with the user, wherein when there is no visual association with the cluster at the position of the rendering of the user, the audio signals from the computing devices are transmitted simultaneously to the computing device associated with the user, wherein a first component of the audio signal is from a first cluster of users participating in a first discussion, and a second component of the audio signal is from a second cluster of users participating in a second discussion, wherein the volume of the first component is based on the distance between the rendering of the user and the representation of the first cluster, and wherein the volume of the second component is based on the distance between the rendering of the user and the representation of the second cluster.
15. The computing device according to claim 9, wherein, The position of the source of the first component is based on the rendered position of the user relative to the position of the representation of the first cluster, wherein the position of the source of the second component is based on the rendered position of the user relative to the position of the representation of the second cluster.
16. A computer-readable storage medium encoded with computer-executable instructions that cause one or more processing units of a system to perform the following operations: Cause a display to include a plurality of rendered user interface arrangements, each rendering being associated with an individual user of a communication session, the user interface arrangements further including clusters, each cluster representing an individual discussion of a subgroup of users, wherein, Render a discussion among a subset of users whose individual renderings are positioned in association with a cluster, the user interface arrangement further including a list of individual discussions, each individual discussion being associated with a graphical element depicting the individual cluster representing the individual discussion; Receive an input indicating selection of the discussion by interaction with the graphical element associated with the discussion, wherein the graphical element is positioned in association with the description of the discussion on the list; In response to an input selecting the discussion from the list: Move the rendering of the user associated with the input to a position indicating the association between the rendering of the user and the rendered cluster representing the discussion among the subset of users, and Modify the access rights of the computing device associated with the user to receive an audio signal from the computing devices of the subset of users having a rendering positioned in association with the cluster.
17. The computer-readable storage medium according to claim 16, wherein, The instructions further cause the one or more processing units to perform the following operations: Receive a selection of a second graphical element associated with the discussion, wherein the second graphical element is positioned in association with the description of the discussion on the list; and In response to the selection of the second graphical element of the discussion from the list, modify the access rights of the computing device associated with the user to send an audio signal to the computing devices of the subset of users having a rendering positioned in association with the cluster.
18. The computer-readable storage medium according to claim 16, wherein, The instructions further cause the one or more processing units to perform the following operations: Receive a selection of a second graphical element associated with the discussion, wherein the second graphical element is positioned in association with the description of the discussion on the list; and In response to the selection of the second graphical element of the discussion from the list: Cause a transition from the user interface arrangement including the plurality of renderings to a second user interface arrangement including individual renderings of the subset of users, having graphical elements depicting the cluster and graphical elements representing other clusters, and Modify the access rights of the computing device associated with the user to send an audio signal to the computing devices of the subset of users having a rendering positioned in association with the cluster.
19. The computer-readable storage medium according to claim 16, wherein, The input indicating selection of the discussion includes receiving an input indicating a topic, wherein the method further includes: Determine that the topic is relevant to the discussion, wherein in response to determining that the topic is relevant to the discussion, move the rendering representing the user and modify the access rights of the computing device associated with the user to receive an audio signal from the devices of the subgroup of users.
20. The computer-readable storage medium according to claim 16, wherein, Indicating that the selected input of the discussion includes receiving an input indicating an identifier of a discussion participant, wherein the method further includes: Determining that the discussion participant identified in the input is participating in a discussion of a subgroup of the user, wherein in response to determining that the discussion participant identified in the input is participating in the discussion of the subgroup of the user, moving a rendering representing the user and modifying access permissions of the computing device associated with the user to receive an audio signal from a device of the subgroup of the user.