Persistent participant prioritization across communication sessions
By introducing persistent participant priority function across communication sessions in the collaborative system, the problem of insufficient stability and visibility of auxiliary video streams in the collaborative system is solved, and the reliability and participation of the meeting is improved.
Patent Information
- Application Number
- CN202380071358.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-31
- Filing Date
- 2023-09-15
- Publication Date
- 2025-05-13
AI Technical Summary
Existing collaborative systems, in the case of processing where language translation or sign language interpretation is required, have insufficient stability and visibility of auxiliary video streams, resulting in reduced productivity and participation of conference participants.
By introducing persistent participant prioritization functionality across communication sessions in the system, users can set priority for assistants with specific roles (such as sign language interpreters, translators) and automatically store and apply them in user settings, so that the assistant's video stream can be automatically displayed in a specified area throughout the communication session without being interrupted by predetermined events.
It realizes a stable display of auxiliary personnel video streams in the meeting, reduces the trouble of users finding and selecting auxiliary personnel in the meeting, improves the reliability and participation of the meeting, and reduces the inefficient use of computing resources.
Smart Images

Figure CN119999181A_ABST
Abstract
Description
Background Art
[0001] There are many collaboration systems that allow users to communicate. For example, some systems allow people to collaborate by sharing content using video streaming, shared files, chat messages, etc. Some systems also allow people to edit documents simultaneously while also enabling them to communicate using video and audio streams. Users can also establish a communication session at a specific time (e.g., a time slot for an online meeting) and share a live video stream that can show people and content simultaneously.
[0002] Although existing collaboration systems provide feature sets that allow people to conduct meetings via live video streams, some of these systems still have many shortcomings. For example, some existing systems do not have effective features to accommodate people who need language translators or sign language interpreters. In such cases, the meeting attendees will have an assistant, such as a translator or interpreter, join the meeting. The assistant can then listen to the meeting, observe the shared content and video stream, and provide explanations of its observations. For these tasks, it is important for the meeting attendees to have a clear view of their assistants. If the assistant's video rendering moves or resizes during the meeting, it may be difficult for the attendees to keep up with the flow of the meeting. If this happens, important information may be missed.
[0003] Some existing systems provide limited features that limit the movement of video streams. For example, some current solutions allow conference attendees to select video streams, for example, video streams can be "pinned" into position. A "pinned" video is a video that is selected to be fixed to a specific position of a user interface during one or more selection operations. Although this solution is helpful in some cases, there are many instances in which these selected streams can be resized, moved, or removed altogether. In some illustrative examples, when the selected stream is covered or dominated by other prioritization features (such as "spotlight"), or when other events (such as low bandwidth detection, etc.) occur, the selected stream is moved or resized. Spotlight occurs when a conference participant wants to highlight a person in the meeting for others to see. If the conference host spotlights a specific person, the video of some of the participants' assistants may be interrupted. This interruption of the conference assistant's video stream can cause a loss of productivity and engagement for people who rely on their assistants, particularly in situations where sign language interpretation or language translation is required.
[0004] Additionally, when people are pinned as meeting attendees, that selection is only for the video stream for that particular meeting. Therefore, every time people join a meeting, they must locate their assistant's video stream and manually select that stream. This is a cumbersome task that can cause people to miss part of the meeting or cause disruption to other participants. Furthermore, this manual requirement to select an assistant's stream sometimes requires their assistant to be online for them to make the selection. This may require people to wait for his or her assistant to join. This collaboration requires a specific order for people to join the meeting, which interrupts other events and the overall flow of the meeting. These problems are further exacerbated by the fact that people may need multiple assistants. All of these issues can lead to a loss of productivity and engagement, which ultimately leads to an inefficient use of computing resources. Summary of the invention
[0005] The technology disclosed herein enables a system to provide persistent participant prioritization across communication sessions. Users can establish priorities for assistants with specific roles (e.g., sign language interpreters, translators, etc.). This prioritization can be stored in user settings, which cause the system to automatically maintain the display of its assistant's video stream in a specified area of the user interface throughout the communication session. The system also utilizes the user settings so that the assistant's prioritization is maintained across multiple communication sessions. Therefore, each time a user joins a meeting, the display of the assistant's video stream is automatically positioned in the specified area without the need for the user to input a "pin" to the display of another participant. The persistent display of the video stream is not interrupted by predetermined events of the communication session, such as a presenter sharing content, detecting low network bandwidth, detecting an active speaker, a participant joining a session, etc.
[0006] In some configurations, the system can provide multi-level prioritization, enabling some users to apply to "normal pins" and other users to apply to "super pins." User settings can define individual priorities assigned to individual groups of participants. A first priority (e.g., super pin) causes the system to display renderings of a first group of participants (such as sign language interpreters) within a first designated area (e.g., a main stage of a user interface). The user settings can cause the system to limit movement of renderings of the first group of participants in the event of a first category of state change of the communication session, e.g., detection of an active speaker, detection of a user joining a meeting, detection of a user leaving a meeting, etc. The user settings can also cause the system to limit movement of renderings of the first group of participants in the event of a second category of state change of the communication session, detection of low bandwidth, display of shared content, etc.
[0007] A second priority level (e.g., a normal pin) causes the system to display a rendering of a second set of participants (such as select team members) within a second designated area of the user interface (e.g., a secondary surface). The user setting can restrict movement of the second set of participants in response to a first category of status change (e.g., active speaker detected) while allowing movement of the second set of participants in response to a second category of status change (e.g., low bandwidth detected, shared content displayed, etc.) of the communication session.
[0008] The technology disclosed in this article provides many technical benefits. In one example, the technology disclosed in this article provides reliable accessibility features. If a participant of a meeting needs a sign language interpreter, the system is able to maintain the display of its sign language interpreter during multiple interruptions. This has many benefits over ordinary pins. For example, a specific event (e.g., low bandwidth is detected) does not disrupt the display of the video stream of the language interpreter. This allows users to view the interpretation of the content of the meeting with increased reliability on some existing systems. In addition, users do not have to go through the process of selecting a language interpreter to be pinned during a meeting. The automatic selection and persistent display of the assistant eliminates the need for a conference participant to manually identify another user as an assistant and provide input to pin the display of the other user. This can save a lot of computing resources because each time a conference participant joins a meeting, the conference participant does not interrupt the meeting or miss any content.
[0009] By providing participant prioritization across communication sessions, the system can promote user engagement. By promoting user engagement and avoiding user fatigue, especially in communication systems, users can exchange information more effectively. This helps to reduce the occurrence of shared content being missed or ignored when the user is distracted or distracted. The promotion of user engagement and avoiding user fatigue can reduce the occurrence of users needing to extend meetings or resend missed information. More efficient communication of shared content can also help avoid the need for external systems (such as mobile phones for text messaging and other messaging platforms). This can help reduce the reuse of networks, processors, memory or other computing resources. The disclosed technology also uses automation of user settings to provide improved human interaction with the system. This enables the system to be utilized in a more effective manner by reducing the display of unwanted menus, reducing objects selected by mistake, or reducing operations triggered by mistake.
[0010] Features and technical benefits other than those explicitly described above will be apparent by reading the following detailed description and reviewing the associated drawings. This summary is provided to introduce a selection of concepts in simplified form, which will be further described in the specific embodiments below. This summary is not intended to identify the key or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. For example, the term "technology" may refer to (one or more) systems, (one or more) methods, computer-readable instructions, (one or more) modules, algorithms, hardware logic and / or (one or more) operations allowed by the context described above and the entire document. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The detailed description is described with reference to the accompanying drawings. In the drawings, the leftmost digit(s) of a reference number identifies the drawing in which the reference number first appears. The same reference numbers in different drawings indicate similar or identical items. References to individual items in a plurality of items can use reference numerals with letters of the alphabetical sequence to refer to each individual item. Generic references to items can use specific reference numerals without the alphabetical sequence.
[0012] Figure 1A Shown is a persistent display of a video of an assistant automatically configured in a designated area of a user interface in response to determining that a role of the assistant corresponds to a prerequisite of a user associated with the user interface.
[0013] Figure 1B A video of the assistant is shown persistently displayed over the course of other participants joining the communication session.
[0014] Figure 1C It shows how users of a communication session can be pinned into positions in a user interface.
[0015] Figure 1D A video of the assistant being persistently displayed through a process in which a user performs a common pin on another user is shown.
[0016] Figure 1E A video of the assistant is shown persistently displayed, while other videos of other users are rearranged based on the other users' voice activity.
[0017] Figure 1F It shows how a second user of a communication session can be pinned into a position in a user interface.
[0018] Figure 1G A video showing how the assistant can be persistently displayed through the process of a second user being pinned.
[0019] Figure 1H It shows how to display the assistant's video persistently while turning other videos of other users into still images in response to detecting low bandwidth.
[0020] Figure 2 An example is shown of how user A's video may be persistently displayed while other videos of other users are being transitioned during content sharing.
[0021] Figure 3 An example is shown of how user A's video is persistently displayed and moved while other videos of other users are transitioned during content sharing.
[0022] Figure 4 An example of a user settings and mapping data structure is shown.
[0023] Figure 5 is a flow chart illustrating aspects of a routine for utilizing the disclosed techniques.
[0024] Figure 6 is a computer architecture diagram illustrating an illustrative computer hardware and software architecture of a computing system capable of implementing various aspects of the techniques and technologies presented herein.
[0025] Figure 7 is a computer architecture diagram illustrating a computing device architecture of a computing device capable of implementing various aspects of the techniques and technologies presented herein. DETAILED DESCRIPTION
[0026] Figures 1A to 1H Various aspects of a system 100 for providing participant prioritization across communication sessions are illustrated. The system utilizes settings that are maintained across multiple communication sessions for people participating in several meetings. The settings define individual prerequisites for participants. For example, the settings can define prerequisites that indicate that a participant needs a sign language interpreter or that a participant needs hearing assistance. Whenever a communication session (e.g., a meeting) begins, the system can access the settings so that the system can identify assistants for the participants before the start of each meeting. This avoids the need for participants to search for and select assistants that can help the participant's prerequisites. This avoids the need for participants to collaborate with other users and alleviates the need for participants to perform multiple manual inputs, which can distract all participants and be inefficient with respect to computing resources.
[0027] In some embodiments, the system is able to identify a particular person as an assistant by identifying a role that corresponds to the participant's prerequisites. For example, if the participant's prerequisites indicate that the participant is hearing-impaired, and another person has a role as a "sign language interpreter," the system can identify the other person as the participant's assistant. In response to selection of the assistant, the system then automatically displays a live video stream of the assistant on a user interface rendered on the participant's computer.
[0028] The communication session can be in the form of an online conference, a broadcast, or any other gathering that includes a start time and an end time. Figure 1A As shown in , the communication session can be managed by a system 100 comprising a plurality of computers 11 , each computer corresponding to an individual user 10 . For purposes of illustration, a first user 10A, Mike Taylor, is associated with a first computer 11A, a second user 10B, Tracki Isaac, is associated with a second computer 11B, a third user 10C, Doug Wright, is associated with a third computer 11C, a fourth user 10D, MJ Price, is associated with a fourth computer 11D, a fifth user 10E, Kat Martin, is associated with a fifth computer 11E, a sixth user 10F, Miguel Jones, is associated with a sixth computer 11F, a seventh user 10G, Kristal McKinney, is associated with a seventh computer 11G, an eighth user 10H, Jessica Kline, is associated with an eighth computer 11H, a ninth user 10I, Monica Larsson, is associated with a ninth computer 11I, a tenth user 10J, Charlotte Davis, is associated with a tenth computer 11J, an eleventh user 10K, Annika Andersson, is associated with an eleventh computer 11K, and a twelfth user 10L, Isla Scogins, is associated with a twelfth computer 11L. The users may also be referred to as "user A", "user B", etc., respectively.
[0029] Each user can be displayed in the user interface as a two-dimensional 2D image, or each user can be displayed in the user interface as a three-dimensional representation, such as an avatar. The 3D representation can be a static model or a dynamic model that is animated in real time in response to user input. Although this example illustrates a user interface in which users are displayed as 2D images, it can be appreciated that the technology disclosed herein can be applied to other forms of representation, video, or other types of rendering. The computer 11 can be in the form of a desktop computer, a head-mounted display unit, a tablet computer, a mobile phone, etc. The system can generate a user interface showing various aspects of a communication session with each of the users. In Figure 1A In the example of FIG. 1 , the first user interface arrangement 101A can include multiple renderings of one or more users 10. The renderings can include renderings of two-dimensional (2D) images, which can include pictures or live video feeds of users.
[0030] In this example, the user interface is presented on a display device associated with the tenth user 10J, Charlotte Davis, of the tenth computer 11J. Charlotte is referred to herein as the "viewer" of the user interface displayed on the tenth computer 11J. The first user interface arrangement 101A includes a first area 120A, also referred to herein as a designated area 120A or a primary tabletop 120A. The first user interface arrangement 101A also includes a second area 120B, also referred to herein as a secondary area 120B or a secondary tabletop 120B. The first user interface arrangement 101A also includes another rendering of a video stream 151J, which shows a self-view of the tenth user 10J. The video stream 151J can be displayed in the second area 120B and is restricted to be displayed in the first area 120A. The first area is reserved only for video streams of users with roles corresponding to the prerequisites of the viewer. The viewer in this example is the tenth user 10J of the tenth computer 11J.
[0031] When a tenth user 10J (User J) joins the communication session, the system automatically accesses the preferences of User J. The preferences can indicate that User J requires assistance, for example, User J indicates that they are hearing impaired and require a prerequisite for assistance. In response to the indication, the system can cause a rendering of a video stream 151A of a selected user (e.g., the first user 10A (User A)) to be displayed within a designated area 120A of the user interface 101A. Figure 2 Describing in more detail, in response to determining that user A has a prerequisite role (eg, sign language interpreter) corresponding to user J (eg, an indication that user J is hearing impaired), user A can be selected for display in primary tabletop 120A.
[0032] Figures 1B to 1H An example of how the assistant can be persistently displayed in a designated area even when the communication session transitions through different types of state changes is shown. For example, Figure 1BIt is shown how to persistently display the assistant's video through the process that other participants are joining the communication session. In this example, when other participants join the communication session, the rendering of each user can be rearranged in the secondary area 120B. When each participant joins, a part of the secondary area 120B (e.g., the left side) can be filled with the rendering of the participant with the live video stream. This can include a rendering 151B of the second user 10B, a rendering 151C of the third user 10C, and a rendering 151D of the fourth user 10D. Another part of the secondary area 120B (e.g., the middle) can be filled with a still image of a user working with a computer that does not generate a live video stream. These users are referred to as audio-only users in this article. For example, the fifth user 10E is displayed as a representation 152E in text form, and the sixth user 10F is displayed as a representation 152F in still image form. As these other users join the conference, the first user 10A is persistently displayed in the primary console 120A, and the rendering of the first user 10A is not moved even though the other users may be rearranged as each individual joins the conference.
[0033] Figure 1C and Figure 1D It shows how a video of an assistant can be persistently displayed through a process where a user J performs a normal pin on another user D. Figure 1C 1 shows how the fourth user 10D of the communication session is pinned to a fixed position in the user interface. In this example, the user selects the menu item "Pin for me" and as shown in Figure 1D , this input causes the system to position the rendering 151D of the fourth user 10D in a different location in the secondary area 120B than the other users. The user is also able to select "Pin for Everyone" (a "spotlight" of the video) and cause the system to pin the selected video of the local computer and all other participants in the meeting. For purposes of illustration, the spotlighted video is also referred to herein as a video with a "normal pin." When the system selects a video using a normal pin, such as "Pin for Me" or "Pin for Everyone," the system locks the video in the secondary deck for a specific action, which is described below as a first category of state change. When the system uses a normal pin to position a video, the video has a second priority over any video that is persistently displayed in the primary deck 120A using a super pin. Thus, as shown in Figure 1D As shown in , videos selected as normal pins will not result in and are limited to being persistently displayed as moved, resized, or removed videos.
[0034] Once selected for normal pinning, the pinned video (e.g., the rendering 151D of the fourth user 10D) is restricted from being moved or resized in response to detecting a first category of state changes. The first category of state changes can include state changes involving detection of an active speaker. When an active speaker is detected, the system can rearrange other video renderings that are not pinned. The rearrangement of other video renderings can involve moving the most active speaker to a prominent position within the secondary surface 120B. For example, as in Figure 1D and Figure 1E As shown in the transition between, when the system detects that user G has a threshold communication level and user H and user I do not have a threshold communication level, user G moves from the left side of the queue to the right side of the queue. The threshold communication level can include a threshold speech volume, a threshold number of words, a threshold word rate (number of words per unit time), a threshold word count, a threshold level of physical movement captured by a camera, or a combination of these factors. When user G meets the criterion, user G causes the video of the non-pin to be moved. Although some video renderings (such as renderings of user B, user C, and user G) are rearranged, the video of user D's pin remains fixed because it is restricted from being moved or resized in this type of state change. In addition, in response to a state change in which a state change of the first category is detected, user A, who is super pinned as an assistant, is also restricted from being moved or resized.
[0035] Figure 1F and Figure 1G A video showing how the assistant is persistently displayed through a process in which user J applies another normal pin to another user (user G), similar to the example described above, once the normal pin is applied to user G's rendering 151G, the rendering is fixed in place in response to detecting a first category of status change (e.g., active speaker, user joining the conversation, etc.). Regardless of the number of renderings fixed in place using normal pins, any rendering fixed using super pins is restricted from moving or is restricted from decreasing to a size less than a threshold size limit.
[0036] Reference now Figure 1H , an example showing another difference between a normal pin and a super pin is shown and described below. In this example, when the system detects a state change from a state change of the second category, the system allows the video rendering subject to the normal pin to be moved, resized or removed. However, when the system detects a state change of this type, the video subject to the super pin is restricted from moving, resizing or removing. In an illustrative example, the state change of the state change of the second category can involve detecting one or more computers experiencing connectivity problems, for example, detecting low network bandwidth.
[0037] Figure 1H How to display the video of user A permanently, while other videos of other users are removed and replaced with still images in response to detecting low bandwidth. In this example, the system detects that at least one computer of the system is experiencing connectivity problems. This can include low bandwidth between each computer in the computer or low bandwidth relative to at least one computer 11 of the system 100. When the bandwidth of at least one computer drops below the threshold, the system limits the video of user A to be moved, removed or resized. However, when the bandwidth of at least one computer drops below the threshold, the system allows the live video feed of the user to be moved, reduced or removed in the secondary table. Even when the video is kept under the ordinary pin, this may also occur. This particular example shows that the live video feed of the user in the secondary table 120B is removed and replaced by other representations, such as their initial image or a still image showing the user. However, in this type of event, for example, when low bandwidth is detected, the live video feed that is persistently displayed in the main table 120A is maintained.
[0038] Reference now Figure 2 and Figure 3 , other examples showing additional differences between ordinary pins and super pins are shown and described. In these examples, when the system detects a state change from a state change of the second category, the system allows video renderings that are not pinned or subject to ordinary pins to be moved, resized or removed. However, when the system detects this type of state change, videos subject to super pins are restricted from moving, resizing or removing. In an illustrative example, the state change of the second category of state changes may involve detecting one or more computers that share content with the communication session.
[0039] Figure 2How the video of user A is persistently displayed, while other videos of other users are converted in response to detecting shared content. In this example, one of the presenters (e.g., user G) starts sharing a slide set. In response to detecting shared content, the system maintains the position of the assistant's rendering within the main table 120A. The aspect ratio of the assistant's rendering can be modified, but the position of at least one selected point of the image (e.g., the lower left corner of the image or the upper left corner of the image) is maintained. The size of the assistant's rendering can also be reduced to a threshold size to allow viewers to have a concurrent view of the rendering 201 of the shared content and the assistant's rendering. Other live stream renderings specified with ordinary pins or live stream renderings that are not pinned are repositioned and / or resized. For example, user D and user G are moved to the lower right portion of the user interface. Other live stream renderings that do not include videos specified with ordinary pins or super pins are repositioned, resized or removed.
[0040] Figure 3 Another example of how to persistently display user A's video while transitioning other users' other videos in response to detecting a second category of state change (e.g., detection of shared content) is shown. In this example, one of the presenters (e.g., user G) begins sharing a slideshow. In response to detecting shared content, the system maintains the minimum size of the assistant's rendering and moves the rendering 151A to a new location. Thus, although the persistently displayed user is moved and resized, it is restricted from being reduced beyond a predetermined minimum size. The minimum size is also referred to herein as the "threshold minimum size" of the rendering 151A for the assistant 10A or a user having a role corresponding to a prerequisite for viewing the user. This restriction on reduction allows the user to clearly understand and view the assistant while simultaneously displaying shared content.
[0041] In some embodiments, the threshold minimum size can be based on the device type or screen size. For example, for a desktop computer or a device with a twenty-two-inch monitor, the threshold minimum size can be a predetermined percentage of the screen, for example, 50% of one dimension of the screen. However, for a tablet or mobile device or a device with a five-inch screen, the threshold minimum size can be a larger predetermined percentage of the screen, for example, 90% of one dimension of the screen. Although the screen dimension is used in this example, other units of measurement can be utilized. For example, if the device screen has less than a threshold number of pixels (e.g., two million), the system can use a first threshold minimum size; and if the device screen has more than a threshold number of pixels, the system can use a second threshold minimum size for the rendering of the assistant 10A, wherein, in this example, the first threshold minimum size is greater than the second threshold minimum size. The threshold minimum size can be based on the number of pixels or one or more dimensions of the rendering. This allows the system to use more of the screen for rendering the assistant 10A for devices with smaller screens. The threshold minimum size can be applied to any rendering 151A of the assistant 10A or a user with a role corresponding to the prerequisite of the viewing user.
[0042] The rendering of the assistant has a higher priority over the rendering of other video streams. An assistant or user (such as user A) selected to be persistently displayed is defined herein as a user with one or more prerequisite roles corresponding to the viewing user (e.g., user J, the viewer of the user interface). Movement of the persistent display can also be limited to a threshold distance. Thus, when an assistant is persistently displayed in the primary deck 120A, the system limits movement of the rendering to a threshold vertical distance or a threshold horizontal distance. Such an embodiment allows a viewer to maintain a view of their designated assistant. Figure 3 1 and 2. It is also shown that other live stream renderings specified by ordinary pins or other renderings that are not pinned are relocated, resized and / or removed. For example, the renderings of user D and user G can be moved to the lower right portion of the user interface. Other live stream renderings that are not pinned are relocated, resized or removed. For example, user H or user I can be moved to the lower right portion of the user interface or removed together.
[0043] In addition to the differences described above, there are many other differences between regular pins and super pins. For example, the video selected for a regular pin is not retained across communication sessions. Users must select a video for a regular pin for each meeting they join. However, the video selected for a super pin is retained across communication sessions. Persistence across communication sessions can be achieved by storing data that associates a particular assistant with a participant. For example, once user A is identified as user J's assistant, user A's identity can be stored in a setting indicating that user A is user J's assistant, and whenever user J joins a meeting, the setting is accessed and the user interface displayed to user J will automatically display a rendering of user A in the designated area 120A.
[0044] In other embodiments, such as in Figure 4 As shown in , persistence across communication sessions can be achieved by storing user settings 400 that associate the prerequisites 410 with one or more user identities of the individual user 10 . Figure 4 An example of a user setting 400 and a mapping data structure 401 is shown. Typically, the user settings can define prerequisites for a particular user. For example, user J has a prerequisite indicating that they need hearing assistance, user B has a prerequisite indicating that they need a French translator, and user C has a first prerequisite indicating that they need hearing assistance, for example, listing that they are hearing impaired. Each user can have multiple assistants. For example, some users can indicate that they need two sign language interpreters. The mapping data structure of 401 can define multiple profiles 430. Each profile can identify the role of each user. For example, user K can participate in a meeting as a Spanish translator, coordinator, or administrator; user L can participate in a meeting as a French translator or a Spanish translator; and user A can participate in a meeting as a sign language interpreter.
[0045] The settings 400 are stored in a manner that allows the system 100 to access the settings each time a user joins a meeting. When a meeting participant (such as user J) joins a meeting, the system accesses the user settings and determines whether one of the prerequisites 410A associated with the meeting participant (user 10J) corresponds to a role 410A of another user (such as user A). When the role of a particular user is determined to correspond to the prerequisite of the meeting participant, the system selects the particular user as an assistant for the meeting participant. The system then persistently displays a rendering of the assistant in a user interface, as described herein.
[0046] Settings 400 can also be referred to herein as "accessibility settings." The accessibility settings can be any data structure, document, or other form of data that defines a person's needs and associates those needs with their identity. For example, profile or registration data at the operating system level or application level can indicate that a user has drivers and devices for specific accessibility needs. This data can be used to indicate that a person has a prerequisite for hearing impairment, etc. In another example, email or communication data indicating a person's accessibility needs can also be used to indicate the user's prerequisites and invoke the operations disclosed herein. If the user has a specific application installed on their phone, for example, a sign language application, such data can also be used for the user's prerequisites and invoke the operations disclosed herein. Settings 400 can be at any stack or level, such as OS, user profile, application level, etc.
[0047] In some embodiments, the system can limit the number of assistants for a particular user. For example, the system can limit user J to only two assistants. This limitation allows the system to provide an understandable display for each assistant, as a large number of assistants may result in a small rendered display that is difficult to see.
[0048] Figure 5 5 is a diagram illustrating aspects of a routine 500 for providing persistent participant prioritization across communication sessions. It should be understood by those of ordinary skill in the art that the operations of the methods disclosed herein are not necessarily presented in any particular order, and that it is possible and contemplated to perform some or all of the operations in an alternative order. For ease of description and illustration, the operations have been presented in the order presented. Operations may be added, omitted, performed together, and / or performed simultaneously without departing from the scope of the appended claims.
[0049] It should also be understood that the illustrated method can start or end at any time and does not need to be performed as a whole. Some or all of the operations of the method and / or substantially equivalent operations can be performed by executing computer-readable instructions included on a computer storage medium, as defined herein. As used in the specification and claims, the term "computer-readable instructions" and its variants are widely used herein to include routines, applications, application modules, program modules, programs, components, data structures, algorithms, etc. Computer-readable instructions can be implemented on various system configurations, including single-processor or multi-processor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based programmable consumer electronics, combinations thereof, etc. Although the example routines described below operate on a system (e.g., one or more computing devices), it is appreciated that the routines can be executed on any computing system, which may include any number of computers working together to perform the operations disclosed herein.
[0050] Therefore, it should be appreciated that the logical operations described herein are implemented as a sequence of computer-implemented actions or program modules running on a computing system such as described herein and / or as interconnected machine logic circuits or circuit modules within the computing system. The implementation is a matter of choice depending on the performance and other requirements of the computing system. Therefore, the logical operations may be implemented in software, firmware, dedicated digital logic, and any combination thereof.
[0051] In addition, Figure 5 The operations illustrated in the other figures can be implemented in association with the example user interfaces and systems described herein. For example, the various devices and / or modules described herein can generate, send, receive, and / or display data associated with the content of a communication session (e.g., live content, broadcasted events, recorded content, etc.) and / or a rendering UI including a rendering of one or more participants of a remote computing device, avatar, channel, chat session, video stream, image, virtual object, and / or application associated with the communication session.
[0052] Routine 500 includes operation 501, wherein the system detects that a user of a computer joins a communication session. The communication session can be in the form of an online meeting, and a user such as user J can join using a communication application configured to display a live video stream of multiple users.
[0053] At operation 503, the system accesses settings 400 maintained across multiple communication sessions for user 10J. The settings define individual prerequisites 410 for user 10J. Access to the settings is performed automatically by the system without user input. Use of the settings causes the system to select at least one selected user 10A having a role 420A corresponding to at least one prerequisite 410A for user 10J. In response to user 10J joining the communication session, the system 100 automatically performs the selection. Figure 4 An example of the arrangement is shown in FIG.
[0054] At operation 505, the system analyzes a data structure 401 that associates individual users (e.g., user 10A, user 10K, and user 10J) with one or more roles 420. Analysis of the data structure enables the system to identify one or more user profiles 430A of at least one selected user 10A having a role 420A that corresponds to at least one prerequisite 410A of user 10J. For example, a role such as a sign language interpreter may not be determined to correspond to a particular prerequisite, such as requiring hearing assistance.
[0055] Operation 505 can involve keyword matching, pseudo-keyword matching or by using other types of matching processes that can include heuristic-based operations. For example, keywords can be stored in association with a person's role, and when the role has a keyword match or a threshold context match with a participant's prerequisite, the person can be assigned as a persistently displayed assistant for the participant. In some configurations, the system can determine that a candidate's role corresponds to at least one prerequisite of the user by identifying at least one of the keyword matches between the role and at least one prerequisite. Additionally or alternatively, the system can perform phrase matching between the role and at least one prerequisite. Additionally or alternatively, character matching between the role and at least one prerequisite can be performed by applying historical data to heuristic-based operations. Therefore, the matching between the user's prerequisite and the role corresponding to the prerequisite can be improved by iteration of routine 500.
[0056] Machine learning techniques can be utilized so that future iterations of routine 500 can improve the discovery and selection of each assistant. The user can provide feedback to the system by approving or rejecting each selected assistant, for example, they can be rated. This feedback can be used to refine the process of selecting one or more assistants in each iteration of the routine.
[0057] At operation 507, the system enables persistent display of the selected assistant. In some configurations, the system enables display of a rendering of at least one video stream 151A of a selected user 10A within a designated area 120A of a user interface 101A rendered on a device 11J associated with a user 10J having at least one prerequisite 410A, the rendering corresponding to the stream of the at least one selected user, which is not displayed on display devices of other users (e.g., users 10A to 10I and user 10L) that do not have prerequisites corresponding to the role of the at least one selected user. Thus, persistent display of the selected assistant may be displayed only for users having prerequisites corresponding to the role of the selected assistant, and other users who do not have prerequisites corresponding to the role of the selected assistant are restricted from persistent display of the assistant. Thus, in Figure 1A-Figure 4 In the example of , Charlotte's computer 11J persistently displays a rendering of Mike 10A's video. However, since Tracki and Doug do not have a requirement for a sign language interpreter, their respective computers 11B and 11C do not persistently display Mike's rendering in the designated area.
[0058] For purposes of illustration, the "persistent display" of the video remains displayed in the event of a first category of state changes and a second category of state changes for the communication session. The first category of state changes can include detecting an active speaker, detecting a user joining a meeting, detecting a user leaving a meeting, etc. The second category of state changes can include detecting low bandwidth for the communication session. Detecting low bandwidth for the communication session occurs when one or more of the computers 11 detects a network connection that is below a threshold network transmission rate. The second category of state changes can also include detecting shared content. This can occur when a participant in a meeting shares content for display on the computing devices of other participants.
[0059] In the case of a state change from a state change of the first category, the video stream selected for ordinary pins can be restricted from being moved, resized or removed. However, in the case of a state change from a state change of the second category, the video stream selected for ordinary pins can be moved, resized or removed. Ordinary pins can also include "spotted" videos. When a person selects a video and a selection from one user pins the video on the device of another user, the spotlighting of the video occurs. Therefore, if user J spotlights user G's video, user G's video will be pinned to a fixed position in the secondary table on the computing device of the other participants in the communication session. When a spotlighted video is detected, the system will give priority to the super-pinned video over the spotlighted video. Therefore, when the spotlighted video is introduced into the user interface, the system restricts the super-pinned video from being moved, resized or removed from the user interface, wherein the super-pinned video is a video of the assistant displayed in response to determining that the assistant has a role corresponding to the prerequisite of the user viewing the user interface.
[0060] The system can also display multiple assistants for a particular user. For example, two sign language interpreters can be assigned to user J. This can occur if a user preference indicates the need for two sign language interpreters. In this case, the main table will be populated with two video renderings, and each video rendering will be permanently displayed. These video renderings will not be interrupted by the state changes described in this article, which include detecting active speakers, detecting new people joining the meeting, detecting low bandwidth, and detecting content sharing.
[0061] In some configurations, the system can limit the number of assistants a particular user can have. For example, even when the user has the prerequisites for three sign language interpreters, the system can be configured to limit the display to two sign language interpreters. This limitation provides a technical benefit of the fact that the system can maintain the persistent display of each assistant at an appropriate size. In some embodiments, the limit can vary based on the device type or screen size. For example, for a desktop computer, the limit may be three assistants, but for a smaller device such as a tablet or phone, the limit may be two assistants.
[0062] At operation 509, at the end of the communication session, for example, after the end of the conference, the routine can enter a waiting state until the user (e.g., user J) joins another conference. At this point, the routine returns to operation 501, where another user is selected as an assistant.
[0063] For the purpose of illustration, the persistent display of the selected video can also mean that the video is not moved and / or resized beyond a threshold level. Therefore, the persistent display of the video can be moved but not exceed a threshold distance, or the persistent display of the video can be resized but not below a threshold size. This means that the persistent display of the video can be reduced, but the reduction is limited. The reduction limit can be a percentage of the screen, for example, not less than 50% of the screen or user interface. The reduction limit can include a limit of pixel measurement or any other type of limit to keep the desired size of the persistent rendering.
[0064] In some embodiments, some renderings are restricted from movement. For purposes of illustration, a rendering that "moves" in response to detecting a change in state includes moving the rendering from a first position to a second position. This can include selecting a point in the rendering, for example, the lower left corner of the video rendering, determining the location (X, Y) of the point using a coordinate system, and then moving the point to a new position. Thus, when a rendering is restricted from moving, the selected point of the rendering can be restricted from moving from a first position to a second position. Thus, the rendering can be resized, for example, the number of pixels can be reduced, but the rendering can be restricted from moving relative to the selected point, for example, the lower left corner remains fixed in its original position. When it comes to the rendering of an assistant (e.g., a person with a role corresponding to the viewer's prerequisite), the video can be reduced in size to a minimum size but not reduced beyond the minimum size.
[0065] One or more of the operations described above can also include a data structure with a preference hierarchy (stack / level) where the viewer can select a first set of participants ("super pins" for language interpreters) that are locked into the primary deck. The viewer can also select a second set of participants (normal "pins") that are locked into the secondary deck. Participants with normal pins are locked in place as others who are speaking rotate in the secondary deck. However, participants with normal pins can be moved or removed in response to specific events (e.g., low bandwidth, shared content, etc.).
[0066] The operations can include accessing a data structure (201) defining individual priority levels for individual groups (110) of participants (11) assigned to one or more communication sessions (user settings), accessing the data structure maintained across multiple communication sessions, wherein a first priority level causes the system (100) to display a rendering of a first group of participants (110A) (interpreters) having a specified role within a first designated area (120A) of the user interface (101), the data structure defining the priority levels limited in response to a state change of the communication session (interpreters locked in state change: low bandwidth conference). A permission to restrict movement of renderings of the first group of participants (110A) within a first designated area (120A), a second priority level causing the system to display renderings of the second group of participants (110B) within a second designated area (120B), wherein the permission allows movement or removal of renderings of the second group of participants (110B) within the second designated area (120B) in response to a change in the state of the communication session (users of ordinary pins are able to move or be removed during a state change: low bandwidth), and the permission restricts movement of renderings of the second group of participants (110B) during other state changes. Users of ordinary pins can be moved based on user activity: active speaker promoted, etc. Similarly, in some embodiments, the video of an ordinary pin is smaller than the video of a super pin, such as an explainer. Individuals of ordinary pins including spotlight video are not maintained in different meetings. Users set the video of ordinary pins in each meeting.
[0067] The operations also include causing a display of the first user interface arrangement (101A) to include a rendering of a first group of participants (110A) having a first priority, a rendering of a second group of participants (110B) having a second priority, and a rendering of other participants (110C), wherein the rendering of the first group of participants (110A) is positioned in a first designated area (120A), and the second group of participants (110B) is positioned in a second designated area (120B), and the rendering of the other participants (110C) is positioned in the second designated area (120B). This is shown in the figure as the sign language interpreter "super-pinned" in the primary deck 120A, and the other pinned users are in the secondary deck 120B.
[0068] The operations also include receiving input indicating a change in the state of the communication session, wherein the change in state includes at least one of: an indication that at least one participant is sharing content for display on a device of a participant (11) of the communication session or that a bandwidth of the communication session is below a bandwidth threshold.
[0069] The operations also include using the permissions of the data structure (201) to cause a transition from the first user interface arrangement (101A) to the second user interface arrangement (101B) to include rendering of a first group of participants (110A) having a first priority, wherein the rendering of a first participant (151A) in the first group of participants (110A) is maintained at least a threshold size within a first designated area (120A), wherein the rendering of the second group of participants (110B) and the rendering of the other participants (110C) are resized to accommodate the state change (content 201 in Figure 2 A / 2B is shown, low bandwidth or content sharing), where the arrangement of the renderings of the second group of participants (110B) is maintained within the second designated area (120B), while the renderings of other participants (110C) are rearranged due to user activity.
[0070] Figure 6 600 is a diagram illustrating an example environment 600 in which a system 602 can implement the technology disclosed herein. It should be appreciated that the subject matter described above can be implemented as a computer-controlled device, a computer process, a computing system, or an article such as a computer-readable storage medium. The operation of the example method is illustrated in individual blocks and is summarized with reference to these blocks. The method is illustrated as a logical flow of blocks, each of which can represent one or more operations that can be implemented in hardware, software, or a combination thereof. In the context of software, the operation represents a computer-executable instruction stored on one or more computer-readable media, which, when executed by one or more processors, enables the one or more processors to perform the operation.
[0071] Typically, computer executable instructions include routines, programs, objects, modules, components, data structures, etc. that perform specific functions or implement specific abstract data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be performed in any order, combined in any order, subdivided into multiple sub-operations, and / or performed in parallel to implement the described processes. The described processes can be performed by resources associated with one or more devices such as one or more internal or external CPUs or GPUs, and / or one or more hardware logic such as a field programmable gate array ("FPGA"), a digital signal processor ("DSP"), or other types of accelerators.
[0072] All methods and processes described above may be embodied in software code modules executed by one or more general-purpose computers or processors and fully automated therethrough. The code modules may be stored in any type of computer-readable storage medium or other computer storage device, such as those described below. Some or all of the methods described may alternatively be embodied in dedicated computer hardware, such as the computer hardware described below.
[0073] Any routine description, element or block in the flowcharts described herein and / or depicted in the accompanying drawings should be understood to potentially represent a module, segment or portion of code, which includes one or more executable instructions for implementing a specific logical function or element in the routine. Alternative implementations are included within the scope of the examples described herein, where elements or functions may be deleted or performed from the order shown or discussed, including substantially synchronously or in reverse order, depending on the functionality as will be understood by those skilled in the art.
[0074] In some implementations, system 602 can be used to collect, analyze, and share data displayed to users of communication session 604. As illustrated, communication session 603 can be implemented between a plurality of client computing devices 606(1)-606(N) (where N is a number having two or greater values) that are associated with or part of system 602. Client computing devices 606(1)-606(N) enable users, also referred to as individuals, to participate in communication session 603.
[0075] In this example, communication session 603 is hosted by system 602 over one or more networks 608. That is, system 602 can provide services that enable users of client computing devices 606(1) to 606(N) to participate in communication session 603 (e.g., via live viewing and / or recorded viewing). Thus, "participants" of communication session 603 can include users and / or client computing devices (e.g., multiple users can be in a room participating in a communication session via the use of a single client computing device), each of which can communicate with other participants. Alternatively, communication session 603 can be hosted by one of client computing devices 606(1) to 606(N) using peer-to-peer technology. System 602 can also host chat conversations and other team collaboration functionality (e.g., as part of an application suite).
[0076] In some implementations, such chat conversations and other team collaboration functions are considered external communication sessions distinct from the communication session 603. The computing system 602 that collects participant data in the communication session 603 may be able to link to such an external communication session. Thus, the system may receive information that enables connection to such an external communication session, such as date, time, session details, etc. In one example, a chat conversation can be conducted in accordance with the communication session 603. Additionally, the system 602 may host a communication session 603 that includes at least a plurality of participants co-located at a meeting location (such as a conference room or auditorium) or located at different locations.
[0077] In the example described herein, client computing devices 606(1) to 606(N) participating in a communication session 603 are configured to receive and render communication data for display on a user interface of a display screen. The communication data can include a collection of various instances or streams of live content and / or recorded content. The collection of various instances or streams of live content and / or recorded content can be provided by one or more cameras (such as a video camera). For example, individual streams of live or recorded content can include media data associated with a video feed provided by a video camera (e.g., audio and visual data capturing the appearance and voice of users participating in the communication session). In some embodiments, the video feed can include such audio and visual data, one or more still images, and / or one or more avatars. The one or more still images can also include one or more avatars.
[0078] Another example of an individual stream of live or recorded content can include media data including avatars of users participating in a communication session and audio data capturing the user's voice. Yet another example of an individual stream of live or recorded content can include media data including a file displayed on a display screen and audio data capturing the user's voice. Thus, the various streams of live or recorded content within the communication data enable remote conferencing to be facilitated between a group of people and content to be shared within a group of people. In some embodiments, the various streams of live or recorded content within the communication data can originate from a plurality of co-located video cameras positioned in a space, such as a room, to record or stream a presentation including one or more individuals making the presentation and one or more individuals consuming the presented content.
[0079] Participants or attendees can view the content of the communication session 603 as the activity occurs or alternatively via a recording at a later time after the activity occurs. In the example described herein, the client computing devices 606(1) to 606(N) participating in the communication session 603 are configured to receive and render communication data for display on a user interface of a display screen. The communication data can include a collection of various instances or streams of live and / or recorded content. For example, an individual stream of content can include media data associated with a video feed (e.g., audio and visual data that captures the appearance and voice of users participating in the communication session). Another example of an individual stream of content can include media data that includes an avatar of a user participating in the conference session and audio data that captures the user's voice. Yet another example of an individual stream of content can include media data that includes content items displayed on a display screen and / or audio data that captures the user's voice. Thus, the various streams of content within the communication data enable facilitating a conference or broadcast presentation between a group of people dispersed in remote locations.
[0080] Participants or attendees of a communication session are people within range of a camera or other image and / or audio capture device, so that the person's movements and / or sounds generated when the person is viewing and / or listening to the content shared via the communication session can be captured (e.g., recorded). For example, a participant may be sitting in a crowd viewing the shared content live at a broadcast location where a stage presentation is taking place. Alternatively, a participant may be sitting in an office conference room, viewing the shared content of a communication session with other colleagues via a display screen. Further, a participant may be sitting or standing in front of a personal device (e.g., a tablet, smart phone, computer, etc.) viewing the shared content of the communication session individually in their office or home.
[0081] Figure 6System 602 includes device(s) 610. Device(s) 610 and / or other components of system 602 can include distributed computing resources that communicate with each other and / or with client computing devices 606(1)-606(N) via one or more networks 608. In some examples, system 602 can be a stand-alone system responsible for managing various aspects of one or more communication sessions, such as communication session 603. As an example, system 602 can be managed by an entity such as SLACK, WEBEX, GOTOMEETING, GOOGLE HANGOUTS, etc.
[0082] The network(s) 608 may include, for example, a public network such as the Internet, a private network such as an institutional and / or personal intranet, or some combination of private and public networks. The network(s) 608 may also include any type of wired and / or wireless network, including, but not limited to, a local area network ("LAN"), a wide area network ("WAN"), a satellite network, a cable network, a Wi-Fi network, a WiMax network, a mobile communication network (e.g., 3G, 4G, etc.), or any combination thereof. The network(s) 608 may utilize communication protocols, including packet-based and / or datagram-based protocols, such as the Internet Protocol ("IP"), the Transmission Control Protocol ("TCP"), the User Datagram Protocol ("UDP"), or other types of protocols. In addition, the network(s) 608 may also include a plurality of devices that facilitate network communications and / or form the hardware foundation of the network, such as switches, routers, gateways, access points, firewalls, base stations, repeaters, backbone devices, etc.
[0083] In some examples, network(s) 608 may also include devices that enable connection to a wireless network, such as a wireless access point (“WAP”). Examples support connection via a WAP that sends and receives data via various electromagnetic frequencies (e.g., radio frequencies), including WAPs that support the Institute of Electrical and Electronics Engineers (“IEEE”) 802.11 standards (e.g., 802.11g, 802.11n, 802.11ac, etc.), as well as other standards.
[0084] In various examples, (one or more) devices 610 may include one or more computing devices that operate in a cluster or other grouping configuration to share resources, balance loads, increase performance, provide failover support or redundancy, or for other purposes. For example, (one or more) devices 610 may belong to various types of devices, such as conventional server-type devices, desktop computer-type devices, and / or mobile devices. Therefore, although illustrated as a single type of device or server-type device, (one or more) devices 610 may include a variety of device types and are not limited to a particular type of device. (One or more) devices 610 may represent, but are not limited to: a server computer, a desktop computer, a web server computer, a personal computer, a mobile computer, a laptop computer, a tablet computer, or any other type of computing device.
[0085] A client computing device (e.g., one of client computing devices 606(1)-606(N)) (each of which is also referred to herein as a "data processing system") can be of various types of devices, which can be the same as or different from device(s) 610, such as a conventional client-type device, a desktop-type device, a mobile-type device, a dedicated-type device, an embedded device, and / or a wearable-type device. Thus, client computing devices can include, but are not limited to, desktop computers, gaming consoles and / or gaming devices, tablet computers, personal data assistants (“PDAs”), mobile phone / tablet hybrid devices, laptop computers, telecommunication devices, computer navigation-type client computing devices such as satellite-based navigation systems including global positioning system (“GPS”) devices, wearable devices, virtual reality (“VR”) devices, augmented reality (“AR”) devices, implanted computing devices, automotive computers, web-enabled televisions, thin clients, terminals, Internet of Things (“IoT”) devices, workstations, media players, personal video recorders (“PVRs”), set-top boxes, cameras, integrated components (e.g., peripherals) for inclusion in computing devices, appliances, or any other type of computing device. Furthermore, client computing devices may include combinations of the previously listed examples of client computing devices, such as, for example, a desktop computer-type device or a mobile-type device in combination with a wearable device, etc.
[0086] Client computing devices 606(1) to 606(N) of various classes and device types can represent any type of computing device having one or more data processing units 692, which are operably connected to a computer-readable medium 694, such as via a bus 616, which in some cases can include one or more of the following: a system bus, a data bus, an address bus, a PCI bus, a Mini-PCI bus, and any various local, peripheral and / or independent buses.
[0087] The executable instructions stored on the computer-readable medium 694 may include, for example, an operating system 619 , a client module 620 , a profile module 622 , and other modules, programs, or applications that can be loaded and executed by the data processing unit(s) 692 .
[0088] The client computing devices 606(1)-606(N) may also include one or more interfaces 624 to enable communication between the client computing devices 606(1)-606(N) and other networked devices (such as device(s) 610) via the network(s) 608. Such network interface(s) 624 may include one or more network interface controllers (NICs) or other types of transceiver devices to send and receive communications and / or data via the network. In addition, the client computing devices 606(1)-606(N) may include input / output ("I / O") interfaces (devices) 626 that enable communication with user input devices, such as peripheral input devices (e.g., game controllers, keyboards, mice, pens, voice input devices such as microphones, video cameras for obtaining and providing video feeds and / or still images, touch input devices, gesture input devices, etc.) and / or output devices, such as peripheral output devices (e.g., displays, printers, audio speakers, tactile output devices, etc.). Figure 6 Client computing device 606 ( 1 ) is illustrated as being connected in some manner to a display device (eg, display screen 629 (N)) that is capable of displaying a UI in accordance with the techniques described herein.
[0089] exist Figure 6In the example environment 600 of FIG. 600 , client computing devices 606 ( 1 ) to 606 (N) can use their respective client modules 620 to connect to each other and / or to (one or more) other external devices to participate in a communication session 603 or to contribute activities to the collaborative environment. For example, a first user can use client computing device 606 ( 1 ) to communicate with a second user of another client computing device 606 ( 2 ). When executing client module 620 , users can share data, which can cause client computing device 606 ( 1 ) to connect to system 602 and / or other client computing devices 606 ( 2 ) to 606 (N) via network (s) 608 .
[0090] The client computing devices 606(1) to 606(N) may use their respective profile modules 622 to generate participant profiles (in Figure 6 602) and provide the participant profile to other client computing devices and / or device(s) 610 of system 602. A participant profile may include one or more of the following: an identity of a user or a group of users (e.g., name, unique identifier (“ID”), etc.), user data such as personal data, user data such as location (e.g., IP address, room in a building, etc.) and technical capabilities, etc. The participant profile may be used to register a participant for a communication session.
[0091] As in Figure 6 , device(s) 610 of system 602 include a server module 630 and an output module 632. In this example, server module 630 is configured to receive media streams 634(1) to 634(N) from individual client computing devices, such as client computing devices 606(1) to 606(N). As described above, the media streams can include video feeds (e.g., audio and visual data associated with a user), audio data to be output with a presentation of an avatar of the user (e.g., an audio-only experience that does not send the user's video data), text data (e.g., a text message), file data, and / or screen sharing data (e.g., a document, a slide deck, an image, a video displayed on a display screen, etc.), etc. Thus, server module 630 is configured to receive a collection of various media streams 634(1) to 634(N) (the collection is referred to herein as "media data 634") during a live view of communication session 603. In some scenarios, not all client computing devices participating in communication session 603 provide media streams. For example, the client computing device may be a consuming device or a “listening” device only, such that it only receives content associated with the communication session 603 , but does not provide any content to the communication session 603 .
[0092] In various examples, the server module 630 can select aspects of the media stream 634 to be shared with individual client computing devices of the participating client computing devices 606(1) to 606(N). Thus, the server module 630 can be configured to generate session data 636 based on the stream 634 and / or pass the session data 636 to the output module 632. The output module 632 can then transmit communication data 639 to the client computing devices (e.g., the client computing devices 606(1) to 606(3) participating in the live view of the communication session). The communication data 639 can include video, audio, and / or other content data provided by the output module 632 based on content 650 associated with the output module 632 and based on the received session data 636. The content 650 can include the stream 634 or other shared data, such as an image file, a spreadsheet file, a slide deck, a document, etc. The stream 634 can include a video component depicting an image captured by the I / O device 626 on each client computer. The content 650 also includes input data from each user, which can be used to control the direction and position of the presentation. The content can also include instructions for sharing data and identifiers of recipients of the shared data. Therefore, the content 650 is also referred to as input data 650 or input 650 in this article.
[0093] As shown, output module 632 transmits communication data 639(1) to client computing device 606(1), and communication data 639(2) to client computing device 606(2), and communication data 639(3) to client computing device 606(3), etc. The communication data 639 sent to the client computing devices can be the same or can be different (e.g., the positioning of the flow of content within a user interface can vary from one device to the next).
[0094] In various embodiments, the device(s) 610 and / or the client module 620 can include a GUI presentation module 640. The GUI presentation module 640 can be configured to analyze the communication data 639 for delivery to one or more of the client computing devices 606. Specifically, at the device(s) 610 and / or the client computing device 606, the UI presentation module 640 can analyze the communication data 639 to determine an appropriate manner for displaying the video, image, and / or content on the display screen 629 of the associated client computing device 606. In some embodiments, the GUI presentation module 640 can provide the video, image, and / or content to a presentation GUI 646 rendered on the display screen 629 of the associated client computing device 606. The presentation GUI 646 can be rendered on the display screen 629 by the GUI presentation module 640. The presentation GUI 646 can include the video, image, and / or content analyzed by the GUI presentation module 640.
[0095] In some implementations, the presentation GUI 646 may include multiple sections or grids that may render or include videos, images, and / or content for display on the display screen 629. For example, a first section of the presentation GUI 646 may include a video feed of a presenter or individual, and a second section of the presentation GUI 646 may include a video feed of a personal consumption conference information provided by the presenter or individual. The GUI presentation module 640 may populate the first section and the second section of the presentation GUI 646 in a manner that appropriately mimics an ambient experience that the presenter and individual may share.
[0096] In some embodiments, the GUI presentation module 640 can zoom in or provide a zoomed view of the individual represented by the video feed to highlight the individual's reactions to the presenter, such as facial features. In some embodiments, the presentation GUI 646 can include video feeds of multiple participants associated with a meeting, such as a general communication session. In other embodiments, the presentation GUI 646 can be associated with a channel, such as a chat channel, an enterprise team channel, etc. Thus, the presentation GUI 646 can be associated with an external communication session that is different from a general communication session.
[0097] Figure 7A diagram showing example components of an example device 700 (also referred to herein as a "computing device") configured to generate data for some user interfaces disclosed herein is illustrated. Device 700 can generate data that can include one or more portions of a video, image, virtual object, and / or content that can be rendered or included for display on display screen 629. Device 700 can represent one of the device(s) described herein. Additionally or alternatively, device 700 can represent one of client computing devices 606.
[0098] As illustrated, device 700 includes one or more data processing units 702, computer readable media 704, and communication interface(s) 706. The components of device 700 are operably connected, for example, via bus 709, which may include one or more of a system bus, a data bus, an address bus, a PCI bus, a Mini-PCI bus, and any of a variety of local, peripheral, and / or independent buses.
[0099] As used herein, data processing unit(s) (such as data processing unit(s) 702 and / or data processing unit(s) 692) may represent, for example, a CPU-type data processing unit, a GPU-type data processing unit, a field programmable gate array (“FPGA”), another type of DSP, or other hardware logic components that may be driven by a CPU in some cases. For example, and not limitation, exemplary types of hardware logic components that may be utilized include application specific integrated circuits (“ASICs”), application specific standard products (“ASSPs”), systems on chips (“SOCs”), complex programmable logic devices (“CPLDs”), and the like.
[0100] As used herein, computer-readable media (such as computer-readable media 704 and computer-readable media 694) can store instructions that can be executed by (multiple) data processing units. Computer-readable media can also store instructions that can be executed by external data processing units, such as by an external CPU, an external GPU, and / or can be executed by an external accelerator, such as an FPGA type accelerator, a DSP type accelerator, or any other internal or external accelerator. In various examples, at least one CPU, GPU, and / or accelerator is incorporated into the computing device, and in some examples, one or more of the CPU, GPU, and / or accelerator is external to the computing device.
[0101] Computer readable media (also referred to herein as computer readable media) may include computer storage media and / or communication media. Computer storage media may include one or more of volatile memory, non-volatile memory and / or other persistent and / or secondary computer storage media, removable and non-removable computer storage media implemented in any method or technology for storing information such as computer readable instructions, data structures, program modules or other data. Thus, computer storage media includes media in tangible and / or physical form included in a device and / or hardware component that is part of the device or a hardware component external to the device, including, but not limited to: random access memory (“RAM”), static random access memory (“SRAM”), dynamic random access memory (“DRAM”), phase change memory (“PCM”), read-only memory (“ROM”), erasable programmable read-only memory (“EPROM”), electrically erasable programmable read-only memory (“EEPROM”), flash memory, compact disk read-only memory (“CD-ROM”), digital versatile disk (“DVD”), optical cards or other optical storage media, magnetic cassettes, magnetic tape, magnetic disk storage, magnetic cards or other magnetic storage devices or media, solid-state memory devices, storage arrays, network attached storage devices, storage area networks, hosted computer storage devices, or any other storage memory, storage devices, and / or storage media that can be used to store and maintain information for access by a computing device. Computer storage media may also be referred to herein as computer-readable storage media, non-transitory computer-readable storage media, non-transitory computer-readable media, computer-readable storage media, computer-readable storage devices, or computer storage media.
[0102] In contrast to computer storage media, communication media may embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism. As defined herein, computer storage media does not include communication media. That is, computer storage media itself does not include communication media consisting only of a modulated data signal, a carrier wave, or a propagated signal.
[0103] The communication interface(s) 706 may represent, for example, a network interface controller ("NIC") or other type of transceiver device for sending and receiving communications over a network. In addition, the communication interface 706 may include one or more cameras and / or audio devices 722 to enable the generation of video feeds and / or still images, etc.
[0104] In the illustrated example, computer-readable medium 704 includes data storage 708. In some examples, data storage 708 includes a data storage, such as a database, a data warehouse, or other type of structured or unstructured data storage. In some examples, data storage repository 708 includes a corpus and / or a relational database having one or more tables, indexes, stored procedures, etc. to enable data access, including, for example, one or more of a Hypertext Markup Language ("HTML") table, a Resource Description Framework ("RDF") table, a Web Ontology Language ("OWL") table, and / or an Extensible Markup Language ("XML") table.
[0105] The data store 708 may store data for the operation of processes, applications, components, and / or modules stored in the computer-readable medium 704 and / or executed by the (one or more) data processing units 702 and / or (one or more) accelerators. For example, in some examples, the data store 708 may store session data 710 (e.g., Figure 6 The data store 708 may also include context data 714, such as content including video, audio, or other content for presentation and display on one or more of the display screens 629. The hardware data 711 may define aspects of any device, such as multiple display screens of a computer. The context data 714 may define any type of activity or status associated with individual users 10A-10L, each associated with each of the multiple video streams 634. For example, the context data may define the level of a person in an organization, how each person's level relates to the level of other people, the performance level of a person, or any other activity or status information that may be used to determine the location of a person's rendering within a virtual environment. This contextual information can also be fed into any model to help emphasize keywords spoken by a person of a certain level, highlight the UI when background voices of a person of a certain level are detected, or change the emotion display in a specific way when a person of a certain level is detected to have a certain emotion.
[0106] Alternatively, some or all of the above data may be stored on a separate memory 716 on one or more data processing units 702, such as a memory on a single board, a CPU-type processor, a GPU-type processor, an FPGA-type accelerator, a DSP-type accelerator, and / or another accelerator. In this example, the computer-readable medium 704 also includes an operating system 718 and an application programming interface 710 (API) configured to disclose functions and data of the device 700 to other devices. In addition, the computer-readable medium 704 includes one or more modules, such as a server module 730, an output module 732, and a GUI presentation module 740, although the number of modules shown is only an example and the number may vary. That is, the functionality described herein in conjunction with the illustrated modules may be performed by a smaller number of modules or a larger number of modules on one device or spread across multiple devices.
[0107] The following claims further serve to define the present disclosure:
[0108] Item A: A computer-implemented method for providing persistent participant prioritization across communication sessions, the method being for execution on a system (100), the method comprising: accessing settings (400) maintained across multiple communication sessions for a user (10J), the settings defining individual prerequisites (410) for the user (10J), the access to the settings being automatically performed by the system without user input, wherein the access to the settings and the selection of at least one selected user (10A) having a role (420A) corresponding to at least one prerequisite of the user (10J) are automatically performed by the system (100) in response to the user (10J) joining the communication session; analyzing a data structure (401) associating individual users (10A, 10K, 10L) with one or more roles (420), wherein the analysis of the data structure Identifying one or more user profiles (430A) of the at least one selected user (10A) having a role (420A) corresponding to the at least one prerequisite (410A) of the user (10J); and causing a rendering of a video stream (151A) of the at least one selected user (10A) to be displayed within a designated area (120A) of a user interface (101A) rendered on a device (11J) associated with the user (10J), wherein the user (10J) has the at least one prerequisite (410A) corresponding to the role (420A) of the at least one selected user (10A), and wherein the display of the rendering of the video stream of the at least one selected user (10A) is not displayed on display devices of other users (10A-10I and 10L) that do not have settings including the prerequisite corresponding to the role of the at least one selected user.
[0109] Item B: A computer-implemented method as described in Item A, wherein the system accesses the settings in response to the user joining the communication session so that the system can automatically display the at least one selected user in the user interface without requiring input from the user, and the permission for the settings is configured to allow the system to access the settings before the user joins the session so that the system can display the at least one selected user in the designated area when the user joins the communication session.
[0110] Clause C: A computer-implemented method according to clauses A to B, wherein, in response to a predetermined state change of the communication session, one or more licenses restrict movement of the rendering of the video stream (151A) of the at least one selected user (10A) within the designated area (120A) of the user interface, wherein the state change includes at least one of the following: detecting that a data rate of at least one computing device of the communication session is below a data rate threshold.
[0111] Clause D: A computer-implemented method as described in clauses A to C, wherein, in response to a predetermined state change of the communication session, one or more license-restricted user interfaces are configured to cause movement of the rendering of the video stream (151A) of the at least one selected user (10A) within the designated area (120A), wherein the state change includes at least one of: detecting shared content provided by at least one participant of the communication session for display on a primary area of one or more computing devices participating in the communication session.
[0112] Clause E: A computer-implemented method as described in clauses A to D, wherein, in response to a predetermined state change of the communication session, the rendering of the video stream (151A) of the at least one selected user (10A) within the designated area (120A) of one or more permission-restricted user interfaces is moved, wherein the state change includes at least one of: detection of shared content provided by at least one participant of the communication session, or modification of the rendering of other participants of the communication session based on a new user joining the communication session or an active speaker reaching a threshold of voice activity.
[0113] Clause F: The computer-implemented method of clauses A to E, further comprising: determining that the role corresponds to the at least one prerequisite of the user by identifying at least one of: a keyword match between the role and the at least one prerequisite, a phrase match between the role and the at least one prerequisite, or a character usage match between the role and the at least one prerequisite by applying historical data to a heuristic-based operation.
[0114] Clause G: A computer-implemented method as described in clauses A to F, wherein, in response to a predetermined state change of the communication session, the rendering of the video stream (151A) of the at least one selected user (10A) within the designated area (120A) of one or more permission-restricted user interfaces is moved, wherein the state change includes at least one of the following: detection of shared content provided by at least one participant of the communication session, or modification of the rendering of other participants of the communication session based on a new user joining the communication session or an active speaker reaching a threshold of voice activity, wherein the rendering of a second group of users displayed in a second designated area is configured to be modified in response to the state change.
[0115] In summary, although various configurations have been described in language specific to structural features and / or method actions, it should be understood that the subject matter defined in the attached representations is not necessarily limited to the specific features or actions described. Instead, the specific features and actions are disclosed as example forms of implementing the claimed subject matter.
Claims
1. A computer-implemented method for providing persistent participant prioritization across a communication session, the method being for execution on a system, the method comprising: accessing settings maintained across multiple communication sessions for a user, the settings defining individual prerequisites for the user, the accessing of the settings being automatically performed by the system without user input, wherein the accessing of the settings and the selection of at least one selected user having a role corresponding to at least one prerequisite for the user are automatically performed by the system in response to the user joining a communication session; analyzing a data structure that associates individual users with one or more roles, wherein the analyzing of the data structure identifies one or more user profiles of the at least one selected user having the role corresponding to the at least one prerequisite of the user; and Causing a rendering of a video stream of at least one selected user to be displayed within a designated area of a user interface rendered on a device associated with the user, wherein the user has the at least one prerequisite corresponding to the role of the at least one selected user, and wherein the display of the rendering of the video stream of the at least one selected user is not displayed on display devices of other users that do not have settings including the prerequisite corresponding to the role of the at least one selected user.
2. The computer-implemented method of claim 1 , wherein: The system accesses the settings in response to the user joining the communication session to enable the system to automatically display the at least one selected user in the user interface without requiring input from the user, and the permission for the settings is configured to allow the system to access the settings before the user joins the session to enable the system to display the at least one selected user in the designated area when the user joins the communication session.
3. The computer-implemented method of claim 1 , wherein: In response to a predetermined state change of the communication session, one or more permissions restrict movement of the rendering of the video stream of the at least one selected user within the designated area of the user interface, wherein the state change includes at least one of the following: detecting that a data rate of at least one computing device of the communication session is below a data rate threshold.
4. The computer-implemented method of claim 1 , wherein: In response to a predetermined state change of the communication session, one or more permissions restrict movement of the rendering of the video stream of the at least one selected user within the designated area of the user interface, wherein the state change includes at least one of the following: detecting shared content provided by at least one participant of the communication session for display on a primary area of one or more computing devices participating in the communication session.
5. The computer-implemented method of claim 1 , wherein: One or more permissions restrict movement of the rendering of the video stream of at least one selected user within the designated area of the user interface in response to a predetermined state change of the communication session, wherein the state change comprises at least one of: detection of shared content provided by at least one participant of the communication session, or modification of the rendering of other participants of the communication session based on a new user joining the communication session or an active speaker reaching a threshold of voice activity.
6. The computer-implemented method of claim 1 , further comprising: The role is determined to correspond to the at least one prerequisite of the user by identifying at least one of: a keyword match between the role and the at least one prerequisite, a phrase match between the role and the at least one prerequisite, or a character usage match between the role and the at least one prerequisite by applying historical data to a heuristic-based operation.
7. The computer-implemented method of claim 1 , wherein: In response to a predetermined state change of the communication session, one or more permissions restrict movement of the rendering of the video stream of the at least one selected user within the designated area of the user interface, wherein the state change includes at least one of the following: detection of shared content provided by at least one participant of the communication session, or modification of the rendering of other participants of the communication session based on a new user joining the communication session or an active speaker reaching a threshold of voice activity, wherein the rendering of a second group of users displayed in a second designated area is configured to be modified in response to the state change.
8. A computing device for providing persistent participant prioritization across a communication session, comprising: one or more processing units; as well as A computer-readable storage medium having computer-executable instructions encoded thereon to cause the one or more processing units to: accessing settings maintained across multiple communication sessions for a user, the settings defining individual prerequisites for the user, the accessing of the settings being automatically performed by the system without user input, wherein the accessing of the settings and the selection of at least one selected user having a role corresponding to at least one prerequisite for the user are automatically performed by the system in response to the user joining a communication session; analyzing a data structure that associates individual users with one or more roles, wherein the analyzing of the data structure identifies one or more user profiles of the at least one selected user having the role corresponding to the at least one prerequisite of the user; and Causing a rendering of a video stream of the at least one selected user to be displayed within a designated area of a user interface rendered on a device associated with the user having the at least one prerequisite corresponding to the role of the at least one selected user, wherein the display of the rendering of the video stream of the at least one selected user is not displayed on display devices of other users who do not have the prerequisite corresponding to the role of the at least one selected user.
9. The computing device of claim 8, wherein: The system accesses the settings in response to the user joining the communication session to enable the system to automatically display the at least one selected user in the user interface without requiring input from the user, and the permission for the settings is configured to allow the system to access the settings before the user joins the session to enable the system to display the at least one selected user in the designated area when the user joins the communication session.
10. The computing device of claim 8, wherein: In response to a predetermined state change of the communication session, one or more permissions restrict movement of the rendering of the video stream of the at least one selected user within the designated area of the user interface, wherein the state change includes at least one of the following: detecting that a data rate of at least one computing device of the communication session is below a data rate threshold.
11. The computing device of claim 8, wherein: In response to a predetermined state change of the communication session, one or more permissions restrict movement of the rendering of the video stream of the at least one selected user within the designated area of the user interface, wherein the state change includes at least one of the following: detecting shared content provided by at least one participant of the communication session for display on a primary area of one or more computing devices participating in the communication session.
12. The computing device of claim 8, wherein: One or more permissions restrict movement of the rendering of the video stream of the at least one selected user within the designated area of the user interface in response to a predetermined state change of the communication session, wherein the state change comprises at least one of: detection of shared content provided by at least one participant of the communication session, or modification of the rendering of other participants of the communication session based on a new user joining the communication session or an active speaker reaching a threshold of voice activity.
13. The computing device of claim 8, wherein: The instructions also cause the one or more processing units to determine that the role corresponds to the at least one prerequisite of the user by identifying at least one of: a keyword match between the role and the at least one prerequisite, a phrase match between the role and the at least one prerequisite, or a character usage match between the role and the at least one prerequisite by applying historical data to a heuristic-based operation.
14. The computing device of claim 8, wherein: In response to a predetermined state change of the communication session, one or more permissions restrict movement of the rendering of the video stream of the at least one selected user within the designated area of the user interface, wherein the state change includes at least one of the following: detection of shared content provided by at least one participant of the communication session, or modification of the rendering of other participants of the communication session based on a new user joining the communication session or an active speaker reaching a threshold of voice activity, wherein the rendering of a second group of users displayed in a second designated area is configured to be modified in response to the state change.
15. A computer-readable storage medium having encoded thereon computer-executable instructions for causing one or more processing units of a system to: accessing settings maintained across multiple communication sessions for a user, the settings defining individual prerequisites for the user, the accessing of the settings being performed automatically by the system without user input, wherein said accessing of said settings and selecting of at least one selected user having a role corresponding to at least one prerequisite of said user are automatically performed by said system in response to said user joining a communication session; analyzing a data structure that associates individual users with one or more roles, wherein the analyzing of the data structure identifies one or more user profiles of the at least one selected user having the role corresponding to the at least one prerequisite of the user; and Causing a rendering of a video stream of the at least one selected user to be displayed within a designated area of a user interface rendered on a device associated with the user having the at least one prerequisite corresponding to the role of the at least one selected user, wherein the display of the rendering of the video stream of the at least one selected user is not displayed on display devices of other users who do not have the prerequisite corresponding to the role of the at least one selected user.