Persistent display of shared content for priority participants with communication sessions

By dynamically adjusting the display position and size of the assistant video stream, the problem of the assistant video stream being blocked by shared content in the existing collaborative system is solved, and user engagement and productivity are improved.

CN120226312APending Publication Date: 2025-06-27MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380078929.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-14
Filing Date
2023-09-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the case where existing collaborative systems require language translation or sign language translation, the assistant's video stream is easily blocked or adjusted by the display of shared content, resulting in a decrease in participant's productivity and participation.

Method used

The system dynamically adjusts the assistant's video stream display position and size by automatically detecting the type of shared content and user settings to ensure that it remains continuously visible and unblocked during the shared content display.

Benefits of technology

The priority display of assistant video streams is realized, which improves user engagement and productivity, and reduces the inefficient utilization of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226312A_ABST
    Figure CN120226312A_ABST
Patent Text Reader

Abstract

A system provides preferential participants and persistent display of shared content. In a virtual conference, the system may automatically and adaptively arrange video locations for particular participants based on the participants' roles and shared content having particular data types. For example, a user may have assistance settings indicating that help is needed. The system can automatically display a persistent display of the video stream of the user's assistant in a user interface. When content is shared, the system analyzes a rendering of the content to identify areas where the shared content is not displayed. Subsequently, the system dynamically configures the user interface such that a persistent display of the video stream of the user assistant is located in an area where the shared content is not displayed. These features enable a user to have a continuously visible view of the assistant during display of the shared content.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] There are various collaboration systems that allow users to communicate. For example, some systems allow people to collaborate by sharing content such as video streams, shared files, chat messages, etc. Some systems also allow people to edit documents simultaneously while enabling them to communicate using video and audio streams. Users can also establish communication sessions at specific times and share real-time video streams that can display people and content simultaneously.

[0002] Although existing collaboration systems provide a feature set that allows people to conduct meetings via live video streams, some of these systems still have many drawbacks. For example, some existing systems do not have effective features to accommodate people who require language translation or sign language interpretation. In such cases, a meeting participant can have an assistant join the meeting, such as a translator or interpreter. Then, the assistant can listen to the meeting, observe the shared content and video stream, and provide an interpretation of their observations. For these tasks, it is important for the meeting participants to have a clear view of their assistant. If the video rendering of the assistant's live video stream moves or resizes during the meeting, the participants may have difficulty following the flow of the meeting. If this occurs, significant information may be missed.

[0003] Some existing systems provide some features that can restrict operations for repositioning video streams. For example, some current solutions allow meeting participants to select a video stream, and for example, the video stream can be "pinned" to a location. Although this solution can help in some cases, there are many situations where these selected streams can be resized, moved, or removed entirely. In an illustrative example, when content is shared during a meeting, the selected stream depicting a participant can be moved or resized. When a slide file is shared during an online meeting, such content is typically displayed on the main stage of the user interface. Even if the video stream is pinned, this arrangement usually reduces the rendering of other users to a small size or removes it entirely. This type of rearrangement, especially the rearrangement of the video stream of a person's assistant, can lead to a decrease in productivity and engagement of people who rely on the assistant, especially in cases where sign language interpretation or language translation is required. These problems and others can lead to a decrease in productivity and engagement, which ultimately results in inefficient use of computing resources. Summary of the Invention

[0004] The techniques disclosed herein enable a system to provide a persistent display of a priority participant and shared content. In a virtual meeting, the system can automatically and adaptively position the video of a particular attendee based on the shared content type and user settings, and based on the role of the attendee (such as "sign language interpreter"). The system can dynamically move the display of the video of a particular attendee and adjust its size to mitigate situations where the rendering of the video of a particular attendee is blocked or visually impaired by the display of the shared content. For example, a first user can have a user setting indicating a user in need of assistance, such as "hearing impairment" in the accessibility settings. When this first user is in a meeting, the system can automatically display a persistent display of the video stream of the user's assistant in a designated area of the user interface. When shared content is presented during the meeting, the system determines whether the content type meets one or more criteria. If the content type meets one or more criteria, the system analyzes the user interface to identify a first set of areas that display the shared content and a second set of areas that do not display the shared content. The system then dynamically configures the user interface such that the persistent display of the video stream of the user's assistant is located in the second set of areas that do not display the shared content. Such a feature enables the user to have a consistent view of the assistant during the display of the shared content.

[0005] In some configurations, detecting shared content having a particular data type can cause the system to transition from a normal operation mode to a content tracking mode. In the normal operation mode, the system can automatically display the video stream of an assistant having a role corresponding to the prerequisites of the meeting participants. When the system is in the normal operation mode, the video stream of the assistant can be in a static position, e.g., on the main stage of the user interface. When a user shares content with other users, the system induces the content tracking mode, in which the system continuously analyzes the rendering of the shared content that meets one or more criteria and determines the areas of the display content within the rendering and other areas that do not display content. The system then displays the video stream of the assistant in the areas of the user interface that do not display content. This helps to mitigate the overlap between the video stream of the assistant and the display of the shared content while also keeping the video stream of the assistant at a size sufficient to allow the user to understand the gestures of their assistant.

[0006] When a user shares content with a specific data type and / or when the shared content is displayed in a specific area of the user interface, the shared content may meet the criteria for triggering a content tracking mode. For example, when a user shares content with a specific data type (e.g., video, text document, or slide file), detecting the shared content can trigger the content tracking mode. In such an example, if the user shares other types of content that do not meet one or more criteria (e.g., chat messages, contact cards, or still images displayed within a chat thread), the system may not trigger the content tracking mode. In another example, when content is shared in a predetermined area (e.g., the main stage of the user interface), detecting the shared content can trigger the content tracking mode. If content is shared in other areas of the user interface (e.g., within a chat thread, sub-stage, etc.), such an implementation may not trigger the content tracking mode.

[0007] The criteria for triggering a content tracking mode may also include a combination of detected events. For example, when a user shares content with a specific data type (e.g., video, text document, or slide) that is displayed within the main stage of the user interface, detecting the shared content can trigger the content tracking mode. In such an implementation, other types of data (e.g., chat messages, contact cards, or still images shared within a chat thread or sub-stage) will not trigger the content tracking mode.

[0008] The techniques disclosed herein provide many technical benefits. In one example, the techniques disclosed herein provide reliable assistive features. If the participants in a meeting require sign language interpretation, the system can maintain the display of their sign language interpreter during multiple interruptions. This has many benefits over traditional front-end pinning. For example, certain events (e.g., detecting shared content) do not disrupt the display of the language translation video stream. This allows users to view the translation of the meeting content with higher reliability than some existing systems. Additionally, users do not have to go through the process of selecting which language translation to pin to the front-end display during the meeting. The automatic selection and persistent display of the assistant eliminates the need for meeting participants to manually identify another user as an assistant and provide input to pin the display of that other user to the front-end. This can save a lot of computing resources because meeting participants do not interrupt the meeting or miss any content every time they join the meeting.

[0009] By providing participant prioritization across communication sessions and for providing a persistent display of prioritized participants during the display of content, the system can enhance user engagement. By enhancing user engagement and avoiding user fatigue, especially in a communication system, users can exchange information more effectively. This helps mitigate the situation of missing or overlooking shared content when users are distracted or disengaged. Enhancing user engagement and avoiding user fatigue can reduce the occurrence of users needing to extend meetings or resend missed information. More effective communication of shared content can also help avoid the need for external systems such as mobile phones for texting and other messaging platforms. This can help reduce the reuse of network, processor, memory, or other computing resources. The disclosed techniques also use automation of user settings to provide improved human interaction with the system. This enables the system to be utilized in a more effective manner by reducing the display of unwanted menus, reducing misselected objects, or reducing erroneously triggered operations.

[0010] By reading the following detailed description and reading the associated drawings, features and technical benefits other than those explicitly described above will be apparent. The present invention content is provided to introduce a selection of concepts that are further described below in the detailed description in a simplified form. The present invention content is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. For example, the term "technique" may refer to a system, method, computer-readable instructions, module, algorithm, hardware logic, and / or operation as permitted by the context described above and the entire document. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The detailed description is described with reference to the accompanying drawings. In the drawings, the leftmost digit of a reference numeral identifies the drawing in which the reference numeral first appears. The same reference numerals in different drawings indicate similar or identical items. References to individual items among multiple items may use reference numerals with letters in an alphabetical sequence to refer to each individual item. General references to an item may use a specific reference numeral without an alphabetical sequence.

[0012] Figure 1A A persistent display of a video of an assistant automatically configured in a specified area of a user interface in response to determining that the role of the assistant corresponds to a prerequisite of a user associated with the user interface is shown.

[0013] Figure 1B A user interface for sharing content among users in a communication session is shown.

[0014] Figure 1C An example of an area of a user interface is shown, where a first set of identified areas is determined to include shared content, and a second set of identified areas is determined not to include shared content.

[0015] Figure 1D Shows a video of an assistant that is persistently displayed in one of the areas identified in the second set that is determined not to include shared content.

[0016] Figure 1E Shows a video of an assistant that is moved when the shared content is updated.

[0017] Figure 2 Shows an example of how the video of the assistant is modified to accommodate large-scale or full-screen display of the shared content.

[0018] Figure 3 Shows an example of how the input device is used to move the video of the assistant during content sharing among users.

[0019] Figure 4A Shows a user interface layout where the communication application is minimized and the communication application displays a monitor with shared content on the screen sharing of the desktop.

[0020] Figure 4B Shows a user interface layout where the communication application is minimized and the communication application displays a monitor with shared content on the screen sharing of the desktop, where the monitor is moved by the input of the presenter interacting with the desktop of the presenter's computer.

[0021] Figure 4C Shows a user interface layout where the communication application is minimized and the communication application displays a monitor with shared content on a word processing application.

[0022] Figure 4D Shows a user interface layout where the communication application is minimized and the communication application displays a monitor with shared content on a word processing application, where the monitor is moved by the input of the presenter interacting with the word processing application running on the presenter's computer.

[0023] Figure 5A Shows an example of an area that can be used to analyze the shared content to determine the area for displaying the video of the assistant.

[0024] Figure 5B Shows how the rendering of the shared content can be analyzed to determine the area for displaying the video of the assistant.

[0025] Figure 5C Shows how the selected area can be used to determine the area for displaying the video of the assistant.

[0026] Figure 6 Shows an example of user settings and mapping data structures.

[0027] Figure 7is a flowchart showing aspects of a routine for utilizing the disclosed techniques.

[0028] Figure 8 is a computer architecture diagram showing an illustrative computer hardware and software architecture of a computing system capable of implementing the techniques and aspects of the techniques presented herein.

[0029] Figure 9 is a computer architecture diagram showing a computing device architecture of a computing device capable of implementing the techniques and aspects of the techniques presented herein. DETAILED DESCRIPTION

[0030] Figures 1A to 1E shows aspects of a system 100 that provides a persistent display of shared content and a priority participant during a communication session. The system automatically displays a selected participant, such as a sign language interpreter. The system can also dynamically move and resize the display of the selected participant to mitigate instances where the rendering of the selected participant is blocked by the display of the shared content or is visually impaired. As Figure 1A shown, the system can start in a normal operation mode, where a first user interface includes a rendering of a video stream of a selected participant 10A (such as a sign language interpreter) displayed in a main area. In the normal operation mode, the display of the selected participant does not move based on the display of the shared content. When content is shared by another user and displayed in the main area of the user interface, as Figure 1B shown, the system induces a content tracking mode. Figure 1C shows an example of how shared content can be analyzed during the content tracking mode to identify regions that include the shared content and other regions that do not include the shared content. Figure 1D shows how the system displays the video of the selected participant within a region that does not include the shared content during the content tracking mode, and Figure 1E shows how the system moves the video of the selected participant when the content is updated.

[0031] The communication session can be in the form of an online meeting, a broadcast, or any other gathering that includes a start time and an end time. As Figure 1AAs shown, a communication session can be managed by a system 100 that includes multiple computers 11, each computer 11 corresponding to a single user 10. For illustrative purposes, the first user 10A, Mike Taylor, is associated with the first computer 11A, the second user 10B, Traci Isaac, is associated with the second computer 11B, the third user 10C, Doug Wright, is associated with the third computer 11C, the fourth user 10D, MJ Price, is associated with the fourth computer 11D, the fifth user 10E, Kat Martin, is associated with the fifth computer 11E, the sixth user 10F, Miguel Jones, is associated with the sixth computer 11F, the seventh user 10G, Krystal McKinney, is associated with the seventh computer 11G, the eighth user 10H, Jesica Kline, is associated with the eighth computer 11H, the ninth user 10I, Monica Larsson, is associated with the ninth computer 11I, the tenth user 10J, Charlotte Davis, is associated with the tenth computer 11J, the eleventh user 10K, Anika Andersson, is associated with the eleventh computer 11K, and the twelfth user 10L, Isla Scogins, is associated with the twelfth computer 11L. These users can also be referred to as "User A", "User B", etc. respectively.

[0032] Each user can be displayed as a two-dimensional 2D image in the user interface, or each user can be displayed as a three-dimensional representation, such as an avatar. The 3D representation can be a static model or a dynamic model that is animated in real time in response to user input. Although this example shows a user interface in which users are displayed as 2D images, it is understood that the techniques disclosed herein can be applied to other forms of representation, video, or other types of rendering. The computer 11 can be in the form of a desktop computer, a head-mounted display unit, a tablet computer, a mobile phone, etc. The system can generate a user interface that shows aspects of the communication session to each user. In Figure 1A the example of, the first user interface arrangement 101A can include multiple renderings of one or more users 10. The renderings can include the rendering of two-dimensional (2D) images, which can include pictures of people or live video feeds.

[0033] In this example, the user interface is rendered on the display device of the tenth computer 11J, which is associated with the tenth user 10J Charlotte Davis. Charlotte is referred to herein as the "viewer" or "viewing user" of the user interface displayed on the tenth computer 11J. The first user interface layout 101A includes a first region 131A, also referred to herein as the designated region 131A or the main stage 131A. The first user interface layout 101A also includes a second region 131B, also referred to herein as the secondary region 131B or the secondary stage 131B. The first user interface layout l0lA also includes another rendering of the video stream 151J showing the self-view of the tenth user 10J. This video stream 151J can be displayed in the second region 131B and is restricted to be displayed in the first region 131A. The first region is reserved only for video streams of users with roles corresponding to the prerequisites of the viewer. The first region is also reserved for displaying content shared by users with a predetermined data type, such as recorded videos, slide content, text documents, spreadsheet content, etc.

[0034] When the tenth user 10J (User J) joins the communication session, the system automatically accesses User J's preferences. The preferences can indicate that User J needs assistance, for example, User J has indicated that they have a hearing difficulty and prefer the prerequisite of participating in the meeting with an assistant. In response to this indication, the system can cause a rendering of the video stream 151A of the selected user (e.g., the first user 10A (User A)) to be displayed within the designated region 131A of the user interface 101A. As described in more detail below with respect to FIG. 5, in response to determining that User A has a role (e.g., sign language interpreter) corresponding to the prerequisite of User J (e.g., the indication that User J has a hearing difficulty), User A can be selected to be displayed in the main stage 131A.

[0035] The system is configured to persistently display the assistant in the designated region even when the communication session transitions through different types of operating modes (such as a transition between a normal operating mode where the system only shares the video feeds of individual users and a second operating mode where the system shares the rendering of shared content with a specific content type and the video feeds of individual users). As Figure 1A shown, the system can be in the normal operating mode, where the video rendering 151A of the assistant is located within the main stage level 131A of the first user interface layout 101A. The normal operating mode can be caused when users do not share content to be displayed on the computing devices of other participants.

[0036] Then, the system can receive an input to cause a transition from the normal operating mode to the content tracking mode. This input can be a user input identifying the shared content or system operation, where the content is automatically displayed on one or more client devices 11 of the participants 10 in the communication session 603.Figure 1B An example rendering of content shared among participants 10 in a communication session is shown.

[0037] In response to an input that causes a transition from a normal operation mode to a content tracking mode, the system can analyze the rendering of shared content 120 to identify a first set of regions 141 of the shared content 120 that includes a threshold level and a second set of regions 142 of the shared content 120 that does not include the threshold level. For illustrative purposes, the first set of regions 141 of the shared content 120 that includes the threshold level is also referred to herein as "unavailable regions 141", while the second set of regions 142 of the shared content 120 that does not include the threshold level is also referred to herein as "available regions 142". As described in more detail below, a variety of suitable techniques can be used to perform the analysis of the shared content to identify the regions.

[0038] In some configurations, if the shared content has a specific content type, the system can cause only the content tracking mode. For example, if the input identifies shared content having a data type from a first class of data types (e.g., slides, word processing documents, spreadsheets, videos, or images), the system can transition only to the content tracking mode. Thus, if other types of content, chat messages, audio streams, etc. are shared, the system can not transition to the content tracking mode.

[0039] After the system identifies the available regions and the unavailable regions, the system can place the video rendering of the assistant 10A in one or more of the available regions 142A. As Figure 1D shown, when the system is in the content tracking mode, the system can cause a second user interface layout 101B to be displayed. The second user interface layout 101B can include a rendering 151A of the video stream of a selected participant 10A located within one or more of the available regions (such as region 142A). The rendering 151A of the video stream of the selected participant 10A is located within the available region 142A such that the rendering 151A of the selected participant is displayed simultaneously with the shared content 120 located within the first set of regions 141. This arrangement enables user J to have a continuously visible view of the assistant during the display of the shared content. This arrangement helps to reduce the overlap between the video stream of the assistant and the display of the shared content, while also keeping the video stream of the assistant at a size sufficient to allow the user to understand the assistant's gestures.

[0040] Whenever the shared content is updated, the analysis of the shared content can be performed. For example, if the shared content involves a video, the analysis can be performed for each frame or for the objects depicted in each video change. If the shared content involves a new section of a slide or a document, the analysis of the displayed content can be performed each time an input is received to select the new slide or the new section of the document. For example, when the content transitions from Figure 1B the first slide toFigure 1E When on the second slide, the system analyzes the shared content of the second slide to identify a new available area 142E that does not include content at the threshold level, and the system moves the video rendering 151A from the first position 181A to a second position 181B within the new available area 142E.

[0041] In some configurations, when a predetermined set of conditions is detected, the size of the rendering 151A of the assistant 10A can be reduced. In an illustrative example, if the system determines that the rendering of the shared content does not contain any available area of a predetermined size to accommodate the rendering 151A of the assistant 10A, the system can reduce the rendering 151A of the assistant 10A. As Figure 2 shown, the rendering 151A can be reduced to a minimum size threshold. The minimum size threshold is used to allow the rendering to be in a visible size such that a viewing user can easily view the gestures performed by the assistant. When the rendering size is reduced, the position of the rendering can be placed near the boundary of the rendering of the shared content to minimize the occurrence of the rendering 151A covering significant content near the center of the rendering of the shared content. This position can also include an analysis of the case where a keyword is detected in the rendering of the shared content. If one or more keywords from a predetermined keyword list are detected, the rendering 151A of the assistant is placed in an area that does not include the display of one of the predetermined keywords or any other priority content. Priority content such as a predetermined keyword, an image of a predetermined person or object can be specified by any one of the participants via a preference file, a chat message or any other form of communication generated by the user. Thus, the system can position the rendering 151A of the assistant so as not to overlap with a first area containing priority content, and position the rendering 151A of the assistant to overlap with a second area that does not contain priority content.

[0042] In one example, during the analysis of the rendering of the shared content, the system can determine that one or more dimensions of the areas in the set of available areas 142 are less than a threshold dimension. In response to determining that one or more dimensions of the areas in the set of available areas 142 are less than the threshold dimension, when the system is in the content tracking mode, the system can reduce at least one dimension of the rendering 151A of the video stream of the selected participant 10A to the threshold minimum size.

[0043] In some embodiments, the threshold minimum size may be based on the device type or screen size. For example, for a desktop computer or a device with a twenty-two-inch monitor, the threshold minimum size may be a predetermined percentage of the screen, such as 50% of one dimension of the screen. However, for a tablet computer or a mobile device or a device with a five-inch screen, the threshold minimum size may be a greater predetermined percentage of the screen, such as 90% of one dimension of the screen. Although screen dimensions are used in this example, other units of measurement may be utilized. For example, if the device screen has less than a threshold number of pixels, such as two million, the system may use a first threshold minimum size; and if the device screen has more than the threshold number of pixels, the system may use a second threshold minimum size to render the assistant 10A, where in this example, the first threshold minimum size is greater than the second threshold minimum size. The threshold minimum size may be applied to a plurality of pixels or one or more dimensions of the rendering. This allows the system to use a higher percentage of the screen for rendering the assistant 10A for smaller screen devices. The threshold minimum size may be applied to any rendering 151A of the assistant 10A, or a user having a role corresponding to the prerequisites of the viewing user.

[0044] In some configurations, as Figure 3 shown, the system may enable a user to use an input device to move the rendering 151A of the assistant. This input control may be enabled during a content tracking mode (e.g., when shared content is being displayed). This may include receiving a control input at the system from a computing device 11J communicating with a display screen presenting a second user interface arrangement 101B. The control input includes coordinates within the second user interface arrangement 101B and may be received from a pointing device such as a mouse, trackball, touchpad, touchscreen, eye tracking device, etc.

[0045] The system may then move the rendering 151A of the video stream of the selected participant 10A according to the coordinates indicated by the control input. The control input is configured to control the movement of the rendering 151A of the video stream of the selected participant 10A during a content tracking mode. The control input is restricted from controlling the movement of the rendering 151A of the video stream of the selected participant 10A during a normal operation mode.

[0046] The rendering 151A of the assistant may persist during a transition of the system operation state between multiple levels of the operating system. For illustrative purposes, the rendering 151A of the assistant may also be referred to herein as "monitor 151A". The monitor 151A may be in such as Figures 1A to 3During the use of the communication application shown. When the communication application is minimized, the monitor can also be displayed. The system can be configured to persistently display monitor 151A in the following cases: when the system displays the user interface for the communication application, when the system displays the desktop of the operating system without the user interface for the communication program, and when the system displays the user interface of another application (such as a word processing application or a slide board application) without the user interface for the communication application. Figures 4A - 4C Examples of these operating states are shown.

[0047] Figure 4A Another user interface layout 101C is shown, where the communication application is minimized. For example, the communication application is not displayed and the computer displays another application or the operating system desktop. As Figure 4A shown, when user B is sharing a view of their computer desktop, the computers of other users include the computer of user J. However, what is unique for user J and other users with the prerequisite role corresponding to user A is that in addition to the display of user B's computer desktop, the system also displays the monitor of user 10A on the rendering of the desktop. The size and position of monitor 151A are determined by the techniques for rendering 151A of the assistant disclosed herein. The monitor can be displayed in an area that does not contain threshold-level content or priority content.

[0048] User J can provide input from the corresponding computer 11J to move monitor 151A. In such an input, since user J is not the person sharing the content, the input only moves the monitor for that one computer of user J. For example, the input at computer 11J only moves the monitor displayed at computer 11J, even if other users have a display of monitor 151A showing the same assistant.

[0049] However, when the user sharing the content (such as user B sharing their desktop) provides an input to move the cursor, the input can move the monitor for all other users who are viewing the desktop with a display of monitor 151A showing the assistant. Thus, the input provided by user B to move the cursor or edit the content can move monitor 151A displayed on the screen of user J's computer J. Thereby, when the presenter (user B) is selecting or editing content or selecting a different application, the viewing user (user J) who has already had the assistant displayed does not need to provide an input to make their monitor 151A track the activities of the presenter. The monitor tracks important content shared by the presenter based on the presenter's actions. This feature is shown in Figure 4B where the presenter (user B) moves the cursor, and monitor 151A displayed on the screen of user J tracks the cursor movement.

[0050] Figure 4CShows another example of the monitor 151A displayed in the user interface layout 101C, where the communication application interface is minimized. In this case, the presenter user B has selected a text document displayed in a word processing application to be shown on the screens of other users. In this example, the cursor 471 is in the form of a text input cursor. Thus, as Figure 4C and Figure 4D shown, when the presenter user B types content 120 in the word processing application, the system automatically moves the monitor 151A that is persistently displayed on the computing device of user J to track the movement of the text input cursor. Based on the content being added or the content highlighted by the cursor, this tracking feature allows user J to view relevant or up-to-date content. Thus, in any operating mode, the monitor showing the first user user A that is persistently displayed to the second user user J along with the shared content is automatically moved by the system based on the input provided by the third user user B who shares the screen of the display operating system desktop, where the second user user J has prerequisites corresponding to the role of the first user.

[0051] This feature can be beneficial for other types of applications that show content in a specific location. For example, consider a scenario where a user is sharing their screen with a developer studio and other users in the meeting are viewing the screen. Using the techniques disclosed herein, the assistant's monitor can track the cursor of the user sharing the screen. This allows those viewing the assistant to see details such as code changes while also having a view of their assistant near those details.

[0052] Analyzing the rendering to identify available and unavailable regions can include any suitable technique. Figures 5A to 5C An illustrative example is shown in Figure 5A As shown, the system can generate data for a pattern 410 that defines multiple undefined regions 140 (e.g., a grid of undefined regions). In this example, each region in the region grid includes an undefined region 140, e.g., a region that is not defined as an unavailable region 141 or an available region 142.

[0053] Each undefined region can be of a predetermined shape and size, or each region can be based on the shape and size of the rendering object, such as the characters and pictures of a content file. The regions of the pattern 410 can be as small as a single pixel, which enables per-pixel analysis, or each undefined region can include a larger predetermined area, e.g., a part of the main stage or a part of the secondary stage, etc.

[0054] Then, the system can analyze the undefined regions 140 to determine whether each undefined region contains a threshold level of content. In an illustrative example, the system can analyze each region to determine the multiple pixels used to display content and the multiple pixels not used to display content, e.g., the pixels used to display the background. A contrast threshold can be used to determine the differences in color, brightness, or other display attributes to distinguish the pixels used to display content from the pixels used to display the background. If the ratio of the number of pixels used to display content to the number of pixels not used to display content exceeds a threshold ratio within a specific region, the system determines that the specific region contains a threshold level of content.

[0055] Figure 5B An example of analyzing three regions is shown. For example, the unavailable region 141X is selected. Assuming an example threshold ratio of 0.10, this region includes sixteen pixels used to display content and eighty-four pixels not used to display content, or 16 out of 100 pixels used to display content (ratio of.16). Thus, assuming 16 pixels exceed the threshold level of content, this region is the unavailable region 141X. In the same example threshold ratio of 0.10, the system will also identify the first available region 142X and the second available region 142Y because these regions have two pixels (ratio of 0.02) and zero pixels (ratio of 0.00) used to display content, respectively. Thus, these regions do not have a threshold level of content. Other undefined regions of the pattern 410 of the region can be analyzed in a similar manner, and the results are shown in Figure 5C which Figure 5A each undefined region 140 is designated as an unavailable region 141 or an available region 142, as Figure 5C shown. Although these examples use the same threshold to determine the threshold level of content for each region, it can be understood that different threshold levels can be used. For example, a first threshold can be used to check a first set of regions, and a second threshold can be used to check a second set of regions, where the first threshold and the second threshold are different.

[0056] In some embodiments, the system can display a rendering 151A of the assistant 10A in one of the available regions 142. However, as Figure 5C shown, if the size of any of the available regions 142 is not appropriately adjusted to accommodate the rendering 151A of the assistant 10A associated with the threshold minimum size of the rendering 151A, the system can cluster the available regions 142 to define a larger available region 142A. The rendering 151A of the assistant 10A can be placed in this larger available region 142A or any other region having dimensions corresponding to the threshold minimum size of the rendering 151A.

[0057] In some embodiments, the system may maintain user settings that persist across meetings. This persistence may be achieved by storing user settings that associate preconditions with one or more user identities of individual users. Figure 6 An example of user settings 400 and a mapping data structure 401 is shown. Generally, user settings may define preconditions for certain users. For example, user J has a precondition indicating that they require a hearing aid, user B has a precondition indicating that they require a French translation, and user C has a first precondition indicating that they require a hearing aid, e.g., listing that they have hearing difficulties. Each user may have multiple assistants. For example, some users may indicate that they require two sign language interpreters.

[0058] Each user may also be associated with a content type 411. For example, user J may have a preference indicating an association with data types such as PPTX, XLSX, DOC, and MPEG. Thus, when a user shares content having these data types, the system may trigger a content tracking mode based on these settings. Content type 411 may also be referred to herein as a "data type", a "content data format", or a "content data type". Content type 411 refers to different data packets for displaying content to a display screen and / or the format for such display. For example, an instant message passed in a chat thread is a content data format or a content data type, and a power point file for displaying slides on a presentation window is another content data format or another content data type.

[0059] Other users have other content types that can be used to trigger a content tracking mode. A normal operation mode that may involve an automatic selection of an assistant may be triggered in response to determining that an assistant (e.g., user A) has a role corresponding to a precondition of another user participating in a meeting (e.g., connecting to a communication session).

[0060] The mapping data structure 401 may define multiple profiles 430. Each profile may identify the role of each user. For example, user K may be automatically selected as an assistant for users having a precondition indicating a need for a Spanish translation, a mediator, or an administrator; user L may be automatically selected as an assistant for users having a precondition indicating a need for a French translation or a Spanish translation; and user A may be automatically selected as an assistant for users having a precondition indicating a need for a sign language interpreter.

[0061] The setting 400 is stored in a manner that allows the system 100 to access the setting each time a user joins a meeting. When a meeting participant such as user J joins a meeting, the system accesses the user settings and determines whether one of the prerequisites 410A associated with the meeting participant user 10J corresponds to the role 410A of another user such as user A. When the role of a particular user is determined to correspond to the prerequisite of the meeting participant, the system selects that particular user as the assistant for the meeting participant. The system then persistently displays the rendering of the assistant in the user interface as described herein.

[0062] The setting 400 can also be referred to herein as an "assistive setting". An assistive setting can be any data structure, document, or other form of data that defines a person's needs and associates those needs with their identity. For example, operating system-level or application-level profile or enrollment data can indicate that a user has drivers and equipment for specific assistive needs. This data can be used to indicate prerequisites such as a person having a hearing difficulty. In another example, email or communication data indicating a person's assistive needs can also be used to indicate the user's prerequisites and trigger the operations disclosed herein. If a user has a particular application installed on their phone, such as a sign language application, such data can also be used for the user's prerequisites and trigger the operations disclosed herein. The setting 400 can be at any stack or level, such as the OS, user profile, application level, etc.

[0063] In some embodiments, the system can limit the number of assistants for a particular user. For example, the system can limit user J to only two assistants. This limitation allows the system to provide an understandable display for each assistant, as a large number of assistants may result in smaller renderings that may be difficult to see.

[0064] Figure 7 FIG. is a diagram showing aspects of a routine 500 for providing persistent participant prioritization across communication sessions. Those of ordinary skill in the art will understand that the operations of the methods disclosed herein are not necessarily presented in any particular order, and that it is possible and contemplated to perform some or all of the operations in an alternative order. For ease of description and illustration, the operations have been presented in the order shown. Operations can be added, omitted, performed together, and / or performed simultaneously without departing from the scope of the appended claims.

[0065] It should also be understood that the methods shown can be started or ended at any time and need not be executed in their entirety. Some or all of the operations of the method and / or substantially equivalent operations can be performed by executing computer-readable instructions included on a computer storage medium, as defined herein. The term "computer-readable instructions" and its variants, as used in the specification and claims, are used herein broadly to include routines, applications, application modules, program modules, programs, components, data structures, algorithms, etc. Computer-readable instructions can be implemented on a variety of system configurations, including single-processor or multi-processor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based programmable consumer electronics, combinations thereof, and the like. Although the example routines described below run on a system (e.g., one or more computing devices), it can be understood that the routines can be executed on any computing system that can include any number of computers working in concert to perform the operations disclosed herein.

[0066] Accordingly, it should be understood that the logical operations described herein are implemented as a series of computer-implemented acts or program modules running on a computing system such as that described herein, and / or are implemented as interconnected machine logic circuits or circuit modules within the computing system. The implementation is a matter of choice depending on the performance and other requirements of the computing system. Thus, the logical operations can be implemented in software, firmware, special purpose digital logic, and any combination thereof.

[0067] In addition, the operations shown in FIG. 5 and other figures can be implemented in association with the example user interfaces and systems described herein. For example, the various devices and / or modules described herein can generate, send, receive, and / or display data associated with the content of a communication session (e.g., live content, broadcast events, recorded content, etc.) and / or include a presentation UI rendering of one or more participants of a remote computing device, avatar, channel, chat session, video stream, image, virtual object, and / or application associated with the communication session.

[0068] In operation 502, the system may receive an input that causes a state change from the normal operation mode to the content tracking mode. The input may identify shared content 120 having a data type. The shared content may include files such as slides, text files, videos, etc. The input that causes the state change may be an input from viewing user User J, or the input that causes the state change may be an input from a presenter such as User B for the shared content, or the system may identify the content type of the shared content and determine whether it meets a criterion, and then automatically cause the state change. For example, User J may provide an input such as a voice input, a touch command, a menu selection, or any other type of input to cause a state change from the normal operation mode to the content tracking mode. In another example, a presenter, such as User B who wants to share a screen or shared content on the main stage, may provide an input that causes a state change from the normal operation mode to the content tracking mode by sharing only the screen or by sharing a file or a view of the file.

[0069] This input may be received while the system is operating in the normal operation mode. During the normal operation mode, the system displays a first user interface layout 101A including a rendering 151A of the video stream of selected participant 10A. When the system 100 is in the normal operation mode, the rendering 151A of the video stream of selected participant 10A is located at a predetermined position. The rendering 151A of the assistant includes a video image of a user (e.g., user 10A) having a role determined to correspond to a prerequisite for a participant (e.g., user J) in need of assistance. During the normal operation mode, the display of the rendering 151A of the video stream of selected participant 10A is not moved or positioned relative to any displayed content.

[0070] During operation 502, when the content tracking mode is activated, the system may access a data structure 400 that defines preferences for controlling the position of the rendering of the video stream of selected participant 10A. The preferences may identify one or more content data types 411 for enabling the activation of the content tracking mode. When the shared content meets one or more criteria regarding the one or more data types 411, the activation of the content tracking mode causes the system to control the position of the rendering of the video stream of selected participant 10A relative to the display of the shared content. The one or more data types 411 may be specified in the profile of a conference participant or in a system preference file.

[0071] In response to the input, the routine proceeds to operation 504, where the system determines whether the content identified in the input meets one or more criteria. One or more criteria can be defined in a user profile or other system preferences. For example, as shown in the user settings of FIG. 5, if the input provided by User B includes PPTX and one or more preferences of the participants (including the preferences of User J) include the PPTX data type, the routine can proceed to operation 506. However, if the input provided by User B includes a chat message and the chat message does not match the data type indicated in the user profile or other system preferences, the routine will return to operation 502, where the system will wait for another input.

[0072] In operation 506, the system can perform an operation and transition to a content tracking mode. When in the content tracking mode, the system changes the license of each client computer from one operation mode to another operation mode, in the one operation mode, the rendering of the assistant is locked in place, in the other operation mode, the rendering of the assistant is allowed to move to different positions within the user interface relative to the display of the shared content.

[0073] In operation 508, in response to an input that causes a state change from the normal operation mode to the content tracking mode, the system can analyze the rendering of the shared content 120 to identify a first set of regions 141 that includes the shared content 120 and a second set of regions that do not include the shared content 120. Before identifying and selecting the display regions, the rendering does not have to be displayed on the device. The rendering can start in memory for analysis. Various methods for identifying the location of the displayed content can be used to determine the selected regions.

[0074] In operation 510, when the system is in the content tracking mode, the system can cause a second user interface layout 101B to be displayed. The second user interface layout 101B can include a rendering 151A of the video stream of the selected participant 10A positioned relative to a region 142A in a second set of regions 142 that do not include the shared content 120. The second user interface layout 101B can also display the shared content in the first set of regions 141. In some configurations, the display of the second user interface layout 101B with the selected participant 10A is only displayed on the computer of the user who needs the assistant. The user who needs the assistant is defined herein as a user with a prerequisite, where the prerequisite corresponds to the role of the assistant. The computers of users who do not have a prerequisite corresponding to the role of the assistant do not cause the tracking mode, and thus those computers do not display the second user interface layout 101B with the selected participant 10A.

[0075] The following terms are for the above and Figure 7Supplement to the disclosure of the operations shown. Each clause includes features that can be specifically combined with any other clause. For example, a clause that depends on clause A can be interpreted as depending on other clauses, such as clause I.

[0076] Clause A: A computer-implemented method for controlling the rendering position of the video stream of a selected participant (10A) in a communication session (603) during the display of shared content (120), the method being for execution on a system (100), the method comprising: accessing a data structure (400) defining preferences for controlling the rendering position of the video stream of the selected participant (10A) when a content tracking mode is activated, the preferences identifying one or more content data types (411) for causing the activation of the content tracking mode, wherein the activation of the content tracking mode causes the system to control the rendering position of the video stream of the selected participant (10A) to minimize the overlap between the rendering of the video stream of the selected participant (10A) and the rendering of the shared content (120) having one or more content data types (411); for example, in FIG. 5, depending on the shared content type and user settings (such as "hearing assistance required"), the system adaptively places the video of a particular attendee (such as "sign language interpreter") based on the role of the attendee; when the system is in a normal operation mode, causing a first user interface layout (101A) to be displayed. The first user interface layout (101A) includes the rendering (151A) of the video stream of the selected participant (10A), and when the system (100) is in a normal operation mode, the rendering (151A) of the video stream of the selected participant (10A) is located at a predetermined position; Figure 1A Shows a system starting in a normal operation mode, where the video rendering of the assistant is static within the UI; receiving an input that causes a state change from the normal operation mode to the content tracking mode, wherein the input identifies the shared content (120) having one or more content data types (411) to be displayed to one or more client devices (11) of the participants (10) in the communication session (603); Figure 1B Shows a user input or system control indicating content sharing of a particular data type for activating the tracking mode; in response to an input that causes a state change from the normal operation mode to the content tracking mode and based on the shared content (120) having one or more content data types (411), analyzing the rendering of the shared content (120) to identify a first set of regions (141) of the shared content (120) at a display threshold level and a second set of regions (142) of the shared content (120) below the display threshold level; Figure 1CShows that the analysis and recognition of shared content includes areas with content and areas without content; and causes a second user interface layout (101B) to be displayed when the system is in content tracking mode, the second user interface layout (101B) including the rendering (151A) of the video stream of a selected participant (10A) in an area (142A) of a second set of areas (142) of the shared content (120) that do not include a threshold level, wherein the video stream of the selected participant (10A) is displayed simultaneously with the shared content (120) located within the first set of areas (141); Attached Figure 1D Shows positioning the rendering of the assistant's video in an area without content.

[0077] Clause B: The method according to the other clauses further includes: accessing settings (400) that persist across multiple communication sessions of a user (10J), the settings defining individual prerequisites (410) for the user (10J), and the system automatically performing access to the settings without user input. Wherein, the access to the settings and the selection of at least one selected user (10A) having a role (420A) corresponding to at least one prerequisite (410A) of the user (10J) are automatically performed by the system (100) in response to the user (10J) joining the communication session; and analyzing a data structure (401) that associates individual users (10A 10K 10L) with one or more roles (420), wherein the analysis of the data structure identifies one or more user profiles (430A) of at least one selected user (10A) having a role (420A) corresponding to at least one prerequisite (410A) of the user (1OJ), wherein analyzing the data structure (401) to identify at least one selected user (10A) causes the first user interface layout (101A) to include the rendering (151A) of the video stream of the selected participant (10A); The user settings determine the match between the role of a participant (e.g., sign language interpreter) and the participant's prerequisites.

[0078] Clause C: The method according to the other clauses further includes: receiving a control input from a computing device (11J) in communication with a display screen showing the second user interface layout (101B), wherein the control input includes coordinates within the second user interface layout (101B); and moving the rendering (151A) of the video stream of the selected participant (10A) according to the coordinates indicated by the control input, wherein the control input is configured to control the movement of the rendering (151A) of the video stream of the selected participant (10A) during the content tracking mode, and wherein the control input is restricted from controlling the movement of the rendering (151A) of the video stream of the selected participant (10A) during the normal operation mode. Figure 3 Shows how any user of the communication session can adjust the mouse control of the assistant's rendering.

[0079] Clause D: The method according to other clauses further includes: receiving an update to the shared content (120); in response to receiving the update to the shared content (120), identifying a new region (142E) that does not include the update to the shared content (120); and moving the rendering (151A) of the video stream of the selected participant (10A) to the new region (142E) that does not include the update to the shared content (120). Figure 1E Shows the dynamic movement of the assistant when the content changes.

[0080] Clause E: The method according to other clauses further includes: determining that one or more dimensions of a region (142E) from a second set of regions (142) that does not include the shared content (120) are less than a threshold dimension for the rendering (151A) of the video stream of the selected participant (10A); in response to determining that one or more dimensions of a region (142E) from the second set of regions (142) that does not include the shared content (120) at a threshold level are less than the threshold dimension for the rendering (151A) of the video stream of the selected participant (10A), reducing at least one dimension of the rendering (151A) of the video stream of the selected participant (10A) to a threshold minimum size when in content tracking mode. Figure 2 Shows resizing the rendering of the assistant during large or full-screen content display.

[0081] Clause F: The method according to other clauses further includes: determining that one or more dimensions of a region (142E) from a second set of regions (142) that does not include the shared content (120) are less than a threshold dimension for the rendering (151A) of the video stream of the selected participant (10A); in response to determining that one or more dimensions of a region (142E) from the second set of regions (142) that does not include the shared content (120) at a threshold level are less than the threshold dimension for the rendering (151A) of the video stream of the selected participant (10A), positioning the rendering (151A) of the video stream of the selected participant (10A) to a position that minimizes the overlap between the displayed portion of the shared content (120) and the rendering (151A) of the video stream of the selected participant (10A). Figure 2 : Moving the rendering of the assistant during large or full-screen content display, for example, to the side or bottom of the content.

[0082] Clause G: The method according to other clauses, wherein the normal operating mode causes the system to position the rendering of the video stream at a static location, where the rendering (151A) of the video stream of the selected participant (10A) does not move according to the position of the rendering of the shared content.

[0083] Clause H: According to the method described in other clauses, wherein the selection (411) of shared content having a data type other than the one or more content data types does not cause a state change from the normal operation mode to the content tracking mode. For data types other than the selected content data types, the tracking mode is not enabled (411). For example, when someone sends an instant message, the tracking mode is not enabled.

[0084] Clause I: A computer-implemented method for controlling the rendering (151A) position of the video stream of a selected participant (10A) in a communication session (603) during the display of shared content (120) shared by a presenter, the method being for execution on a system (100), the method comprising: accessing a data structure (400) defining preferences for controlling the rendering position of the video stream of the selected participant (10A) when the content tracking mode is activated, the data structure identifying one or more content data formats (411) for causing the activation of the content tracking mode, wherein the activation of the content tracking mode causes the system to control the rendering position of the video stream of the selected participant (10A) during the sharing of the shared content (120) by the presenter, thereby controlling the overlap between the rendering of the video stream of the selected participant (10A) and the rendering of the shared content (120) having one or more content data formats (411), wherein the selected participant (10A) is different from the presenter sharing the shared content; causing a first user interface arrangement (101A) to be displayed when the system is in the normal operation mode. The first user interface arrangement (101A) includes the rendering (151A) of the video stream of the selected participant (10A), and when the system (100) is in the normal operation mode, the rendering (151A) of the video stream of the selected participant (10A) is located at a predetermined position; receiving an input that causes a state change from the normal operation mode to the content tracking mode; in response to the input that causes a state change from the normal operation mode to the content tracking mode and based on the shared content (120) having one or more content data formats (411), and based on the role of the selected participant (10A) corresponding to the prerequisite of the user (10J): analyzing the rendering of the shared content (120) to identify a first set of regions (141) of the shared content (120) that display a first threshold level and a second set of regions (142) of the shared content (120) that do not display a second threshold level; and causing a second user interface arrangement (101B) to be displayed when the system is in the content tracking mode, the second user interface arrangement (101B) including the rendering (151A) of the video stream of the selected participant (10A) in a region (142A) located in the second set of regions (142) that do not include the shared content (120) at the second threshold level, wherein the video stream of the selected participant (10A) is displayed simultaneously with the shared content (120) located within the first set of regions (141).

[0085] Clause J: The method according to other clauses further includes: accessing settings (400) that persist across multiple communication sessions of a user (10J), where the settings define separate preconditions (410) for the user (10J), and the access to the settings is automatically performed by the system without user input; in response to the user (10J) joining a communication session, automatically selecting, by the system (100), at least one selected user (10A) having a role (420A) corresponding to at least one precondition (410A) of the user (10J); analyzing a data structure (401) that associates individual users (10A, 10K, 10L) with one or more roles (420); and identifying one or more user profiles (430A) of at least one selected user (10A) having a role (420A) corresponding to at least one precondition (410A) of the user (10J), wherein the identification of the at least one selected user (10A) causes a second user interface layout (101B) to include a rendering (151A) of the video stream of the selected participant (10A), and wherein the selected participant is not displayed on the computing device of a user not associated with the preconditions of one or more roles corresponding to the individual user.

[0086] Clause K: The method according to other clauses further includes: receiving a control input from a computing device (11J) communicating with a display screen showing the second user interface layout (101B), where the control input includes coordinates within the second user interface layout (101B); and moving the rendering (151A) of the video stream of the selected participant (10A) according to the coordinates indicated by the control input, where the control input is configured to control the movement of the rendering (151A) of the video stream of the selected participant (10A) during the content tracking mode, and where the control input is restricted to controlling the movement of the rendering (151A) of the video stream of the selected participant (10A) during the normal operation mode.

[0087] Clause L: The method according to other clauses further includes: receiving an update (120) to shared content; in response to receiving the update (120) to shared content, identifying a new region (142E) that does not include the update (120) to shared content; and moving the rendering (151A) of the video stream of the selected participant (10A) to the new region (142E) that does not include the update (120) to shared content.

[0088] Clause M: The method according to other clauses further includes: determining that one or more dimensions of the region (142E) from the second set of regions (142) that do not include the shared content (120) are less than a threshold dimension for the rendering (151A) of the video stream of the selected participant (10A); in response to determining that one or more dimensions of the region (142E) from the second set of regions (142) that do not include the shared content (120) at a second threshold level are less than a threshold dimension for the rendering (151A) of the video stream of the selected participant (10A), when in the content tracking mode, reducing at least one dimension of the rendering (151A) of the video stream of the selected participant (10A) to a threshold minimum size.

[0089] Clause N: The method according to other clauses further includes: determining that one or more dimensions of the region (142E) from the second set of regions (142) that do not include the shared content (120) are less than a threshold dimension for the rendering (151A) of the video stream of the selected participant (10A); in response to determining that one or more dimensions of the region (142E) from the second set of regions (142) that do not include the shared content (120) at a second threshold level are less than a threshold dimension for the rendering (151A) of the video stream of the selected participant (10A), positioning the rendering (151A) of the video stream of the selected participant (10A) to a position that minimizes the overlap between the displayed portion of the shared content (120) and the rendering (151A) of the video stream of the selected participant (10A).

[0090] Clause O: In the method according to other clauses, the input identifies the shared content (120) having the one or more content data types (411) for display to one or more client devices (11) of the participants (10) of the communication session (603), and the input is received from a computing device associated with the presenter.

[0091] Clause P: In the method according to other clauses, an input that causes a state change from the normal operation mode to the content tracking mode is received from a computing device associated with the user.

[0092] Figure 8FIG. is a diagram illustrating an example environment 600 in which system 602 may implement the techniques disclosed herein. It should be understood that the subject matter described above may be implemented as a computer-controlled apparatus, a computer process, a computing system, or an article of manufacture such as a computer-readable storage medium. Operations of example methods are illustrated in various boxes and are outlined with reference to these boxes. The methods are shown as a logical flow of blocks, each of which may represent one or more operations that may be implemented in hardware, software, or a combination thereof. In a software context, the operations represent computer-executable instructions stored on one or more computer-readable media that, when executed by one or more processors, cause the one or more processors to perform the recited operations.

[0093] Generally, computer-executable instructions include routines, programs, objects, modules, components, data structures, etc. that perform particular functions or implement particular abstract data types. The order of description of the operations is not intended to be construed as limiting, and any number of the described operations may be executed in any order, combined in any order, subdivided into multiple sub-operations, and / or executed in parallel to implement the described process. The described process may be executed by resources associated with one or more devices, such as one or more internal or external CPUs or GPUs, and / or one or more hardware logics, such as field-programmable gate arrays (“FPGAs”), digital signal processors (“DSPs”), or other types of accelerators.

[0094] All of the above methods and processes may be embodied in software code modules executed by one or more general-purpose computers or processors and be fully automated via the software code modules. The code modules may be stored in any type of computer-readable storage medium or other computer storage device, such as those described below. Some or all of the methods may alternatively be embodied in dedicated computer hardware, such as the computer hardware described below.

[0095] Any routine descriptions, elements, or blocks in the flowcharts described and / or depicted in the figures herein should be understood as potentially representing code modules, segments, or portions that include one or more executable instructions for implementing specific logical functions or elements in the routine. Alternative implementations are included within the scope of the examples described herein, where elements or functions may be deleted, or executed out of the order shown or discussed, including substantially synchronously or in reverse order, depending on the functionality involved, as would be understood by those skilled in the art.

[0096] In some embodiments, system 602 can be used to collect, analyze, and share data displayed to the users of communication session 604. As shown, communication session 603 can be implemented between multiple client computing devices 606(1) to 606(N) (where N is a number with a value of 2 or greater) associated with or as part of system 602. The client computing devices 606(1) to 606(N) enable users (also referred to as individuals) to participate in communication session 603.

[0097] In this example, communication session 603 is hosted by system 602 via one or more networks 608. That is, system 602 can provide services that enable the users of client computing devices 606(1) to 606(N) to participate in communication session 603 (e.g., via live viewing and / or recorded viewing). Thus, the "participants" in communication session 603 can include users and / or client computing devices (e.g., multiple users can be in a room participating in a communication session via the use of a single client computing device), and each user can communicate with other participants. Alternatively, communication session 603 can be hosted by one of the client computing devices 606(1) to 606(N) using peer-to-peer technology. System 602 can also host chat conversations and other team collaboration functions (e.g., as part of an application suite).

[0098] In some embodiments, such chat conversations and other team collaboration functions are considered to be external communication sessions distinct from communication session 603. The computing system 602 that collects participant data in communication session 603 can be capable of linking to such external communication sessions. Thus, the system can receive information such as date, time, session details, etc., that enables connection to such external communication sessions. In one example, a chat conversation can be conducted based on communication session 603. Additionally, system 602 can host communication session 603, which includes at least multiple participants co-located at a meeting location (e.g., a conference room or auditorium) or located at different locations.

[0099] In the examples described herein, the client computing devices 606(1) through 606(N) participating in the communication session 603 are configured to receive and render communication data for display on a user interface of a display screen. The communication data can include a collection of various instances or streams of live content and / or recorded content. The collection of various instances or streams of live content and / or recorded content can be provided by one or more cameras, such as video cameras. For example, an individual stream of live or recorded content can include media data associated with a video feed provided by a video camera (e.g., audio and visual data capturing the appearance and speech of a user participating in the communication session). In some embodiments, the video feed can include such audio and visual data, one or more still images, and / or one or more avatars. The one or more still images can also include one or more avatars.

[0100] Another example of an individual stream of live or recorded content can include media data that includes an avatar of a user participating in the communication session and audio data capturing the user's speech. Yet another example of an individual stream of live or recorded content can include media data that includes a file displayed on the display screen and audio data capturing the user's speech. Thus, the various streams of live or recorded content within the communication data enable facilitating a remote conference among a group of people and sharing content within the group. In some embodiments, the streams of live or recorded content within the communication data can originate from multiple co-located cameras located in a space such as a room to record or live a content presented by one or more individuals presenting the content and one or more individuals consuming the presented content.

[0101] Participants or attendees can watch the content of the communication session 603 in real time as the event occurs or, alternatively, watch the content of the communication session 603 via a recording at a later time after the event has occurred. In the examples described herein, the client computing devices 606(1) through 606(N) participating in the communication session 603 are configured to receive and render communication data for display on a user interface of a display screen. The communication data can include a collection of various instances or streams of live and / or recorded content. For example, a single content stream can include media data associated with a video feed (e.g., audio and visual data capturing the appearance and speech of a user participating in the communication session). Another example of a single content stream can include media data that includes an avatar of a user participating in the conference session and audio data capturing the user's speech. Yet another example of a single content stream can include media data and / or audio data, the media data including a content item displayed on the display screen, and the audio data capturing the user's speech. Thus, the various content streams within the communication data enable facilitating a conference or a broadcast presentation among a group of people dispersed at remote locations.

[0102] The participants or attendees of a communication session are people within the range of a camera or other image and / or audio capture device, such that the actions and / or sounds of the people that occur while the people are viewing and / or listening to the content shared via the communication session can be captured (e.g., recorded). For example, a participant may be sitting in a crowd watching shared content live at a broadcast location where a stage presentation is taking place. Or a participant can be sitting in an office conference room, viewing the shared content of a communication session with other colleagues via a display screen. Further still, a participant can be sitting or standing in front of a personal device (e.g., a tablet computer, a smart phone, a computer, etc.), viewing the shared content of a communication session alone in their office or at home.

[0103] Figure 8 System 602 includes device 610. Device 610 and / or other components of system 602 can include distributed computing resources that communicate with each other and / or with client computing devices 606(1) through 606(N) via one or more networks 608. In some examples, system 602 can be an independent system responsible for managing aspects of one or more communication sessions such as communication session 603. As an example, system 602 can be managed by entities such as SLACK, WEBEX, GOTOMEETING, GOOGLE HANGOUTS, etc.

[0104] Network 608 can include, for example, a public network such as the Internet, a private network such as an institutional and / or personal intranet, or some combination of private and public networks. Network 608 can also include any type of wired and / or wireless network, including but not limited to a local area network (“LAN”), a wide area network (“WAN”), a satellite network, a wired network, a Wi-Fi network, a WiMax network, a mobile communication network (e.g., 3G, 4G, etc.), or any combination thereof. Network 608 can utilize communication protocols, including packet-based and / or datagram-based protocols such as the Internet Protocol (“IP”), the Transmission Control Protocol (“TCP”), the User Datagram Protocol (“UDP”), or other types of protocols. In addition, network 608 can also include multiple devices that facilitate network communication and / or form the hardware foundation of the network, such as switches, routers, gateways, access points, firewalls, base stations, repeaters, backbone devices, etc.

[0105] In some examples, network 608 can also include devices that enable connection to a wireless network, such as a wireless access point (“WAP”). Examples support the connection of WAPs that send and receive data at various electromagnetic frequencies (e.g., radio frequency), including WAPs that support Institute of Electrical and Electronics Engineers (“IEEE”) 802.11 standards (e.g., 802.11g, 802.11n, 802.11ac, etc.) and other standards.

[0106] In various examples, device 610 may include one or more computing devices operating in a cluster or other grouped configuration to share resources, balance loads, improve performance, provide failover support or redundancy, or for other purposes. For example, device 610 may belong to various categories of devices, such as traditional server-type devices, desktop computer-type devices, and / or mobile-type devices. Thus, although shown as a single type of device or a server-type device, device 610 may include various device types and is not limited to a specific type of device. Device 610 may represent, but is not limited to, a server computer, a desktop computer, a web server computer, a personal computer, a mobile computer, a laptop computer, a tablet computer, or any other kind of computing device.

[0107] Client computing devices (e.g., one of client computing devices 606(1) to 606(N), each of which is also referred to herein as a "data processing system") may belong to various device categories, which may be the same as or different from those of device 610, such as traditional client-type devices, desktop computer-type devices, mobile-type devices, dedicated-type devices, embedded-type devices, and / or wearable-type devices. Thus, client computing devices may include, but are not limited to, desktop computers, game consoles and / or game devices, tablet computers, personal data assistants ("PDAs"), mobile phone / tablet hybrid devices, laptop computers, telecommunications devices, computer navigation-type client computing devices, such as satellite-based navigation systems, including global positioning system ("GPS") devices, wearable devices, virtual reality ("VR") devices, augmented reality ("AR") devices, implantable computing devices, automotive computers, network-enabled televisions, thin clients, terminals, Internet of Things ("IoT") devices, workstations, media players, personal video recorders ("PVRs"), set-top boxes, cameras, integrated components for inclusion in a computing device (e.g., peripherals), appliances, or any other kind of computing device. Additionally, client computing devices may include combinations of the previously listed examples of client computing devices, such as, for example, a desktop computer-type device or a mobile-type device combined with a wearable device, etc.

[0108] Client computing devices 606(1) to 606(N) of various categories and device types may represent any type of computing device having one or more data processing units 692 operatively connected to a computer-readable medium 694 via a bus 616, which in some instances may include a system bus, a data bus, an address bus, a PCI bus, a Mini-PCI bus, and one or more of any of various local, peripheral, and / or independent buses.

[0109] The executable instructions stored on the computer-readable medium 694 can include, for example, an operating system 619, a client module 620, a profile module 622, and other modules, programs, or applications that can be loaded and executed by the data processing unit 692.

[0110] The client computing devices 606(1) to 606(N) can also include one or more interfaces 624 to enable communication between the client computing devices 606(1) to 606(N) and other networked devices (such as device 610) via the network 608. Such network interfaces 624 can include one or more network interface controllers (NICs) or other types of transceiver devices to send and receive communications and / or data over the network. Additionally, the client computing devices 606(1) to 606(N) can include input / output (“I / O”) interfaces (devices) 626 that enable communication with input / output devices, such as user input devices including peripheral input devices (e.g., game controllers, keyboards, mice, pens, voice input devices such as microphones, cameras for obtaining and providing video feeds and / or still images, touch input devices, gesture input devices, etc.) and / or output devices including peripheral output devices (e.g., displays, printers, audio speakers, haptic output devices, etc.). Figure 8 It is shown that the client computing device 606(1) is connected to a display device (e.g., display screen 629(N)) in a certain way, which can display the UI according to the techniques described herein.

[0111] In Figure 8 the example environment 600, the client computing devices 606(1) to 606(N) can use their respective client modules 620 to connect to each other and / or to other external devices to participate in a communication session 603 or to contribute activities to the collaborative environment. For example, a first user can utilize the client computing device 606(1) to communicate with a second user of another client computing device 606(2). When the client module 620 is executed, the users can share data, which can cause the client computing device 606(1) to connect to the system 602 and / or other client computing devices 606(2) to 606(N) via the network 608.

[0112] The client computing devices 606(1) to 606(N) can use their respective profile modules 622 to generate participant profiles ( Figure 8(not shown in the figure), and provides the participant profile to other client computing devices and / or the device 610 of the system 602. The participant profile may include one or more of the identity of the user or user group (e.g., name, unique identifier ("ID"), etc.), user data such as personal data, machine data such as location (e.g., IP address, room in a building, etc.), and technical capabilities. The participant profile can be used to register the participants of the communication session.

[0113] As Figure 8 shown, the device 610 of the system 602 includes a server module 630 and an output module 632. In this example, the server module 630 is configured to receive media streams 634(1) to 634(N) from respective client computing devices such as client computing devices 606(1) to 606(N). As described above, the media streams may include video feeds (e.g., audio and visual data associated with a user), audio data to be output together with the presentation of the user's avatar (e.g., an audio-only experience without sending the user's video data), text data (e.g., text messages), file data, and / or screen sharing data (e.g., documents, slides, images, videos displayed on a display screen, etc.). Thus, the server module 630 is configured to receive a collection of various media streams 634(1) to 634(N) (this collection is referred to herein as "media data 634") during the live viewing of the communication session 603. In some scenarios, not all client computing devices participating in the communication session 603 provide media streams. For example, a client computing device may be only a consumption or "listening" device, such that it only receives the content associated with the communication session 603 but does not provide any content to the communication session 603.

[0114] In various examples, the server module 630 may select aspects of the media stream 634 to share with each of the participating client computing devices 606(1) through 606(N). Thus, the server module 630 may be configured to generate session data 636 based on the media stream 634 and / or pass the session data 636 to the output module 632. The output module 632 may then transmit communication data 639 to the client computing devices (e.g., client computing devices 606(1) through 606(3) participating in the live viewing of the communication session). The communication data 639 may include video, audio, and / or other content data provided by the output module 632 based on the content 650 associated with the output module 632 and based on the received session data 636. The content 650 may include the media stream 634 or other shared data, such as image files, spreadsheet files, slides, documents, etc. The media stream 634 may include a video component depicting images captured by the I / O device 626 on each client computer. The content 650 also includes input data from each user, which may be used to control the orientation and position of the representation. The content may also include instructions for sharing data and identifiers of the recipients of the shared data. Thus, the content 650 is also referred to herein as input data 650 or input 650.

[0115] As shown, the output module 632 sends communication data 639(1) to the client computing device 606(1), and sends communication data 639(2) to the client computing device 606(2), and sends communication data 639(3) to the client computing device 606(3), etc. The communication data 639 transmitted to the client computing devices may be the same or may be different (e.g., the positioning of the content stream within the user interface may vary depending on the device).

[0116] In various embodiments, the device 610 and / or the client module 620 may include a GUI rendering module 640. The GUI rendering module 640 may be configured to analyze the communication data 639 for delivery to one or more of the client computing devices 606. Specifically, the UI rendering module 640 at the device 610 and / or the client computing device 606 may analyze the communication data 639 to determine an appropriate manner for displaying video, images, and / or content on the display screen 629 of the associated client computing device 606. In some embodiments, the GUI rendering module 640 may provide video, images, and / or content to a rendered GUI 646 presented on the display screen 629 of the associated client computing device 606. The rendered GUI 646 may be caused to be presented on the display screen 629 by the GUI rendering module 640. The rendered GUI 646 may include video, images, and / or content analyzed by the GUI rendering module 640.

[0117] In some embodiments, the presented GUI 646 can include multiple sections or grids that can render or include video, images, and / or content for display on the display screen 629. For example, a first section of the presentation GUI 646 can include a video feed of a presenter or individual, and a second section of the presentation GUI 646 can include a video feed of an individual consuming the meeting information provided by the presenter or individual. The GUI presentation module 640 can populate the first and second sections of the presented GUI 646 in a manner that appropriately mimics the environmental experience that the presenter and individual may be sharing.

[0118] In some embodiments, the GUI presentation module 640 can zoom in or provide a zoomed view of an individual represented by a video feed to highlight the individual's reaction to the presenter, such as facial features. In some embodiments, the presented GUI 646 can include video feeds of multiple participants associated with the meeting, such as a general communication session. In other embodiments, the presented GUI 646 can be associated with a channel, such as a chat channel, an enterprise team channel, etc. Thus, the presented GUI 646 can be associated with an external communication session different from a general communication session. Figure 9 A diagram shows example components of an example device 700 (also referred to herein as a "computing device") configured to generate data for some of the user interfaces disclosed herein. The device 700 can generate data that can include one or more sections that can render or include video, images, virtual objects, and / or content for display on the display screen 629. The device 700 can represent one of the devices described herein. Additionally or alternatively, the device 700 can represent one of the client computing devices 606.

[0119] As shown, the device 700 includes one or more data processing units 702, a computer-readable medium 704, and a communication interface 706. The components of the device 700 are operably connected, for example, via a bus 709, which can include a system bus, a data bus, an address bus, a PCI bus, a Mini-PCI bus, and one or more of any of various local, peripheral, and / or independent buses.

[0120] As used herein, a data processing unit (such as data processing unit 702 and / or data processing unit 692) can represent, for example, a CPU-type data processing unit, a GPU-type data processing unit, a field-programmable gate array ("FPGA"), another type of DSP, or other hardware logic components that can be driven by a CPU in some cases. Illustrative types of hardware logic components that can be utilized, for example but not limited to, include application-specific integrated circuits ("ASICs"), application-specific standard products ("ASSPs"), systems-on-a-chip ("SOCs"), complex programmable logic devices ("CPLDs"), etc.

[0121] As used herein, computer-readable media such as computer-readable medium 704 and computer-readable medium 694 may store instructions executable by a data processing unit. The computer-readable media may also store instructions executable by an external data processing unit such as an external CPU, an external GPU, and / or executable by an external accelerator such as an FPGA-type accelerator, a DSP-type accelerator, or any other internal or external accelerator. In various examples, at least one of the CPU, GPU, and / or accelerator is incorporated into the computing device, while in some examples, one or more of the CPU, GPU, and / or accelerator are external to the computing device.

[0122] Computer-readable media (which may also be referred to herein as computer-readable media) may include computer storage media and / or communication media. Computer storage media may include volatile memory, non-volatile memory, and / or other persistent and / or auxiliary computer storage media, one or more of the removable and non-removable computer storage media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Thus, computer storage media includes tangible and / or physical forms of media that are included in a device and / or hardware component as part of a device or external to a device, including but not limited to random access memory (“RAM”), static random access memory (“SRAM”), dynamic random access memory (“DRAM”), phase change memory (“PCM”), read-only memory (“ROM”), erasable programmable read-only memory (“EPROM”), electrically erasable programmable read-only memory (“EEPROM”), flash memory, compact disc read-only memory (“CD-ROM”), digital versatile disc (“DVD”), optical card, or other optical storage media, magnetic tape cartridge, magnetic tape, disk storage, magnetic card, or other magnetic storage device or media, solid-state memory device, storage array, network-attached storage, storage area network, hosted computer storage, or any other storage memory, storage device, and / or storage media that may be used to store and maintain information for access by a computing device. Computer storage media may also be referred to herein as computer-readable storage media, non-transitory computer-readable storage media, non-transitory computer-readable media, computer-readable storage media, computer-readable storage devices, or computer storage media.

[0123] In contrast to computer storage media, communication media may embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism. As defined herein, computer storage media does not include communication media. That is, computer storage media does not include communication media consisting solely of a modulated data signal, a carrier wave, or a propagated signal itself.

[0124] The communication interface 706 can represent, for example, a network interface controller (“NIC”) or other types of transceiver devices that send and receive communications over a network. Additionally, the communication interface 706 can include one or more cameras and / or audio devices 722 to enable generation of video feeds and / or still images, etc.

[0125] In the example shown, the computer-readable medium 704 includes a data repository 708. In some examples, the data repository 708 includes data repositories such as databases, data warehouses, or other types of structured or unstructured data repositories. In some examples, the data repository 708 includes a corpus and / or a relational database having one or more tables, indexes, stored procedures, etc. to enable data access including one or more of, for example, Hypertext Markup Language (“HTML”) tables, Resource Description Framework (“RDF”) tables, Web Ontology Language (“OWL”) tables, and / or Extensible Markup Language (“XML”) tables.

[0126] The data repository 708 can store data for the operation of processes, applications, components, and / or modules stored in the computer-readable medium 704 and / or executed by the data processing unit 702 and / or the accelerator. For example, in some examples, the data repository 708 can store session data 710 (e.g., such as Figure 8 the session data 636 shown), profile data 712 (e.g., associated with a participant profile), and / or other data. The session data 710 can include the total number of participants (e.g., users and / or client computing devices) in a communication session, the activities that occur in the communication session, a list of invitees to the communication session, and / or other data related to when and how the communication session is conducted or hosted. The hardware data 711 can define aspects of any device, such as multiple display screens of a computer. The session data can also define any type of activity or status related to individual users 10A - 10L, each user associated with an individual video stream among the multiple video media streams 634. For example, the context data can define the level of a person in an organization, how each person's level relates to the levels of other people, the performance level of a person, or any other activity or status information that can be used to determine the rendering location of a person within a virtual environment. This context information can also be fed into any model to help emphasize keywords spoken by a person at a particular level, highlight the UI when the background sound of a person at a particular level is detected, or change the mood display in a particular way when a person at a particular level is detected to have a particular mood.

[0127] Alternatively, some or all of the above data may be stored on a separate memory 716 on one or more data processing units 702, such as a memory on a CPU-type processor, a GPU-type processor, an FPGA-type accelerator, a DSP-type accelerator, and / or another accelerator. In this example, the computer-readable medium 704 further includes an operating system 718 and an application programming interface 710 (API) configured to expose the functionality and data of the device 700 to other devices. Additionally, the computer-readable medium 704 includes one or more modules, such as a server module 730, an output module 732, and a GUI rendering module 740, although the number of illustrated modules is merely an example and the number may vary. That is, the functionality associated with the illustrated modules described herein may be performed by fewer or more modules on one device, or distributed across multiple devices.

[0128] Finally, although various configurations have been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claimed subject matter.

Claims

1. A computer-implemented method for controlling the rendering position of a video stream of a selected participant in a communication session during the display of shared content shared by a presenter, the method being for execution on a system, the method comprising: Accessing a data structure defining preferences for controlling the rendering position of the video stream of the selected participant when a content tracking mode is activated, the data structure identifying one or more content data formats for causing the activation of the content tracking mode, wherein the activation of the content tracking mode causes the system to control the rendering position of the video stream of the selected participant during the sharing of the shared content by the presenter to control the overlap between the rendering of the video stream of the selected participant and the rendering of the shared content having the one or more content data formats, wherein the selected participant is different from the presenter sharing the shared content; Causing a first user interface arrangement to be displayed when the system is in a normal operation mode, the first user interface arrangement including the rendering of the video stream of the selected participant, and when the system is in the normal operation mode, the rendering of the video stream of the selected participant is located at a predetermined position; Receiving an input that causes a state change from the normal operation mode to the content tracking mode; In response to an input that causes a state change from the normal operation mode to the content tracking mode and based on the shared content having the one or more content data formats, and based on the role of the selected participant corresponding to the user's preconditions: Analyzing the rendering of the shared content to identify a first set of regions of the shared content that display a first threshold level and a second set of regions of the shared content that do not display a second threshold level; and Causing a second user interface arrangement to be displayed when the system is in the content tracking mode, the second user interface arrangement including the rendering of the video stream of the selected participant in a region that does not include the second set of regions of the shared content that do not display the second threshold level, wherein the video stream of the selected participant is displayed simultaneously with the shared content located within the first set of regions.

2. The method according to claim 1, further comprising: Accessing settings that persist across multiple communication sessions of a user, wherein the settings define the individual preconditions of the user, and the access to the settings is automatically performed by the system without user input; Selecting at least one selected user having a role corresponding to at least one precondition of the user is automatically performed by the system in response to the user joining a communication session; Analyzing a data structure that associates individual users with one or more roles; and Identify one or more user profiles of the at least one selected user having a role corresponding to the at least one prerequisite of the user, wherein identifying the at least one selected user causes the second user interface arrangement to include rendering of a video stream of the selected participant, wherein the selected participant is not displayed on a computing device of a user not associated with a prerequisite corresponding to the one or more roles of the individual user.

3. The method according to claim 1, further comprising: Receiving a control input from a computing device communicating with a display screen displaying the second user interface arrangement, wherein the control input includes coordinates within the second user interface arrangement; and Moving the rendering of the video stream of the selected participant according to the coordinates indicated by the control input, wherein the control input is configured to control movement of the rendering of the video stream of the selected participant during the content tracking mode, wherein the control input is restricted to control movement of the rendering of the video stream of the selected participant during the normal operation mode.

4. The method according to claim 1, further comprising: Receiving an update to the shared content; In response to receiving the update to the shared content, identifying a new area that does not include the update to the shared content; And Moving the rendering of the video stream of the selected participant to the new area that does not include the update to the shared content.

5. The method according to claim 1, further comprising: Determining that one or more dimensions of an area from a second set of areas not including the shared content are less than a threshold dimension of the rendering of the video stream of the selected participant; In response to determining that one or more dimensions of an area from a second set of areas not including the shared content at the second threshold level are less than the threshold dimension of the rendering of the video stream of the selected participant, reducing at least one dimension of the rendering of the video stream of the selected participant to a threshold minimum size when in the content tracking mode.

6. The method according to claim 1, further comprising: Determining that one or more dimensions of an area from a second set of areas not including the shared content are less than a threshold dimension of the rendering of the video stream of the selected participant; In response to determining that one or more dimensions of an area from a second set of areas not including the shared content at the second threshold level are less than the threshold dimension of the rendering of the video stream of the selected participant, positioning the rendering of the video stream of the selected participant at a location that minimizes overlap between a displayed portion of the shared content and the rendering of the video stream of the selected participant.

7. The method according to claim 1, wherein, The input identifies the shared content having the one or more content data types for display to one or more client devices of participants in the communication session, and the input is received from a computing device associated with the presenter.

8. The method according to claim 1, wherein, The input that causes the state change from the normal operation mode to the content tracking mode is received from a computing device associated with the user.

9. A computing device for controlling the rendering position of a video stream of a selected participant in a communication session during the display of shared content shared by a presenter, the method for execution on a system, the computing device comprising: One or more processing units; And A computer-readable storage medium encoded with computer-executable instructions that cause the one or more processing units to perform the following operations: Access a data structure defining preferences for controlling the rendering position of the video stream of the selected participant when a content tracking mode is activated, the data structure identifying one or more content data formats for causing activation of the content tracking mode, wherein activation of the content tracking mode causes the system to control the rendering position of the video stream of the selected participant during sharing of the shared content by the presenter to control the overlap between the rendering of the video stream of the selected participant and the rendering of the shared content having the one or more content data formats, wherein the selected participant is different from the presenter sharing the shared content; Cause a first user interface arrangement to be displayed when the system is in a normal operation mode, the first user interface arrangement including the rendering of the video stream of the selected participant, and when the system is in the normal operation mode, the rendering of the video stream of the selected participant is located at a predetermined position; Receive an input that causes a state change from the normal operation mode to the content tracking mode; In response to the input that causes the state change from the normal operation mode to the content tracking mode and based on the shared content having the one or more content data formats, and based on the role of the selected participant corresponding to the user's prerequisites: Analyze the rendering of the shared content to identify a first set of regions of the shared content that display a first threshold level and a second set of regions of the shared content that do not display a second threshold level; and When the system is in the content tracking mode, cause a second user interface arrangement to be displayed, the second user interface arrangement including the rendering of the video stream of the selected participant in a region that does not include the second set of regions of the shared content that do not display the second threshold level, wherein the video stream of the selected participant is displayed simultaneously with the shared content located within the first set of regions.

10. The computing device according to claim 9, wherein, The display of the second user interface arrangement displays the shared content in an operating system desktop or in an application, the application being an application that executes independently of a communication application that manages the communication session, wherein when the user interface of the communication application is minimized, the second user interface arrangement displays the shared content, and wherein cursor input provided by a computing device associated with an input identifying the shared content causes movement of the rendering of the video stream of the selected participant.

11. The computing device according to claim 9, wherein, The instructions further cause the one or more processing units to perform the following operations: Receive a control input from a computing device communicating with a display screen displaying the second user interface arrangement, wherein the control input includes coordinates within the second user interface arrangement; and Move the rendering of the video stream of the selected participant according to the coordinates indicated by the control input, wherein the control input is configured to control the movement of the rendering of the video stream of the selected participant during the content tracking mode, and wherein the control input is restricted from controlling the movement of the rendering of the video stream of the selected participant during the normal operation mode.

12. The computing device according to claim 9, wherein, The instructions further cause the one or more processing units to: Receive an update to the shared content; In response to receiving the update to the shared content, identify a new region that does not include the update to the shared content; And Move the rendering of the video stream of the selected participant to the new region that does not include the update to the shared content.

13. The computing device according to claim 9, wherein, The instructions further cause the one or more processing units to: Determine that one or more dimensions of a region from a second set of regions of the shared content that do not include the second threshold level are less than a threshold dimension of the rendering of the video stream of the selected participant; In response to determining that one or more dimensions of a region from a second set of regions of the shared content that do not include the second threshold level are less than the threshold dimension of the rendering of the video stream of the selected participant, when in the content tracking mode, reduce at least one dimension of the rendering of the video stream of the selected participant to a threshold minimum size.

14. The computing device according to claim 9, wherein, The instructions further cause the one or more processing units to: Determine that one or more dimensions of a region from a second set of regions of the shared content that do not include the second threshold level are less than a threshold dimension of the rendering of the video stream of the selected participant; In response to determining that one or more dimensions of a region from a second set of regions of the shared content that do not include the second threshold level are less than the threshold dimension of the rendering of the video stream of the selected participant, position the rendering of the video stream of the selected participant to minimize the overlap between the displayed portion of the shared content and the rendering of the video stream of the selected participant.