Identifying graphics interchange format files for inclusion with content of a video game

CN116782986BActive Publication Date: 2026-08-11SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2026-08-11

Smart Images

  • Figure CN116782986B_ABST
    Figure CN116782986B_ABST
Patent Text Reader

Abstract

Methods and systems for representing the emotions of an audience watching an online video game include capturing interaction data from viewers engaged in watching gameplay of the video game. The captured interaction data is used to cluster the audience into different groups based on emotions detected from the interactions of the viewers within the audience. A Graphical Exchange Format (GIF) file is identified for each group based on the emotions associated with those groups. The GIF, representing the unique emotions of the different audience groups, is forwarded to the viewers' client devices for rendering alongside the content of the video game.
Need to check novelty before this filing date? Find Prior Art

Description

1. Technical Field

[0002] This disclosure generally relates to representing the emotions of a video game audience, and more specifically to methods and systems for displaying expressive icons and / or GIFs that mimic emotions detected from different audience groups of video games. Background Technology

[0003] 2. Description of related technologies

[0004] The video game industry has changed significantly over the years. Specifically, online games and live events, such as esports, have seen tremendous growth in terms of the number of live events, viewership, and revenue. Consequently, as the popularity of online games and live events continues to grow, the number of viewers accessing online games (i.e., gameplay) to watch continues to increase. Due to the distributed nature of video games, viewers can connect to and watch online games from anywhere, comfortably from the comfort of their own homes.

[0005] A growing trend in the video game industry is the development and refinement of unique methods to enhance the experience for viewers of online game content and others (e.g., players, commentators, etc.). For example, to provide a truly immersive gaming experience, viewers have various tools (e.g., user interfaces, recording tools, etc.) to express their emotions and share them with other viewers and / or players. For instance, viewers of online game content may have interactive user interfaces with interactive tools, such as chat interfaces, video / audio content upload tools, etc., to comment on online games and communicate with other users. These interactive tools allow viewers to provide audio comments, video comments, text comments, etc., and share emoticons, Graphics Interchange Format images / files (GIFs), etc.

[0006] While various tools offer viewers a degree of engagement, they cannot truly measure the diverse atmospheres expressed by different viewers, nor the number of viewers expressing each atmosphere. Consider esports, where the number of viewers accessing and watching gameplay online can reach thousands or even millions (depending on the popularity of the video game, the players, etc.), and the number of comments shared by viewers can also reach thousands or millions. To allow viewers to experience the diverse atmospheres detected by different viewers within a video game's audience, they would have to analyze every comment provided by different viewers. Due to the sheer volume, analyzing each comment can be overwhelming, especially when comments are generated in real-time during a live video game broadcast. For viewers, connecting with like-minded viewers makes them enjoy the online game (i.e., the gameplay) as if they were playing and watching it with their friends. The inability to see the emotions of other viewers watching the gameplay prevents viewers from fully enjoying the video game's gameplay.

[0007] The implementation method disclosed herein arose in this context. Summary of the Invention

[0008] This disclosure includes methods and systems related to aggregating interactions with viewers of an online video game and rendering expressive avatars on an image representation of the video game audience, wherein the audience includes multiple viewers who have visited the video game to watch gameplay. Various expressive avatars represent the emotions of the audience as part of the audience. Expressive avatars are identified by collecting interaction data from viewers in real time while they are watching gameplay. The interaction data collected from different viewers may include video of the viewers while they are watching gameplay, audio generated by the viewers, text posted in chat interfaces, social media interfaces, etc., graphics-interchange format images / files (GIFs), emoticons, emojis, etc. The interaction data collected from different viewers is aggregated, and sentiment analysis is performed on the aggregated data to identify salient emotions and thoughts conveyed by the viewers. Since a large number of viewers may be visiting the video game to watch gameplay, processing and expressing the emotions of each individual viewer is practically impractical. Aggregating the interaction data allows for consideration of the emotions of each individual viewer in order to determine the different moods expressed within the audience. Data from sentiment analysis is used to determine the different moods (i.e., emotions and reactions) expressed by viewers within an audience.

[0009] Atmospheres detected from different viewers within the audience are categorized, and viewers expressing the same or similar types of atmospheres are grouped together to define atmosphere factions. Data associated with atmosphere factions is used to generate expressive avatars, which are then forwarded to each viewer's client device for rendering. The expressive avatar representing each atmosphere faction is configured to overlay a corresponding segment of the audience's image representation rendered on the viewer's client device. The avatars included in the audience's image representation identify different dominant atmospheres detected from different viewers within the audience. The size of each expressive avatar is scaled to correspond to the size of the corresponding atmosphere faction. The larger the avatar of a particular atmosphere faction, the more dominant the corresponding emotion associated with that avatar is compared to other emotions identified in the audience (i.e., the degree to which the atmosphere detected from the corresponding avatar dominates other atmospheres).

[0010] Viewer-provided interactions can respond to various activities occurring within the gameplay of the video game or to comments / interactions from other viewers. Interactions can take the form of audio comments, text comments, or chat comments provided via a chat interface. Alternatively, viewer interactions can be in video format. For example, viewer expressions of interaction with activities occurring within the video game or with other viewers can be captured using one or more cameras and transmitted to the game cloud server, said cameras being integrated into or communicatively coupled to the viewer's client device. The game cloud server processes the captured images of viewer expressions to determine the emotions expressed by the viewer.

[0011] The system aggregates various audience interactions and performs sentiment analysis on the interaction data to identify salient emotions and thoughts conveyed by the audience. A word cloud is generated and dynamically updated using keywords identified from text and audio content included in the interaction data. These updated keywords capture salient emotions and thoughts expressing the audience's emotional state. Machine learning algorithms are used to identify keywords in the interaction data that express salient emotions and thoughts of the audience and analyze facial features from audience images captured in the video content. Machine learning algorithms are used to identify various modal data streams included in the interaction data. The identified modal data streams are then processed using unimodal or multimodal methods to identify audience emotions and reactions. Various atmosphere factions are identified by clustering audiences expressing similar atmospheres (i.e., emotions) identified from keywords. Audience clusters can be further refined or tuned based on age, demographics, and other user attributes. User attributes can be obtained from user profiles maintained at the game cloud system.

[0012] Expressive avatars representing different moods (i.e., emotions) are generated for different mood factions. These avatars are generated to include different characteristics, thus providing a visual representation of the dominance of each emotion within the audience (i.e., the crowd). Some characteristics include the magnitude of the dominance of each emotion within the audience, colors reflecting the audience's emotions, such as the temperature and intensity of the hue used to invoke group psychology. The colors used to represent different emotions can be selected by referring to brand color psychology publications / literature available to the system when generating the avatars. An exemplary reference can be found at https: / / en.wikipedia.org / wiki / Color_psychology. By using machine learning, the emotions of a large audience are identified, aggregated, and represented in a scalable manner to provide the audience with visual representations of the different emotions of the audience and the number of audience members expressing different emotions. Avatars expressing various emotions are provided for rendering an image representation of the audience presented alongside video game content from an online game session. One advantage is that the avatars provide the audience with a way to quickly (i.e., almost in real-time) visually measure the distribution of reactions of a large number of game viewers and allow game viewers to compare how their own reactions compare to those of a peer group of viewers. Another advantage is that viewers can quickly identify and connect with like-minded individuals, allowing them to fully enjoy the gameplay from any position. Yet another advantage is that avatars allow online video game players to measure feedback on specific gameplay elements.

[0013] In addition to providing avatars to express various emotions, audience emotions can also be used to identify reaction tracks, which are then incorporated into the avatars presented to the audience during gameplay, whether viewed live or later. Reaction tracks are used to express the emotions of the audience. For example, one type of recognizable reaction track is the laughter track (also called the laughter sound track). Laughter tracks are provided to express happy emotions. Besides providing laughter tracks to express happy emotions, reaction tracks can also be used to express other emotions such as sadness, surprise, anger, neutrality, etc. Appropriate reaction tracks are identified based on the emotions expressed by different avatars within each atmosphere faction identified in the audience watching the video game. In some implementations, reaction tracks are audio tracks that capture the atmosphere of the audience.

[0014] Instead of avatars, or other than avatars, audience interaction data can be used to identify appropriate Graphical Interchange Format (GIF) files to represent different emotions expressed by the audience, and to provide the identified GIFs for rendering alongside the video game content. In one example, the GIF used to represent the emotion of each group can be automatically selected based on the preferences of the audience in a group, or the previous selection of a GIF by one or more audience members in the group, or the popularity of the GIF, etc. In another example, the GIF used to represent the emotion of the group can be automatically selected, and the audience can be provided with options that override the selection. The options can be provided on the user interface along with a subset of GIFs for the audience to choose from. The subset of GIFs presented on the user interface can be based on the type of GIF used to express a specific emotion that the audience previously selected in the video game (i.e., in the interactive interface when the interaction data is provided) or in social interactions, and can be identified from audience preferences maintained in each audience's user profile or in the interaction history. The interaction history can be maintained for each video game, each audience member, each audience group, each interaction session, each emotion, etc.

[0015] In other examples, options can be provided on the user interface for viewers to select their own GIFs to represent the emotions of the group in which the viewer is a member. In this example, links to one or more GIFs can be provided to the viewer for selection. In another example, instead of automatically selecting GIFs and providing options to the viewer, a subset of GIFs representing the emotions associated with each group can be identified and provided on the user interface for one or more viewers of the corresponding group to select. The viewer's selection of GIFs from the subset can be used to represent the group's emotions and provided to the viewer's client device to be rendered alongside the video game content. In an alternative example, each viewer of the group can be allowed to select their chosen GIF to be rendered alongside the video game content presented on their respective client device. In this example, each viewer has the freedom to control the rendering of the GIF on their own client device.

[0016] In one implementation, a method is provided for representing the emotions of an audience watching an online game video game. The method includes capturing interaction data from viewers participating in the gameplay of the video game. The interaction data captured from the viewers is aggregated. The aggregation includes clustering viewers into different groups based on emotions detected from the viewers within the audience. Each viewer group is associated with a unique emotion and a confidence score corresponding to the number of viewers in the corresponding group expressing the unique emotion. An avatar is generated to represent the emotion of each group, wherein the avatar for each group provides a visual representation of the reaction of the corresponding viewer group. The facial expressions of the avatars associated with each group are dynamically adjusted to match the facial expression changes of the viewers in the corresponding group. The avatars representing the unique emotions of the different viewer groups within the audience are presented alongside the content of the video game.

[0017] In one implementation, one or more modal data streams included in the interaction data are identified. The one or more modal data streams are processed to identify the emotions expressed by viewers watching the video game. Viewers are clustered into groups based on the emotions expressed, such that each viewer group is associated with a unique emotion. The one or more modal data streams identified from the interaction data correspond to any or a combination of text data, video data, audio data, chat data, emoticons, emojis, graphic content, or Graphics Exchange Format files collected in real time from viewers watching gameplay of the video game. Video data captures the facial expressions of different viewers while they are watching the video game, while audio data includes audio content and one or more audio features, such as pitch, amplitude, or duration.

[0018] In one implementation, multiple models are generated and trained using machine learning algorithms. Each of the multiple models is trained using data from a specific modality data stream identified from the interaction data. The outputs of the multiple models are aggregated to classify the audience's emotions and determine the probability of each emotion expressed by the audience via the interaction data.

[0019] In one implementation, a machine learning algorithm is used to generate and train the model. The model is trained using a modal data stream identified from the interaction data as input. The model's output is used to classify the audience's emotions and determine the probability of each emotion expressed by the audience via the interaction data.

[0020] In one implementation, the size of the incarnation for each unique emotion is scaled based on a confidence score associated with the corresponding audience group.

[0021] In one implementation, the confidence score associated with each audience group varies depending on the number of viewers detected in the corresponding audience group.

[0022] In some implementations, aggregating interaction data involves generating a word cloud and dynamically updating it using keywords identified through sentiment analysis of the interaction data. Keywords are updated to the word cloud to capture the audience's emotional state at each point in time. Keywords from the word cloud are used as input to machine learning algorithms to generate and train one or more models. The output from one or more models is used to identify the emotions expressed by the audience and the probability of each emotion via the interaction data.

[0023] In one implementation, the interactive data includes one or more of video data, audio data, text comments, or emoji responses provided via the interactive interface. The interactive data is generated or captured in real time by the viewer while watching the video game.

[0024] In one implementation, the facial expressions of the audience in each group change according to changes that occur during gameplay of the video game, and the facial expressions of each avatar associated with the corresponding group are dynamically adjusted to reflect changes detected in the facial expressions of the audience in the corresponding audience group.

[0025] In one implementation, the confidence score associated with each audience group varies depending on the number of viewers detected in the corresponding audience group.

[0026] In one implementation, an interactive timemap is generated and presented to represent the intensity of responses to different emotions detected from different audience groups in a video game. The intensity of the response to each emotion within the interactive timemap varies over time according to changes occurring within the gameplay of the video game. Changes in response intensity captured in the interactive timemap are linked to specific parts of the video game's gameplay that cause these changes in response intensity for the corresponding emotions detected from the respective audience groups. These links allow access to the specific parts of the video game's gameplay to view the interactions that lead to the corresponding changes in response intensity.

[0027] In one implementation, the interactive timeline is a line graph comprising multiple lines, each corresponding to a specific emotion detected from a particular audience group. An avatar corresponding to that specific emotion is rendered on the corresponding line.

[0028] In one implementation, gameplay data for a video game is streamed in real time, and the gameplay recording is stored in a gameplay data storage device and can be used for subsequent streaming at a delayed time. The delayed time used in this application corresponds to rendering the gameplay recording when the video game is replayed at a later time, rather than real-time streaming. An interactive timemap representing the intensity of emotional responses associated with the video game's audience is generated and presented along with the video game content. The interactive timemap is generated in real time and stored in the gameplay data storage device along with the gameplay recording. When the gameplay recording is subsequently streamed at the delayed time, a new interactive timemap is generated by modifying the interactive timemap to include the intensity of responses from multiple viewers captured while the gameplay recording was streamed during the delayed time. The new interactive timemap capturing the intensity of responses from multiple viewers is presented during the playback of the gameplay recording at the delayed time.

[0029] In one implementation, the presented avatar includes a user interface providing viewers with segmentation and formatting options. The segmentation options provide choices for selecting a segment from a plurality of segments defined on the display screen to render the avatar, and the formatting options provide rendering options to be used when rendering the avatar on the display screen. The formatting options include one of a transparency format, an overlay format, or a rendering format.

[0030] In one implementation, the presentation of the avatar includes determining the geographic locations of viewers within each audience group. When viewers in each group are associated with a single geographic location and each audience group is associated with distinctly different geographic locations, a map identifying the geographic locations associated with the different audience groups is presented. The corresponding avatar associated with each respective audience group is overlaid on the geographic location identified on the map as associated with that audience group.

[0031] In one implementation, a specific viewer within each group is identified, the viewer's reaction is captured during a defined game moment in the video game, and the captured viewer's reaction is rendered alongside the video game content.

[0032] In one implementation, a specific viewer within each group is identified by: identifying planned actions in the video game based on the game state; identifying the types of reactions exhibited by different viewers within each group to different actions occurring in the video game; and selecting a specific viewer within each group based on the specific viewer's reactions to different actions, the selection of a specific viewer including predictive amplification to capture the specific viewer's reaction in each group when the action occurs in the video game.

[0033] In one implementation, a specific viewer within each group is selected based on the type and number of comments generated by the remaining viewers in the respective group that relate to the expression of the specific viewer in each group, either randomly or based on expressive responses provided by the specific viewer within the group.

[0034] In one implementation, the interaction data captured from the audience includes reactions to events or actions occurring in the video game, or inverse reactions to the reactions of a specific audience member watching the gameplay of the video game.

[0035] In one implementation, the interaction data captured from the audience of the group includes responses from specific audience members associated with said group. Aggregating the interaction data captured from the audience includes aggregating responses from other audience members in the group to the responses of the specific audience members.

[0036] In one implementation, clustering viewers into different groups includes providing viewers with the option to move from a first group to a second group. This option is provided on a user interface rendered alongside the video game content. The movement causes the viewer to dynamically unassociate from the first group and dynamically associate with the second group. This dynamic association allows the viewer to access interactions from viewers in the second group, while the dynamic unassociation prevents the viewer from accessing interactions from viewers in the first group.

[0037] In an alternative implementation, a method is provided for representing the sentiment of an audience of viewers watching an online game video game. The method includes capturing interaction data from viewers participating in the gameplay of the video game. The interaction data captured from the audience is aggregated, and sentiment analysis of the interaction data is performed. The aggregation includes identifying one or more modal data streams included in the interaction data, processing the one or more modal data streams to identify the sentiment expressed by viewers watching the gameplay of the video game, and clustering viewers into groups based on the sentiment expressed. Each viewer group is associated with a unique sentiment and a confidence score corresponding to the number of viewers in the corresponding group expressing the unique sentiment. An avatar is generated to represent the unique sentiment of each group. The facial expressions of the avatars associated with each group are dynamically adjusted to match the facial expression changes of the viewers in the corresponding group. The avatars representing the unique sentiments of different viewer groups are presented alongside the content of the video game.

[0038] In one embodiment, a method for representing the emotions of an audience watching an online game video game is disclosed. The method includes aggregating interaction data collected from viewers participating in watching gameplay of the video game. The aggregation includes clustering viewers into different groups based on emotions detected from the viewers, such that each viewer group is associated with a unique emotion identified from the interaction data. Reaction audio tracks are identified to correspond to the unique emotions associated with each group. The reaction audio track for each group is presented alongside the content of the video game being watched by the viewer.

[0039] In one implementation, the confidence score for each audience group is determined based on the number of audience members in each group. The volume of the response track for each group is then scaled based on the confidence score of that group.

[0040] In one implementation, changes in the emotions expressed by the audience of each group are detected over time, and the reaction audio tracks of each group are dynamically identified to correspond to changes in the emotions detected in the corresponding group, wherein the changes are detected to be related to changes appearing in the content of the video game.

[0041] In one implementation, the reaction audio tracks for each group are presented to all viewers participating in watching the video game.

[0042] In one implementation, the reaction audio track for each audience group is presented to the audience of the corresponding group, such that the reaction audio track presented to the first audience group is different from the reaction audio track presented to the second audience group.

[0043] In one implementation, a reaction audio track for a specific audience group is presented to all viewers watching the video game, where the specific group is selected based on a confidence score for each group. The confidence score for each group indicates the number of viewers in that group.

[0044] In one implementation, sentiment analysis of the interaction data is performed to identify keywords corresponding to emotions, and a word cloud is generated and dynamically updated with keywords. The keywords in the word cloud capture the audience's emotional state at any given point in time.

[0045] In one implementation, response tracks for a specific audience group are identified by grouping keywords in a word cloud according to emotions defined by the keywords, such that each group of keywords corresponds to a specific emotion. The selected group of keywords representing the emotions of a specific group is used to identify the response tracks for that specific group.

[0046] In one implementation, the affiliation between a viewer and a specific player or team in a video game is identified from the viewer's or player's social graph. A word cloud is generated and updated based on keywords in the interaction data. These keywords correspond to the viewer's emotional response to a specific player or team's gameplay in the video game. The keywords are grouped into multiple groups based on the viewer's affiliation, with each group corresponding to a specific emotion. Selected groups of keywords representing the emotions of a specific group are used to identify the group's reaction audio tracks.

[0047] In one implementation, an artificial intelligence model is used to identify the response audio tracks for each group.

[0048] In one implementation, a request is received from a first group of viewers. The request is to reassign the viewer to a second group instead of the first. In response to the request, the viewer is unassociated from the first group and associated with the second group. This association results in a change in the number of viewers in both groups and provides the viewer with access to interactive data from the second group.

[0049] In one implementation, avatars are generated and associated with each audience group. The facial expressions of each group's avatar are dynamically adjusted based on changes detected in corresponding emotions expressed by the audience of each group. These changes in audience-expressed emotions correspond to events occurring in the video game. Each group's avatar is presented alongside the video game content along with a reaction soundtrack.

[0050] In some implementations, an artificial intelligence model is used to generate a generic avatar for a group of viewers. This generic avatar is associated with the group. This association allows access to interactive data generated by the group's audience. The interactive data is used to determine the dynamics of emotions expressed by the group's audience, where emotions change in response to changes detected in the video game. The generic avatar's facial expressions are dynamically adjusted, and appropriate reaction audio tracks are identified to correspond to changes detected in the group's emotions. The avatar with the adjusted facial expressions and the reaction audio tracks are then forwarded to the viewer's client device for rendering.

[0051] In some implementations, avatars and reaction soundtracks of two opposing groups are presented to the audience simultaneously alongside the video game content. These opposing audience groups provide contrasting reactions to events occurring within the video game.

[0052] In one implementation, the audience is selected from a specific group of viewers associated with a particular emotion in order to provide salient expression of that emotion. The audience is selected based on the emotions expressed by the viewers. Videos of the audience expressing collective emotions during key gameplay moments related to events occurring in the video game are presented. These videos are presented along with a soundtrack of the specific group's reactions.

[0053] In one implementation, the audience of a particular group is selected randomly or through predictive analysis of the audience’s previous facial expressions captured while the audience in the particular group is watching gameplay of a video game.

[0054] In one implementation, a Graphical Interchange Format (GIF) image is identified to provide reaction salience for specific emotions associated with a particular group. The GIF is identified using keywords associated with the emotions of the particular group. The identified GIF is presented to provide reaction salience during key gameplay moments related to events occurring in a video game. The GIF expresses reaction salience based on emotions expressed by an audience of the particular group.

[0055] In another embodiment, a method for representing the emotions of an audience watching an online game video game is disclosed. The method includes capturing interaction data from viewers participating in the gameplay of the video game. The interaction data captured from the viewers is aggregated. The aggregation includes clustering viewers in the audience into different groups based on emotions detected from the viewers within the audience. Each viewer group is associated with a unique emotion identified from the interaction data, and a confidence score is calculated for each group. A confidence score is determined for each viewer group based on the number of viewers in each group expressing a unique emotion associated with said group. Reaction audio tracks are identified to correspond to unique emotions associated with a specific viewer group identified within the audience, wherein the specific viewer group is identified based on the confidence score of each group. The reaction audio tracks for the specific viewer group are presented to each group of viewers alongside the content of the video game being watched.

[0056] In one implementation, a method is disclosed for representing the emotions of an audience watching gameplay of a video game. The method includes aggregating interaction data collected from viewers participating in watching gameplay of the video game. The aggregation includes clustering viewers into different groups based on the emotions expressed by the viewers. Each viewer group is associated with a unique emotion identified from the interaction data. A Graphical Interchange Format (GIF) file is identified for each unique emotion expressed within the viewer group. The identified GIF for each unique emotion is associated with the corresponding viewer group, such that each viewer group is associated with a unique GIF. The identified GIF, along with the gameplay content of the video game, is returned to the viewer's client device for rendering.

[0057] In one implementation, changes in the emotions expressed by the audience of each group are detected, and GIFs identified for said groups are dynamically updated. These emotional changes are correlated with changes occurring in the gameplay of the video game, and the group's GIFs are dynamically updated to correspond to the emotional changes of the group's audience.

[0058] In one implementation, the identified GIF for each group is returned to the video game's viewer's client device for rendering on the viewer's image representation. The viewer's image representation is configured to be rendered alongside the video game content.

[0059] In one implementation, the size of the GIF associated with each group is scaled to match the number of viewers in the corresponding group, such that the GIF identified for the first group with the highest number of viewers is presented as larger than that for the second group with fewer viewers than the first group. The size of each GIF is scaled to correspond to the number of viewers in each audience group.

[0060] In one implementation, aggregating the interaction data includes: identifying modal data streams included in the interaction data; processing the modal data streams to identify emotions expressed by viewers watching the video game; and clustering viewers into groups based on the emotions expressed, where each group is associated with a unique emotion. The modal data streams identified from the interaction data correspond to any or a combination of text data, video data, audio data, chat data, emoticons, emojis, or graphic content collected in real-time from viewers watching the video game.

[0061] In one implementation, multiple models are generated and trained using machine learning algorithms. Each of the multiple models is trained using data from a specific modality data stream identified from the interaction data. The outputs of the multiple models are aggregated to identify the emotions expressed by the audience and the probability of each emotion via the interaction data.

[0062] In an alternative implementation, a machine learning algorithm is used to generate and train the model. The model is trained using a modal data stream identified from the interaction data as input. The model's output is used to identify the emotions expressed by the audience and the probability of each emotion via the interaction data.

[0063] In one implementation, identifying GIFs for a specific group includes identifying a subset of GIFs for unique emotions associated with that specific audience group, and presenting said subset of GIFs on a user interface for selection by one or more viewers within that specific group. The subset of GIFs is selected based on previous selections of GIFs for unique emotions by one or more viewers within that specific group.

[0064] In one implementation, associating a GIF with a unique emotion for a specific group includes receiving a specific GIF selected from a subset of GIFs presented on the user interface, and associating the selected GIF with the specific group.

[0065] In one implementation, a subset of GIFs is identified for unique emotions associated with a specific audience group. This subset is identified based on preferences specified by one or more viewers within the specific group. Each GIF in the subset is associated with a confidence indicator representing the number of times a GIF was selected by one or more viewers within the specific group for a unique emotion. Based on the confidence indicator associated with a specific GIF, a specific GIF is automatically selected from the subset to be associated with the specific group.

[0066] In one implementation, a custom option is provided to override the automatically selected GIFs for a specific group. Choosing the custom option results in rendering a subset of GIFs with unique emotion recognition specific to the group, and a selection option to choose alternative GIFs from the subset to associate with the specific group.

[0067] In one implementation, a GIF identified for the unique emotion of each group is formatted and rendered in a segment defined on the display screen of the client device. This segment is identified based on preferences specified by each viewer within each group.

[0068] In one implementation, a rendering format is defined for each identified GIF, wherein the rendering format is one of a transparency format, an overlay format, or a rendering format.

[0069] In one implementation, a unique emotion recognition response audio track is used for each group. The recognized response audio track is returned along with a GIF for rendering on the viewer's client device.

[0070] In one implementation, an identified GIF for each group is provided to be rendered to the audience of said group, such that a GIF corresponding to the emotions of the respective group is presented to each audience group.

[0071] In one implementation, an option is provided that allows viewers to move from a first group to a second group. This option is presented on the interface alongside a list of groups created based on aggregated interaction data. Selecting the option to identify the second group causes the viewer to unassociate from the first group and associate with the second group. Unassociating from the first group prevents the viewer from accessing the first group's interaction data, while associating with the second group provides the viewer with access to the second group's interaction data.

[0072] In one implementation, a word cloud is generated and dynamically updated using keywords identified through sentiment analysis of the interaction data. The keywords updated to the word cloud correspond to the unique emotions expressed by the audience at each point in time.

[0073] In one implementation, keywords in the word cloud are grouped according to the emotions defined by the keywords, with each group corresponding to a different emotion. A group of keywords corresponding to a specific emotion is used to identify GIFs targeting that specific emotion. These GIFs are then associated with a corresponding audience group that provides interactive data from which the keywords targeting that specific emotion were identified.

[0074] In one implementation, the audience cluster is further based on the audience's age, demographics, affiliation with players, affiliation with teams, and user profiles.

[0075] In another embodiment, a method for representing the emotions of an audience of viewers watching an online video game is disclosed. The method includes aggregating interaction data detected from viewers engaging in gameplay of the video game. The aggregation includes identifying one or more modal data streams included in the interaction data. The one or more modal data streams are processed to identify the emotions expressed by viewers watching gameplay of the video game. Viewers are clustered into groups based on the emotions expressed, with each viewer group associated with a unique emotion identified from the interaction data. Graphical Interchange Format (GIF) files are identified for the unique emotions expressed in each viewer group. The GIF identified for each unique emotion is associated with the corresponding viewer group, such that each viewer group is associated with a unique GIF. The GIF identified for each group, along with the content of the video game, is returned to the viewer's client device for rendering.

[0076] In one implementation, clustering the audience involves generating and training one or more models using machine learning algorithms. The one or more models are trained using data from one or more modal data streams identified from the interaction data. The outputs of the one or more models are aggregated to identify different emotions expressed by the audience and the probability of each emotion.

[0077] Other aspects and advantages of this disclosure will become apparent from the following detailed description taken in conjunction with the accompanying drawings, which illustrate the principles of this disclosure by way of example. Attached Figure Description

[0078] This disclosure can be better understood by referring to the following description taken in conjunction with the accompanying drawings, in which:

[0079] Figure 1A simplified block diagram of a system according to an implementation of the present disclosure is shown, the system being configured to execute gameplay of a video game for multiple players and to identify and present emotions detected from different audience groups watching the gameplay of the video game.

[0080] Figure 2 A simplified overview of the different stages of emotion of different viewers within an audience watching gameplay of a video game is illustrated according to an implementation of this disclosure.

[0081] Figure 3 A block diagram is shown of an emotion expression engine, according to an implementation of the present disclosure, for recognizing emotions detected from different viewers and presenting visual representations of the emotions detected from different viewers.

[0082] Figure 4 This illustration provides a broad overview of the different components of an emotion expression engine used during different stages of processing interactive data to generate and present incarnations of different emotions, according to one implementation of this disclosure.

[0083] Figure 5 An exemplary interaction analyzer for collecting and analyzing interaction data of viewers of online games watching video games, according to one implementation of this disclosure, is illustrated.

[0084] Figure 6 An exemplary keyword analysis engine is illustrated, according to one implementation of the present disclosure, for identifying various emotions detected from different viewers.

[0085] Figure 7 An exemplary avatar visualizer is illustrated, according to one implementation of the present disclosure, for generating and scaling avatars representing different emotions detected from different viewers.

[0086] Figure 8A An overview of video image processing as part of a single-modal emotion recognition process is illustrated according to one implementation of this disclosure.

[0087] Figure 8B An overview of video image processing as part of a multimodal emotion recognition process is illustrated according to one implementation of this disclosure.

[0088] Figure 8C A simplified monomodal emotion recognition process is illustrated, implemented according to one implementation of the present disclosure, using an emotion expression engine for recognizing different facial expressions detected from the interactions of viewers watching online games.

[0089] Figure 9A simplified screen representation of the interactive collection phase (i.e., operation) of video from different viewers captured while a viewer is watching an online game of a video game, according to one implementation of this disclosure, is shown.

[0090] Figure 10 The illustration depicts a sample facial feature recognition process for identifying different facial expressions detected from an audience, according to one implementation of the present disclosure.

[0091] Figure 11 The illustration depicts a sample set of emotions identified by analyzing facial features captured in a video game while the viewer is watching an online video game, according to one implementation of this disclosure.

[0092] Figure 12 The illustration depicts a sample atmosphere rating used by an emotion expression engine to identify audience emotions during emotion analysis, according to one implementation of this disclosure.

[0093] Figure 13 A simplified screen view representation of the emotional aggregation phase (i.e., operation) of audience interaction collected during online gameplay of a video game, according to one implementation of this disclosure.

[0094] Figure 14 A simplified screen view representation of the emotion visualization phase performed by an emotion expression engine according to one implementation of this disclosure is illustrated.

[0095] Figure 15 The illustration depicts various interactive inputs from the audience collected from the audience of the audience for generating a word cloud, according to one implementation of the present disclosure, the word cloud being used to identify keywords representing emotions detected from the audience.

[0096] Figure 16 The illustration depicts a sample atmosphere faction defined by an emotion expression engine according to one implementation of this disclosure.

[0097] Figure 17 The illustration depicts, according to one implementation of this disclosure, sample avatars / emoticons representing different mood factions presented to the audience.

[0098] Figure 18 The illustration depicts a method operation for identifying and presenting the emotions of the audience and displaying an avatar representing the expressed emotions next to the content of a video game, according to one implementation of the present disclosure.

[0099] Figure 19 The illustration depicts variations in different components of an emotional expression engine for performing different processing stages of interactive data according to one implementation of the present disclosure to generate and present avatars representing different emotions and to identify appropriate response audio tracks for presentation at the client device.

[0100] Figure 20 An example of a reaction audio track, based on one implementation, is shown that identifies different emotions expressed by viewers of an online game watching a video game.

[0101] Figure 21 An example of a reactive audio track is illustrated according to one implementation, which is identified and presented on a representative image of the audience along with an expressive avatar.

[0102] Figure 22 The illustration depicts various atmosphere factions identified among the audience according to one implementation method, and examples of corresponding response tracks associated with the respective atmosphere factions.

[0103] Figure 23 An alternative example is illustrated, showing an image representation of the audience in one implementation, along with an embodiment of the dominant emotion identified in the audience and a reaction audio track of the most dominant emotion rendered next to the corresponding embodiment.

[0104] Figure 24 The illustration depicts a method operation for recognizing a viewer's emotions and presenting a corresponding reaction audio track alongside the content of a video game, according to one implementation of the present disclosure.

[0105] Figure 25 The illustration depicts variations of different components of an emotional expression engine, according to one implementation of the present disclosure, for processing interactive data detected from viewers to identify Graphical Interchange Format (GIF) files so as to be presented on a client device along with video game content.

[0106] Figure 26 The illustration depicts segments of a GIF, identified on a client device's display screen according to one implementation method, used to render different emotions expressed by the viewer.

[0107] Figure 27 This illustration shows a subset of exemplary GIFs that identify different emotions detected from audience interaction data, based on one implementation method.

[0108] Figure 28 An exemplary view is illustrated, showing an audience image according to one implementation, a subset of GIFs identified for a specific emotion rendered on the viewer's client device, and selection options for selecting a specific GIF from the subset.

[0109] Figure 29 The illustration depicts different emotions being identified and rendered on the display screen of a client device according to one implementation method, as well as customization options for changing the GIF for each rendered emotion.

[0110] Figure 30 The illustration depicts a method for identifying audience emotions and presenting a graphical exchange format file representing the identified emotions alongside the content of a video game, according to one implementation of this disclosure.

[0111] Figure 31 An exemplary information service provider is illustrated, according to one implementation of this disclosure, for processing audience interaction data to present avatars representing different emotions.

[0112] Figure 32 Components of an exemplary server apparatus are illustrated that can be used to carry out various embodiments of the present disclosure. Detailed Implementation

[0113] The following implementations of this disclosure describe methods and systems for generating expressive avatars for a group of viewers watching an online video game and displaying these expressive avatars alongside the content of the video game. The avatars' expressions represent different emotions detected by viewers of the online video game (i.e., gameplay), allowing viewers to quickly measure the distribution of different atmospheres among a large number of viewers. The avatars also allow each viewer to identify peer groups with whom they can associate by comparing their own reactions to those of different viewer groups within the audience. This association allows viewers to “go play” with other viewers expressing similar emotions (i.e., atmosphere), giving the impression that the viewer is playing with their friends or family gathered in a living room, public space, or stadium watching the game. This disclosure allows viewers to experience the emotions of the audience and associate with viewers who experience the emotions they feel, thereby allowing viewers to have a richer gaming viewing experience via digital viewing.

[0114] One of the main drawbacks of conventional digital viewing for viewers is the lack of other viewers to share the gaming experience. Online video games (e.g., esports) allow viewers to watch the online gameplay (i.e., the game) from anywhere and become part of the viewing audience. Viewers connect to online video games from their living rooms, dorm rooms, or any other preferred gathering place to watch the gameplay. However, for viewers watching the gameplay, online viewing lacks the connection with other viewers who would typically be present for a live match in a stadium. This drawback can be overcome by inviting other viewers to a venue (e.g., a sports bar) to watch the online game. Gathering other viewers requires planning, is time-consuming, and requires other viewers to be available, able, and willing to go to the venue during the specified time period. Even when viewers gather at a chosen venue, those typically gathered at that venue may all express the same emotions as those arranged to gather. Viewers at the venue may not fully understand the various emotions detected in the audience, as they have already gathered inside a stadium to watch a live sporting event.

[0115] To mitigate inconvenience and overcome the drawbacks of conventional online viewing, the various implementations described herein provide ways to visually represent (i.e., render) the emotions of different viewers watching online video games from different geographical locations. Especially for the highly popular field of esports, expressive avatars are used to convey the emotions of a large audience. These avatars allow a first viewer to quickly (i.e., almost in real-time) measure the distribution of reactions from a large number of game viewers and compare their reactions to those expressed by an equivalent group. The avatars also allow online game players to measure feedback on specific gameplay mechanics. To provide expressive avatars, an emotional expression engine is used to perform the various steps. Some of these steps include information gathering steps or stages, information aggregation steps or stages, and explanatory steps or stages.

[0116] As used in this application, a viewer is an individual (i.e., a person or user) who watches online events, performances, games, activities, etc. An audience refers to a group or collection of viewers who have gathered to watch online events (public or private events) such as plays, movies, concerts, conferences, games, etc. In the various implementations discussed in this application, a viewer is part of an audience that has gathered to listen to and / or watch gameplay of a video game. A viewer may participate in listening to the commentary of gameplay, or participate in watching gameplay, or both. In this application, emotion and sentiment are used interchangeably to refer to the behavior of a person (e.g., a viewer) conveyed through human interaction or through human facial expressions. Interaction is generally via speech (i.e., verbal) or writing (e.g., text or graphic content, including text comments, Graphics Interchange Format images / files, emoticons, emojis, graphic content, etc.). Facial features are generally used to provide expressions.

[0117] During the information gathering phase, live information from viewers is collected while they are watching gameplay of the video game. Live information may include live video of viewers' faces while watching the game, live audio of them interacting with players and / or other viewers, and text or chat comments, emoticons, emojis, or Graphics Exchange format images or files (GIFs) posted on interactive interfaces such as chat interfaces, message boards, etc. Information collected from each viewer is shared voluntarily based on the respective viewer's choice. For example, a sharing option may be provided to each viewer at a user interface presented alongside the video game content, wherein the sharing option identifies specific interactive data generated by the viewer that the viewer allows the system to collect and use to determine the viewer's mood. For example, viewers may generate audio data, text data, chat data, etc., and viewers may generate reactions to events or actions occurring in the video game. Viewers may allow the system to collect only chat data, only text data, only audio data, or only reactions, or any combination thereof. Alternatively, viewers may allow the system to collect any or a combination of interactive data generated or expressed by the viewer for a specific part of the video game or a specific session of the video game. The type and amount of live information collected from each audience member is based on the sharing options chosen by the audience member.

[0118] Information collected during the information gathering phase is aggregated and analyzed in real time. For example, live video feeds capturing images of the audience's faces are used to perform real-time emotion recognition. Machine learning algorithms are used to identify various modal data streams included in the live video feed and process these modal data streams to identify various emotions from the live feed. For example, machine learning algorithms can be used to identify the emotions expressed by the audience from facial features captured in the live feed. Similarly, audio and text comments undergo sentiment analysis to identify significant emotions and thoughts conveyed by the audience, and similarly, machine learning algorithms are used to analyze emojis, GIFs, and emoticons to identify the emotions expressed by the audience. The results of the analysis of modal data streams such as text data, audio data, and chat data are used to create dynamic word clouds, which are updated using keywords corresponding to different emotions. The resulting word clouds capture the audience's emotional state at different points in time. Machine learning algorithms are further used to cluster the audience and identify atmosphere factions within the audience by identifying and grouping members who express similar atmospheres. Atmosphere factions help players and viewers feel the energy of the audience's members and form stronger connections with specific members within an atmosphere faction than with other atmosphere factions.

[0119] The analysis results are used to visualize the mood of each atmosphere faction and generate expressive avatars to represent the mood of each atmosphere faction. Various characteristics of each atmosphere faction's avatar are dynamically adjusted to change color, size, and expression to reflect the current mood of the corresponding atmosphere faction. In some implementations, expressions in specific avatars within the avatars are highlighted to provide reaction salience. In alternative implementations, the reactions of selected viewers within each atmosphere faction or from a very specific atmosphere faction can be included as reaction salience along with the expressive avatars of different atmosphere factions. The expressive avatars, and in some cases, the reaction saliences of selected viewers within the audience, are rendered alongside or overlaid on the content of the video game rendered on each of the viewer's client devices.

[0120] The emotion expression engine is configured to represent a wide variety of emotions (i.e., emotional states) experienced by the audience. Avatars are scaled to visualize the dominance level of corresponding emotions within the audience, with more dominant emotions rendered larger than less dominant ones. Using machine learning, the emotion expression engine can identify the emotions of a large audience distributed across a wide geographical area in a scalable manner and provide aggregated feedback to the audience in virtually real-time.

[0121] In light of the foregoing overview, specific implementations will be described with reference to several exemplary figures to facilitate understanding of the exemplary embodiments. However, it will be apparent to those skilled in the art that this disclosure may be practiced without some or all of the specific details presently described. In other instances, well-known process operations have not been described in detail so as not to unnecessarily obscure this disclosure.

[0122] Figure 1 An implementation scheme of an overall game cloud system 10 is illustrated, configured to execute one or more instances of a video game for multiple players 101 to play and handle player and spectator interactions generated during gameplay. Players 101a to 101n access instances of the game from multiple client devices. Selection of the video game and requests for gameplay are forwarded from the player 101's client device to a game cloud server 300 over a network 200 (such as the Internet). A game engine 302 on the game cloud server 300 uses user account information stored in a user account database (not shown) to verify the player, and upon successful verification, instantiates one or more instances of the video game on one or more game cloud servers 300. The game engine 302 may be a distributed game engine executing the video game on one or more game cloud servers 300, located within one or more data centers (not shown), which may be located in one geographical location or distributed across multiple geographical locations. Multiple spectators 102a-102m can access gameplay of the video game executed on one or more game cloud servers 300 to watch the online gameplay of the video game. Access to viewers 102a-102m may be restricted or open. If access is restricted, viewers 102a-102m are verified before being provided with access to the gameplay of the video game. Various instances of the game running on different game cloud servers 300 in one or more data centers are used together to provide broad access to players 101a-101n and viewers 102a-102m in a distributed and seamless manner.

[0123] The video game can be a multiplayer online game, where players 101a-101n can be individual players or part of different competing teams. In some embodiments, the video game can include two opposing teams, and players 101a-101n can be part of one of the two teams. In alternative embodiments, the video game can include more than two teams, and players 101a-101n can be part of any of the multiple teams. Player interactions in the video game are forwarded to the game engine 302 to affect the game state. In response to player interactions, updated gameplay data is returned to the client devices of players 101a-101n. The gameplay data is also maintained in the gameplay data storage device 332 for later retrieval. It should be noted that the gameplay data stored in the gameplay data storage device 332 is data from live gameplay and can be retrieved when other viewers select the video game for replay. The game engine encodes the gameplay data and forwards the gameplay data as an encoded video stream to the client devices of players 101a-101n. The client devices of players 101a-101n are configured to receive encoded video streams, decode video streams, and render frames of gameplay content on a display screen associated with the respective client device. The display screen may be part of the client device (e.g., the screen of a mobile device) or may be associated with the client device (e.g., a monitor, television, or other rendering surface). In some implementations, the client device of player 101a-101n may be any connected device with a screen and an internet connection.

[0124] In some implementations, during gameplay, viewers 102a-102m can select a video game to watch the gameplay of the video game played by players 101a-101n. Requests from viewers 102a-102m are transmitted via their respective client devices to the game cloud server 300 through network 200. In response to the requests from viewers 102a-102m, the game server 300 (i.e., the game cloud server 300) forwards gameplay data describing the current game state of the video game as an encoded video stream to the corresponding client devices of viewers 102a-102m. Viewers 102a-102m's client devices receive the encoded video stream, decode the video stream, and render frames of the gameplay data on the display screen associated with the corresponding client device. The client devices of viewers 102a-102m can be any computing device, such as a head-mounted display (HMD), a mobile or portable computing device, a desktop computing device, etc., and the display device can be a display screen associated with the HMD or other mobile computing device (e.g., a mobile phone screen, a tablet computing device, etc.), or it can be a separate display device or display surface communicatively connected to the client devices of viewers 102a-102m, such as a monitor, a television, a display screen, etc. Viewers 102a-102m constitute an audience 103 who are watching and / or listening to the gameplay of a video game played by multiple players 101a-101n.

[0125] Viewers 102a-102m can provide interactions related to video game rendering on the client devices of players 101a-101n and viewers 102a-102m. Interactions can be provided on user interfaces such as chat interfaces or message boards, social media interfaces, etc., in the form of text, emojis, emoticons, or GIFs rendered alongside the video game content. Interactions can also be provided as audio comments captured by a microphone or other audio capture device included in or associated with the client device of viewers 102a-102m. Interactions from viewers 102a-102m can also be in the form of facial expressions of viewers watching gameplay, which can be captured as live video using one or more cameras integrated into the client device of viewers 102a-102m (e.g., head-mounted display, smart glasses, mobile device, etc.) or from an external camera communicatively connected to the client device of viewers 102a-102m. Interactions from viewers in audience 103 are forwarded to the game cloud server 300.

[0126] The emotion expression engine 304 collects interactions from viewers 102a-102m within audience 103, analyzes these interactions to identify the emotions expressed by different viewers 102a-102m, groups viewers 102a-102m within audience 103 according to the expressed emotions, generates an avatar representing each viewer group, and adjusts the facial expressions of each generated avatar to match the emotions detected from the corresponding viewer group. Expressive avatars, such as expressive emojis, are forwarded to the client devices of viewers 102a-102m to overlay a representative image of viewers 102a-102m on the corresponding client devices of audience 103, which is rendered alongside the video game content. The rendered expressive avatars provide viewers with visual representations of the different emotions detected from watching gameplay of the video game.

[0127] Figure 2 This illustration provides a broad overview of the various interaction processing stages of an emotion expression engine 304, implemented according to one method, for processing interactions and generating representative avatars for an audience group 102 that is part of an audience 103 watching gameplay in a video game. The emotion expression engine 304 uses machine learning algorithms to perform various processing stages to identify emotions from the audience's interaction data and to describe those emotions in the form of avatars. The processing stages and the subsequent emotion expression engine 304 can be broadly categorized into three main stages. These three main stages include an interaction collection stage performed by an interaction collection engine 311, an emotion aggregation stage performed by an emotion aggregation engine 312, and an emotion visualization stage performed by an emotion visualization engine 313. The emotion aggregation stage encompasses emotion detection and audience clustering based on the detected emotions.

[0128] Figure 3 A set of sample operations performed by the emotion expression engine 304 at each stage in one implementation was identified. (See also...) Figure 2 and Figure 3 The interactions of audience 102 are captured at the corresponding client devices of multiple audiences 102 in the video game and forwarded as interaction data to the game cloud server 300 via network 200. The emotional expression engine 304, which runs on the game cloud server 300 and is operatively connected to the game engine, collects interaction data from the audience, processes the interaction data, and generates expressive avatars.

[0129] Starting from the interaction collection phase, the emotion expression engine 304 begins processing interaction data collected from the audience. During the interaction collection phase, the interaction collection engine 311 collects in real-time various interactions generated by the audience 102 while the audience is watching gameplay of the video game. Interaction data can take the form of audience reactions to events, actions, or activities occurring in the video game based on input provided by one or more players. Alternatively, reactions from the audience can be in response to interactions provided by other viewers or players in the video game. These interactions are categorized based on type. Different types of interactions that can be collected include live video capturing at least the audience's facial features, audio content of verbal interactions from the audience in response to actions / activities within the video game or as part of interactions with other viewers or players, text or graphic content (e.g., emoticons, GIFs, emojis, etc.), etc. The categorized interactions are stored in the emotion collection database (or simply the "emotion database") 334 and provided as input to the emotion aggregation engine 312 as part of the emotion aggregation phase.

[0130] Live video of the viewer can be captured using various image capture devices integrated into or communicatively coupled to the viewer's client device, such as image sensors, cameras, digital cameras, stereo cameras, etc. The live video of the viewer is used to identify facial features from which potential facial expressions can be inferred. Similarly, audio content can be captured using a microphone or other audio capture device of the client device or communicatively coupled to the viewer's client device. Text or graphic content can be obtained from a chat interface, messaging interface, or social media interface rendered alongside the video game content. It should be understood that the interaction data is collected from the viewer based on their selection of options for sharing certain modal content generated while watching gameplay of the video game. For example, a viewer may explicitly choose to share their chat content instead of their video or audio content. Therefore, in some implementations, a user interface (not shown) with selection options for sharing different modal content generated by the viewer can be provided for the viewer to choose. The viewer can select not to share, share one, some, or all of the different modal content identified from the interaction data by selecting appropriate selection options. Based on the selection options chosen by each viewer, the emotional expression engine 304 collects corresponding modal data streams from the viewer's interaction data for processing.

[0131] During the emotion aggregation phase, the emotion aggregation engine 312 collects interactions from viewer 102 in real time while viewers are watching the online gameplay of the video game (i.e., the gameplay), and analyzes these interactions to identify the emotions expressed by the viewers while watching the gameplay. The emotion aggregation engine 312 uses machine learning algorithms to identify emotions detected from the interaction data collected from the viewers (312a). The machine learning algorithms first identify various modal data streams included in the interactions collected from the viewers. The machine learning algorithms then process the modal data streams by employing a unimodal or multimodal approach. For example, video feeds from each viewer can be processed by machine learning algorithms to identify facial features and perform real-time emotion recognition. Similarly, audio and text content can be processed by performing sentiment analysis to identify keywords representing significant emotions and thoughts conveyed by the viewers.

[0132] Then, machine learning algorithms categorize the various emotions expressed by the audience through interaction to identify atmospheres. Audiences may express emotions to varying degrees. For example, audiences may express happiness to varying degrees, including different amplitudes of smiles, using different keywords (e.g., happy, ecstatic, joyful, delighted, awesome, etc.). The machine learning algorithm identifies the different degrees to which different audiences express each emotion and categorizes the various interactions into atmosphere buckets for each emotion accordingly. The different degrees of emotion in each atmosphere bucket are then aggregated (312b). The emotion aggregation engine 312 uses machine learning algorithms to create word clouds and dynamically updates them using keywords identified from sentiment analysis of textual and verbal interactions (312c). The dynamic word cloud at any given point in time in the video game provides a visual, textual representation of the audience's emotional state. Other graphical content, including Graphical Interchange Format files / images (GIFs), emojis, emoticons, etc., included in the audience's interaction data, is processed by machine learning algorithms in a manner similar to video feed processing to identify the emotions expressed by the audience. The machine learning algorithm uses the emotions identified from keywords in the word cloud, the audience’s video and graphic content (e.g., memes, GIFs, emoticons, other graphic content) to identify the atmosphere in the audience (step 312c), and groups (i.e., clusters) the audience members who express similar atmospheres into atmosphere factions (step 312d).

[0133] In one implementation, once an initial atmosphere faction is formed, viewers remain members of that faction until an explicit request to disassociate from the initial faction is received from the viewers. Therefore, when additional interactions are received from viewers of a particular atmosphere faction, these interactions are processed by different sub-components of the emotion aggregation engine 312 to identify additional emotions expressed by viewers of that particular atmosphere faction. The additional emotions associated with a particular atmosphere faction are dynamically updated to reflect the current emotions of the corresponding group of viewers for that particular atmosphere faction. These additional interactions may be reactions to certain actions or activities occurring in the gameplay of a video game, or reactions to certain interactions generated by players or viewers of a particular atmosphere faction or different atmosphere factions. Grouping viewers with similar atmospheres allows viewers in each group to be associated with specific players within that group and to correspond with other viewers expressing similar atmospheres. Similarly, this grouping allows players to drastically change the energy of their associated audience groups, which can be analogous to supporters at a live match in a stadium. Data associated with atmosphere factions is provided as input to the emotion visualization stage.

[0134] In the emotion visualization phase, the emotion visualization engine 313 generates an avatar for each atmosphere faction and adjusts the facial expressions of each avatar to reflect the emotions of each atmosphere faction. In addition to adjusting facial expressions, the emotion visualization engine 313 also adjusts one or more characteristics of the avatar, such as the avatar's size and color, to reflect the audience's emotions, where larger avatars are used to represent more dominant emotions, and smaller avatars are used to represent less dominant emotions expressed by viewers within the audience. Besides generating avatars, the emotion visualization engine 313 also identifies appropriate reaction tracks to associate with each atmosphere faction. The reaction tracks provide the voices of the audience expressing the atmosphere associated with each atmosphere faction. For example, atmospheres may include happiness, sadness, surprise, anger, neutrality, etc., and reaction tracks associated with these atmospheres captured from the live audience are identified and presented along with representative avatars. In one implementation, the reaction tracks for various atmospheres may be stored within the game cloud server 300 or in a reaction track data storage device (not shown) outside the game cloud server 300 and accessed by the emotion expression engine 304 executing within the game cloud server 300. Reaction tracks stored in the reaction track data storage device can include reaction tracks tailored to different emotions and organized according to context. For example, reaction tracks can capture different emotions expressed in response to different events and contexts. The emotion visualization engine 313 identifies appropriate reaction tracks for the emotions expressed in each atmosphere faction based on the background of a video game.

[0135] The adjusted avatar is returned to the corresponding client device of audience 102 for rendering on the image representation of audience 102's audience 103. The rendering of the expressive avatar provides a visual representation of the emotional distribution experienced by audience 102's audience 103 to both player 101 and audience 103. In some implementations, the avatar is in the form of an emoji. Using machine learning, the emotion expression engine 304 attempts to identify the emotions of a large number of audience members in a scalable manner and provide aggregated feedback in real time, allowing audiences in different groups to accurately measure the crowd emotions of audience 102.

[0136] In addition to returning avatars of the audience associated with different atmosphere factions, reaction tracks for various emotion recognitions associated with different atmosphere factions are also forwarded to the audience's client devices to be rendered along with the avatars. The reaction tracks provide the audience with an additional way to experience the atmosphere of the audience. The avatars provide a visual representation of the atmosphere, and the reaction tracks provide an auditory representation of the atmosphere. The audience can hear and feel the atmosphere of the audience, making them believe they are watching the game in an arena or stadium with a crowd of spectators, rather than watching it alone in their living room.

[0137] Figure 4 The illustration shows some components of the emotional expression engine 304 in one implementation. (See previous references.) Figure 2 and Figure 3 The aforementioned emotion expression engine 304 includes an interaction collection engine 311, an emotion aggregation engine 312, and an emotion visualization engine 313. Some components may include additional sub-components. For example, the emotion aggregation engine 312 may include an interaction analyzer 314, a keyword analysis engine 315, and a visual emotion analysis engine 325. Similarly, the emotion visualization engine 313 may include an avatar visualizer 316.

[0138] Interaction collection engine 311 collects interaction data generated by viewer 102 while the viewer participates in watching the online game, and processes the interaction data to identify different interaction patterns (i.e., types) included therein. The different interaction patterns that can be identified by interaction collection engine 311 can be broadly categorized into viewer live video 311a, viewer audio 311b, and chat comments 311c, which include audio and video components. While the viewer is watching the online game (i.e., gameplay), live video of the viewer is captured using one or more image capture devices oriented towards the viewer. The image capture devices may include one or more of a camera, stereo camera, digital camera, or any other image capture device associated with or available at the viewer's client device, wherein the client device may be a portable computing device, such as a laptop, smartphone, head-mounted display, smart glasses, wearable computing device, tablet computing device, etc., or a desktop computing device. Chat comments may include any and all types of content provided via a chat interface or message board, interactive social media interface, or any other interactive application interface. Chat comments can include text comments, videos or video clips, emojis, emoticons, GIFs, audio clips, etc., provided by viewers in response to events, activities or actions that occur during gameplay of the video game or in response to interactions from other viewers or players.

[0139] In one implementation, the emotion aggregation engine 312 processes interactions identified by the interaction collection engine 311 differently based on modality type. For example, an interaction analyzer 314 within the emotion aggregation engine 312 processes graphical content provided via a chat interface to identify emotions expressed via emoticons, emojis, GIFs, and other graphical content. A probability score is calculated for each emotion identified for the graphical content. The emotion and probability score of the graphical content are provided as input to the avatar visualizer 316.

[0140] The live video feed of the audience can be processed in part by the interaction analyzer 314 and in part by the visual emotion analysis engine 325. For example, the visual emotion analysis engine 325 processes audience images captured in the live video, and the interaction analyzer 314 processes verbal content included in the live video. The visual emotion analysis engine 325 analyzes facial images to identify facial features captured in the images, thereby identifying the emotions expressed by the audience. The visual emotion analysis engine 325 can use machine learning algorithms to identify the attributes of various facial features captured in the images and uses an expression recognition neural network to identify the emotions expressed by the audience. Since the expressions provided by the audience may correspond to more than one emotion (see...),... Figure 11Therefore, identifying the most dominant emotion helps to cluster the audience into appropriate atmosphere factions. Thus, the visual emotion analysis engine 325 calculates a probability score (also called an "emotion probability score") 325a for each emotion identified from the attributes of facial features. Based on the probability scores 325a for each emotion identified from the audience's expressions, the most dominant emotion expressed by the audience is identified. After identifying the emotion of each audience member, the audience is clustered into atmosphere factions, where each atmosphere faction corresponds to a unique emotion. Details of the audience in each atmosphere faction and the emotion associated with each atmosphere faction are provided as input to the avatar visualizer 316.

[0141] Interaction analyzer 314 can process the audio content of the live video to identify keywords included within. Similarly, the text portion of the chat content is processed by keyword analysis engine 315 to identify keywords and use these keywords to identify the emotions conveyed by the audience through the text input. Further details regarding the functionality of interaction analyzer 314 will be available in [reference needed]. Figure 5 The description is provided. Keywords identified by the interaction analyzer 314 through analyzing the text and audio content of the interaction serve as input to the keyword analysis engine 315.

[0142] The keyword analysis engine 315 uses keywords identified by the interaction analyzer 314 to populate a word cloud for identifying emotion-related keywords. The emotion-related keywords in the word cloud are used to identify different emotions expressed by viewers within the audience. As more viewers begin to express certain emotions through interaction, the keyword analysis engine 315 calculates probability scores for keywords corresponding to higher emotions in the word cloud. Various keywords from the word cloud and their corresponding calculated probability scores are provided as input to the emotion visualization engine 313. Emotions detected from various interactions are used to cluster viewers into atmosphere factions, where each atmosphere faction corresponds to a different atmosphere or emotion. Viewer clustering can be further refined or adjusted based on age, demographics, and other user attributes. User attributes can be obtained from user profiles maintained at the game cloud system. Atmosphere factions for different emotion recognitions are provided as input to the emotion visualization engine 313.

[0143] The emotion visualization engine 313 receives various inputs from the emotion aggregation engine 312 and uses these inputs to create an avatar for each atmosphere faction. As part of creating avatars to represent the emotions of each atmosphere faction, the emotion visualization engine 313 calculates a confidence score for each atmosphere faction as the number of viewers in the atmosphere faction who have expressed a dominant emotion or a comparable version of a dominant emotion. Based on the calculated confidence scores, the avatar visualizer 316 creates and scales avatars corresponding to the different atmosphere factions identified in the audience. Adjusting the avatars includes adjusting the avatar's expression, size, and color at least according to the confidence score of the corresponding atmosphere faction.

[0144] The size of each avatar is scaled to correlate with its confidence score, such that the avatar with the highest confidence score is rendered larger than the avatar with the lowest confidence score. Similarly, the color of an avatar can be adjusted to reflect the mood rating of the emotion being expressed. For example, anger might be rendered in red, while happiness might be rendered in green. (See reference...) Figure 12 Further details on atmosphere rating are discussed. In some implementations, the avatar visualizer 316 can selectively identify certain emotions among the emotions to generate expressive avatars, which are then returned to the viewer's client device for rendering on the viewer's image representation. For example, the number of emotions identified by the emotion visualization engine 313 may be excessive, and rendering avatars for all emotions identified in the audience may result in overcrowding on the display screen. Therefore, to prevent such overcrowding of avatars while ensuring proper representation of the emotions of the viewer 102 in the audience 103, the avatar visualizer 316 can select a predefined number of emotions to represent using avatars and generate avatars accordingly. The emotions used for representation can be selected based on their confidence scores associated with the corresponding atmosphere faction. For example, the maximum number of avatars to be presented on the audience can be predefined as 5. Therefore, when more than 5 emotions are identified from the viewer's audience, the avatar visualizer 316 can select the top 5 emotions with the highest confidence scores (i.e., those with higher dominance levels) (i.e., dominant emotions) to generate avatars.

[0145] The expressive avatar is returned to the client device of audience 102 for rendering on the image representation of audience 103 presented alongside the video game content. The expressive avatar provides a visual representation of the dominant emotion within the crowd (i.e., audience 103). Based on this emotional visual representation, viewers may be able to identify the group of audiences with whom they congregate, making it appear as if they are watching a live event (e.g., a game) together in a stadium.

[0146] Figure 5The illustration depicts various components of an interaction analyzer 314, in one implementation, for processing interaction data collected from viewer interactions during online gameplay of a video game. As previously described, interaction data may include chat content, live video, live audio, etc. Chat content is obtained from chat interfaces, instant messaging interfaces, or social media interfaces, through which video game viewers and players communicate to express their thoughts and provide comments. An interaction collection engine 311 collects viewer interactions and forwards them to the interaction analyzer 314 for further processing. The interaction analyzer 314 analyzes the interactions to identify different data patterns contained within them. These modalities correspond to the types of content included in the interactions. In an implementation of watching live gameplay of a video game, viewer interactions are captured at the corresponding client device during live gameplay and transmitted as data streams to an emotion expression engine 304. These data streams may include data of different modalities, such as live video data capturing viewer expressions, text data, audio data, graphic data, etc.

[0147] Interaction analyzer 314 receives and processes each element of the interaction (e.g., chat interaction, live video, and audio content) in real time to identify different modal data streams. Interaction analyzer 314 may include multiple sub-modules to process the different modal data streams identified from the interactions. In one implementation, the data stream related to chat content may be processed by chat comment analyzer 314a, the data stream of live video content may be processed by video content analyzer 314b, and the data stream of audio content may be processed by video content analyzer 314c. A similar process is followed when viewers watch a replay of a video game's gameplay at a later time, collecting interactions from viewers and identifying and processing different data streams.

[0148] Chat comments can include text, emoticons, emojis, GIFs, etc., provided by different viewers as part of their interactions with other viewers or players, or as general comments related to gameplay of a video game or as part of reactions related to viewers or players. Chat comment analyzer 314a identifies different modal data included in chat interactions and processes each modal data stream separately. For example, emoticons, GIFs, and emojis within the chat interaction can be extracted and provided as input to face detection algorithm 320a. Face detection algorithm 320a uses machine learning algorithm 320 to identify facial features and crop images included in emoticons, GIFs, emojis, and other graphic content to include only relevant facial features from the graphic content. The cropped images are provided as input to expression recognition neural network 320c. Face detection algorithm 320a and expression recognition neural network 320c are part of visual sentiment analysis engine 325. Various facial samples are used to train expression recognition neural network 320c. The facial expression recognition neural network 320c, aided by the machine learning algorithm 320, uses trained information to identify salient emotions and thoughts expressed by viewers in emoticons, GIFs, emojis, and other graphic images provided through the chat interface. The machine learning algorithm 320 compares each facial feature (e.g., eyes, nose, mouth, etc.) individually, in combination, and as a whole with the trained data from the facial expression recognition neural network 320 to find the salient emotion of the best-matching graphic image. The identified salient emotions are provided as input to the avatar visualizer 316. Similarly, text content within the chat interaction is extracted and forwarded to the emotion keyword recognition (ID) engine 320d.

[0149] In addition to processing chat content, the interaction analyzer 314 also processes live video of the audience captured in real time while the audience is watching gameplay of a video game. This live video can be captured by one or more image capture devices, such as cameras, facing the audience. The camera can be integrated into the audience's client device or can be an external camera communicatively connected to the client device. The captured video includes at least the audience's face. Besides capturing the audience's image, the camera can also capture the audience's verbal reactions (e.g., swearing, comments, reactions, etc.) while watching the video game. The video content analyzer 314b forwards a portion of the live video to a face detection algorithm 320a, which crops the audience's image to include facial features and forwards the cropped image to an expression recognition neural network 320c. The expression recognition neural network 320c processes the facial features from the cropped audience image in a manner similar to how graphical images from chat content are processed. For example, machine learning algorithm 320, with the assistance of facial expression recognition neural network 320c, identifies each captured facial feature and the overall facial features of the viewer captured in the cropped image, and compares each facial feature and the overall facial features with trained data from the facial expression recognition neural network to identify the salient emotions (and thoughts) of the viewer's facial features that best match.

[0150] In some cases, analysis of facial features may lead to the identification of expressions corresponding to multiple emotions. In this case, machine learning algorithm 320, aided by expression recognition neural network 320c, can calculate a probability score 325a for each emotion identified from facial features of the audience captured in images included in a live video. The salient emotions associated with the audience are determined based on the probability scores 325a for the multiple emotions identified from the audience's facial features. The live video of the audience may include both video and audio components. In some implementations, a unimodal approach is employed by feeding only the video component of the live video to the expression recognition neural network 320c to identify salient emotions detected from the audience's interaction. In other implementations, a multimodal approach is employed by pairing the video component with the corresponding audio component captured in the live video and forwarding the paired content (i.e., video and audio content) to the face detection algorithm 320a for forward transmission to the expression recognition neural network 320c. Similar to unimodal methods that use only cropped images from the video, in multimodal methods, the facial expression recognition neural network 320c uses cropped images from the video and associated audio to identify salient emotions detected from each viewer. The audio data can be used to further refine the salient emotions identified from the facial features of each viewer. The salient emotions of the multiple viewers constituting audience 103 are forwarded as input to the avatar visualizer 316.

[0151] In addition to chat and video content, audio content generated during online gameplay of the video game is processed by an audio content analyzer 314c of the interaction analyzer 314. This audio content may be generated by viewers or players during online gameplay and can be captured using a microphone embedded in the client device or an external audio capture device (e.g., an audio recorder, external microphone, etc.) communicatively coupled to the viewer's client device. Alternatively, the audio content may originate from audio clips included in chat content. In some implementations, the audio content analyzer 314c may process the audio content by applying filters to filter out ambient noise and / or selectively extract specific audio signals from the audio signal. The processed audio content is forwarded to an audio input processor 320b. The audio input processor 320b includes or uses a speech-to-text converter 320b1 to convert the audio into text. The converted text is forwarded by the audio input processor 320b to a keyword recognition engine 320d. In addition to the text content, the audio content is analyzed to identify certain audio features and use these features to detect emotions. For example, regardless of the content, the pitch, amplitude, duration, etc., of the audio may convey different emotions. Therefore, the audio input processor 320b can combine the audio analyzer module (not shown) with a machine learning algorithm to analyze audio content and extract certain audio features, such as pitch, amplitude, duration, etc., and predict the emotion associated with the extracted audio features of the audio content. The emotion identified from the extracted audio features of the audio content is forwarded from the interaction analyzer 314 to the avatar visualizer 316 as input.

[0152] The keyword recognition engine 320d receives text content and converted text content from audio data included in the audio components of chat content and live video. The keyword recognition engine 320d examines the text content and identifies keywords included within it. The text content may include keywords related to expressions associated with certain emotions and / or to discussion topics within the chat interface. It should be noted that any reference to the chat interface can be extended to include an interactive interface through which viewers can communicate with each other and with players of the video game. Discussion topics may be about the game status or gameplay of the video game, or they may be about comments or behaviors of viewers or players, or they may be about other content related to the video game. For example, discussion topics may involve player playstyles, comments related to the gameplay or content of the video game, comments related to viewers or players, comments responding to interactions of viewers or players, etc. Keywords identified by the keyword recognition engine 320d are forwarded to the keyword analysis engine 315 as input for further processing.

[0153] Figure 6The diagram illustrates the various components of the keyword analysis engine 315 used to identify keywords related to emotions and to cluster viewers based on the emotions expressed in their interactions. For example, the keyword analysis engine 315 begins by performing emotion keyword detection 315a by identifying keywords related to basic emotions such as happiness, fear, sadness, anger, surprise, disgust, jealousy, anticipation, loneliness, and trust, as well as keywords related to expressions associated with those emotions. Some keywords identified by the interaction analyzer 314 may not directly express emotions but are expressions that can be associated with them. The keyword analysis engine 315 uses a machine learning algorithm 320 to identify expressive keywords that can be associated with different emotions. For example, keywords such as funny, pleasant, or excited may be associated with happy emotions; keywords such as clingy, moody, or critical may be associated with sad emotions; keywords such as panicked, frightened, or tense may be associated with fear; and keywords such as annoyed or bad-tempered may be associated with anger, and so on. Of course, some expressive keywords may be synonyms for the corresponding emotions, while others may not. Machine learning algorithm 320 uses the history of audience interaction in video games and / or other video games, as well as the context of providing expressive keywords at the interface during the current game session, to correctly identify the emotions associated with the expressive keywords.

[0154] Keywords (i.e., keywords that identify basic or primary emotions and / or expressive keywords associated with emotions) are used to dynamically generate and populate word clouds in real time (315b). The keywords in the word cloud correspond to the current emotions of the audience. When additional interactions are received from the audience, additional keywords are identified from the text, and the word cloud is dynamically updated to reflect the audience's emotions.

[0155] The word cloud is examined to identify the various emotions expressed by viewers (315c) within the audience. As part of identifying these emotions, the keyword analysis engine 315, aided by the machine learning algorithm 320, identifies various keywords related to each emotion and indexes the keywords accordingly. Indexing each keyword is done to identify the emotion to which the keyword belongs and the number of viewers who include the keyword in their interactions during online gameplay of the video game. The index is used to cluster (315d) the keywords into different atmosphere factions according to emotion, where each atmosphere faction corresponds to a unique emotion. Thus, emotion keywords such as happy and expressive keywords such as funny, cheerful, excited, etc., can be clustered together under atmosphere factions associated with happy emotions. Similarly, emotion keywords such as sad and expressive keywords such as clingy, moody, or critical are clustered under atmosphere factions associated with sad or grief emotions, and emotion keywords such as fear, dread and expressive keywords such as panic, horrible, tense are clustered under atmosphere factions associated with fear emotions, and so on. Based on the number of people who have used each keyword to express the associated emotion, the size of the keyword in the word cloud can be adjusted, whereby the size is used to visually represent the number of viewers who have used the keyword to express the corresponding emotion.

[0156] The word cloud provides a visual representation of the various emotions expressed by viewers through text or audio content, and the size of the keywords indicates the number of times different viewers used the keywords during the interaction. In addition to clustering viewers into atmosphere factions, the keyword analysis engine 315 also calculates a confidence score for each atmosphere faction. The confidence score for each atmosphere faction is calculated as the number of viewers in the audience who used keywords or expressed a dominant emotion or a variation of a dominant emotion through facial expressions. The emotion expressed by each viewer is determined using a probability score 325a, which is calculated for each emotion identified from the analysis of facial features captured from images obtained from live video, where the emotion expressed by the viewer is identified as the emotion with the highest probability score 325a.

[0157] Figure 7An avatar visualizer 316 is illustrated in one implementation for generating representative avatars for rendering on an image representation of the audience. The avatar visualizer 316 receives salient feelings and thoughts (i.e., emotions) provided by an interaction analyzer 314 and emotion keywords provided by a keyword analysis engine 315. The emotion keywords and emotions are used to create expressive avatars to represent the unique emotions of each audience cluster (i.e., atmosphere faction) identified among the audience. Additionally, confidence scores associated with different atmosphere factions are used to scale the corresponding expressive avatars. First, the avatar visualizer 316 categorizes emotions (316a) to determine the number of emotions identified from the audience's interactions. When the number of identified emotions is too large, some emotions are clustered together based on the degree of similarity detected among those emotions (316b). For example, primary emotions (e.g., happiness, sadness, fear, anger, etc.) can be identified from the input provided by the interaction analyzer 314 and the keyword analysis engine 315. A similarity score is calculated for each emotion identified from the input. The similarity score for the primary emotion is defined as 1, while the similarity scores for other emotions that are variations of the primary emotion are defined as numbers between 0 and 1. For example, taking emotion keywords, the keyword "happiness," representing the primary emotion, is assigned a similarity score of 1. Other keywords that are variations of happiness, such as smiling, satisfaction, ecstasy, laughter, joy, and excitement, are assigned similarity scores between 0 and 1. In one implementation, the similarity scores for keywords that are variations of the primary emotion can be determined based on the context in which the keywords are used in the interaction. The similarity scores are used to identify the dominant emotion (i.e., the primary emotion) and to categorize different emotion clusters identified from the interaction into atmosphere factions defined for the dominant emotion. Thus, atmosphere factions can be formed by including emotions with similarity scores that differ from the dominant emotion by a predefined percentage (e.g., 5% or 10%) or a predefined number (e.g., 0.005-0.010). Based on the examples above, the atmosphere of happiness or joy can include audiences who express happiness, as well as audiences who express variations of happiness (e.g., smiling, contentment, ecstasy, laughter, pleasure, excitement, etc.).

[0158] As part of clustering viewers into atmosphere factions based on the emotions they express, a confidence score for each atmosphere faction is calculated as the number of viewers expressing emotions associated with that atmosphere faction. Therefore, for each atmosphere faction, the confidence score is the number of viewers expressing the primary emotion and variations of that primary emotion associated with the corresponding atmosphere faction. In addition to clustering emotions and calculating confidence scores for various emotions, the avatar visualizer 316 can also determine the number of emotions detected in the audience. When the number of detected emotions is too large, the avatar visualizer 316 can select only a predefined number of avatars to represent them. For example, if the number of identified / detected emotions is 10 or 12 (e.g., 10 or 12 basic emotions), the avatar visualizer 316 can identify the top 5 emotions to represent using avatars, where 5 can be a predefined number.

[0159] Additionally, the avatar visualizer 316 can determine whether the expressed emotion is inherently positive or negative. In one implementation, the avatar visualizer 316 can determine the positive or negative nature of each expressed emotion by referring to psychological literature available to the emotion expression engine 304. The nature of the emotion can be used to present the avatar in different colors to provide a more intuitive representation of the emotion in the audience. For example, happiness is considered a positive emotion, while anger, sadness, or fear is considered a negative emotion. Positive emotions can be represented in green, while negative emotions can be represented in red. When more than one positive emotion is identified, each avatar representing a positive emotion can be represented as a variation in green intensity, where the most positive emotion has the strongest green and the least positive emotion has a lighter shade of green. A similar variation can be applied when more than one negative emotion is detected in the audience. Alternatively, each basic emotion can be represented by a different color. In addition to color, the avatar visualizer 316 can also identify additional features to be included (i.e., mixed) when generating expressive avatars for each emotion identified by the interaction analyzer 314 and the keyword analysis engine 315. The avatar visualizer 316 then performs emotion mixing 316c by including all features (e.g., color, etc.) for each emotion recognition when generating an expressive avatar of an emotion.

[0160] The avatar visualizer 316 then performs an emotion assessment (316d) by adjusting the expression of the appropriate avatar to match the stated emotion. During the assessment, the avatar visualizer 316 blends various features (e.g., color, size, etc.) for each emotion recognition to generate an avatar that appropriately represents the dominant emotion in each mood faction. In one implementation, the avatar is generated in the form of an emoji. It should be noted that rendering an emoji as an expressive avatar is one way to represent an emotion, and other forms of avatars or representations can also be used.

[0161] Once an avatar for each emotion is generated, the avatar visualizer 316's emoji / avatar scaling engine 316e dynamically scales the avatar generated for each emotion using a confidence score calculated for each atmosphere faction. Dynamic scaling provides a visual indication of the dominance level of each emotion within the audience, with the avatar corresponding to the most dominant emotion being larger than those for other emotions. The ratings and scaled avatars representing different emotions are returned to the viewer's client device for rendering on an image representation of the audience presented alongside the video game content. The emotion expression engine 304, using machine learning algorithm 320, provides a fast (i.e., near real-time) way to measure the distribution of reactions among a large number of game viewers scattered across a wide geographic area. Aggregated visual feedback is intuitive for the viewer, allowing them to visualize the various emotions within the audience. The visual representation also allows viewers to compare their reactions to those of different groups and identify the group within the audience whose emotions most closely match their own (i.e., the atmosphere faction).

[0162] Figure 8A This illustration depicts the process of analyzing video portions of a live video to identify different emotions expressed by the audience, based on a single-modal emotion recognition method using a machine learning algorithm employed by an emotion expression engine 304, according to one implementation. In this method, various video frames captured in the live video capture the audience's facial features. Each video frame captured in the live video is analyzed to identify the embedded facial features. Each video frame is cropped to include the facial features. The machine learning model uses these facial features to classify the audience's different emotions.

[0163] The interaction analyzer module 314 independently processes each modality of content captured from or generated by each viewer to infer the expressed emotion and fuse emotions from different modalities identified from interaction data associated with each viewer. The fused emotion for each viewer is then provided as input to the avatar visualizer 316 to visualize the emotion using an avatar. Figure 8A In the example illustrated, video modal data (i.e., live video) is being processed by the interaction analyzer 314. Other modal data identified in the interaction data can be processed in a similar manner. As shown, the interaction analyzer 314 can use a visual emotion analysis engine 325 along with a machine learning algorithm 320 to process the video modal data included in the live video, wherein the processing includes extracting features (e.g., facial features) of each audience member captured in the live video and independently inferring the corresponding audience member's emotion with the assistance of an expression recognition neural network 320c. Figure 8AIn the implementation illustrated, the facial expression recognition neural network is a single-modal flow deep neural network, trained to use features recognized from a specific modality and predict the emotion of the features based on the viewer's specific modality. The emotions recognized from each modality's data are then combined with emotions recognized from other modalities' data for each viewer. For example, as shown... Figure 8A As shown, emotion is predicted for live video. The inferred emotion from the live video is then fused with emotions predicted based on other modalities included in the audience's interaction data, for example, through a weighted average. Therefore, the emotion prediction for the audience is based on the fusion of predictions of different monomodal stream emotions for that audience. An avatar visualizer 316 receives the fused emotion for each audience member, determines the distribution of various emotions within the audience, and generates appropriate avatars to visualize the audience's emotions within the audience.

[0164] Figure 8B This illustration depicts the process by which a multimodal emotion recognition method, based on a machine learning algorithm 320 utilizing an emotion expression engine 304, analyzes the video and audio portions of a live video to identify different emotions expressed by the audience, according to one implementation. Figure 8B In the example, live video and audio features identified from different modal streams are processed to recognize the emotions expressed by the audience. Audio can be generated by the audience when the video is captured or when the audience interacts with other audience members. In this implementation, features from each modality are used as input to train an expression recognition deep neural network 320c to predict the audience's emotions. The predicted emotions are then fed to an avatar visualizer 316 to generate avatars. The avatar visualizer 316 receives the predicted emotion for each audience member, determines the distribution of various emotions within the audience, and generates an avatar with appropriate size and color that reflects the distribution of the audience's emotions. Figure 8B In the implementation illustrated, the facial expression recognition neural network is a multimodal deep neural network, trained to use features from two different modalities to predict the audience's emotions. Figure 8B The example shown only illustrates two modalities (i.e., live video and audio stream) combined to predict the emotions expressed by the audience, but in reality, more than two interactive data modalities can be used to predict the audience's emotions.

[0165] Figure 8CThis illustration depicts the main steps of a machine learning algorithm employed in a unimodal approach to identify the emotions of an audience watching gameplay in a video game, as implemented in one embodiment. In the illustrated example, the unimodal approach uses live video of the audience as input, captured using a camera or image capture device integrated into or associated with the audience's client device. The live video of the audience captured at the client device is forwarded to a cloud server. An emotion expression engine 304 receives the live video, extracts the audience image included therein, detects the audience's face, and crops the audience image to retain only facial features. The cropped image is then used to recognize the audience's facial expressions, and the expressions are associated with emotions using data trained from an expression recognition neural network.

[0166] Figure 9 A sample screen in one implementation is illustrated, showing images of the audience captured in live video during the interactive collection phase using the interactive collection engine 311. Live video of the audience, as part of the live audience watching the gameplay, is captured using an image capture device, and the live video is made available to the emotional expression engine 304. Figure 9 The sample screen shown illustrates that each viewer exhibits different emotions, but this may not always be the case. In addition to features such as hands and body parts, the viewer's image also captures facial features. Because emotions (i.e., basic emotions) are expressed using facial features, the emotion expression engine 304 identifies facial features from the viewer's image and crops the image to include only the facial features displaying the emotion. The cropped facial features are processed by the interaction analyzer 314 to identify salient emotions detected from the viewer. Figure 9 In order to protect the privacy of the audience, a graphic overlay has been provided on the real face of the audience in the image, and the facial features from the audience's real face in reality are used to determine the emotions expressed by the audience.

[0167] Figure 10 The diagram illustrates a sample facial feature recognition process for identifying audience emotions in one implementation. Figure 10A diagram illustrating cropped images of a sample set of viewers is shown, where the images are cropped to include only facial features. In one implementation, the interaction analyzer 314 uses facial features to identify salient emotions in the viewer. The interaction analyzer 314 examines various aspects of each facial feature, such as eyebrows, eyes, nose, forehead, mouth, etc. (e.g., the knot of eyebrows, wrinkles of the nose, wrinkles on the forehead, the degree of eye opening, the degree of smiling or frowning, etc.), and examines the facial features as a whole to identify the emotions detected from the viewer. In some implementations, each facial feature and the facial features as a whole are examined to identify one or more of at least six basic emotions (e.g., anger, disgust, fear, happiness, sadness, surprise, etc.).

[0168] Figure 11 The illustration depicts images of the audience from a live video feed, used in one implementation to determine the emotions expressed by the audience. (See reference...) Figure 10 As indicated, the face detection algorithm within the emotion expression engine 304 receives live video and extracts images of the audience. The extracted images are then cropped to include only the facial features of the audience (in...). Figure 11 (Represented by blue squares). The cropped image is then analyzed for various facial features to identify emotions detected from the viewer. This analysis involves comparing each facial feature, and combinations of facial features, with corresponding facial features in the training data included in the emotion recognition neural network to identify various emotions. In addition to identifying various emotions, the facial detection algorithm also identifies a matching score for each emotion, where the matching score corresponds to the degree of match between a specific feature in the facial features and a corresponding facial feature in the training data. Figure 11 An example of this is shown, where analysis of the audience's facial features has generated a set of emotions and a corresponding matching score for each identified emotion. Although Figure 11 The implementation method shown in the diagram corresponds to live video, but it can be extended to include the analysis of other graphic images, such as emojis, emoticons, GIFs, and other graphic content. Based on the matching score of the facial feature recognition matching level, it is possible to determine... Figure 11 The expressions captured in the images shown are neutral (i.e., the emotions with the highest matching scores).

[0169] Figure 12A sliding atmosphere rating scale is illustrated in one implementation, which can be used to visually represent each emotion when generating corresponding avatars. Emotions can have positive or negative atmosphere ratings. Based on the atmosphere rating of each emotion, the emotion expression engine can limit the color of the generated avatar representing the corresponding emotion. For example, anger is shown as having a negative atmosphere rating, and therefore the avatar representing anger can be represented in red. Similarly, happiness or ecstasy is shown as having a positive atmosphere rating, and therefore can be represented in green. In one implementation, the intensity of the color representing the avatar can be adjusted according to the atmosphere rating. Figure 12 In the example shown, both anger and doubt have negative atmosphere ratings, but the negative atmosphere rating for anger can be greater than that for doubt. Therefore, the embodiment of anger can be defined using a darker red, while the embodiment of doubt can be defined using a lighter red. Figure 12 The implementation illustrated in the diagram also shows sample color codes used to represent various emotions based on the atmosphere rating associated with the mood. It should be noted that other color schemes can be used to represent emotions, including using different colors to represent each unique emotion.

[0170] Figure 13 A representation of the emotion aggregation step (also referred to as the "emotion aggregation stage") 312 in one implementation is illustrated, wherein certain interactions in the audience's interaction are used to dynamically generate a word cloud and populate the word cloud with keywords representing emotions and expressive keywords related to emotions collected from the audience during gameplay of the video game. Figure 13 In the example illustrated, live video footage of the audience is used to fill the word cloud. In one example, the sentiment expression engine 304 employs a unimodal approach to fill the word cloud using information from the live video footage of the audience. It should be noted that various implementations are not limited to unimodal methods but can also include multimodal methods, where live video, audio content, and chat content can be processed simultaneously to fill the word cloud. Keywords in the word cloud are used to identify the emotions of the audience group (i.e., the audience). Figure 13 In the implementation shown in the diagram, some keywords in the word cloud are rendered more prominently than others (i.e., visually emphasized). This is likely to indicate the level of dominance of that keyword in audience interaction (i.e., the number of audience members who have expressed a particular keyword during their interaction). It should be noted that the word cloud itself is not actually rendered on any display screen of any client device, but is displayed for illustrative purposes. Figure 13 The image shows a visual representation of which keywords are more dominant than others in audience interaction. Keywords in the word cloud represent emotions or expressive keywords associated with emotions. Figure 9 Similarly, in Figure 13In the exemplary illustration, a graphic overlay is provided on the actual face of the audience captured in the audience image to protect the audience's privacy, while facial features from the actual face of the audience captured in the live video are used to fill the word cloud.

[0171] Figure 14 An exemplary representation of the emotion visualization phase 313 of the emotion expression engine 304 in one implementation is illustrated. As shown in the exemplary representation, once a word cloud is generated, the emotion expression engine 304 clusters the audiences who generated the keywords based on the emotions represented by the keywords, such that each audience cluster is associated with a unique emotion. Based on the clusters of audiences, the emotion rendering engine 304 identifies an incarnation for each emotion, adjusts the expression of the incarnation, blends various features identified for each emotion, and forwards the blended incarnation for each emotion for rendering on the audience's client device. Figure 14 It shows the target from Figure 9 The interactive collection phase captures exemplary live video footage of the audience, identifying different emotions and generating various avatar representations. For illustrative purposes only, in... Figure 14 In the example illustrated, the avatar is shown as being in Figure 9 The images of the audience depicted in the illustrations have a one-to-one correlation, while in reality, the correlation between the avatar and the audience is actually one-to-many. Despite this... Figure 14 In the example illustrated, only the results of live audience video are shown for filling the word cloud and generating avatars. However, in reality, other modal data generated by the audience, such as text data, video data, audio data, emojis, GIFs, emoticons, and other graphic content, are also used to fill the word cloud and / or identify avatars representing emotions. Figure 9 Similarly, in Figure 14 In the exemplary illustration, a graphic overlay is provided on the actual face of the audience captured in the audience image to protect the audience's privacy, while facial features from the actual face of the audience captured in the image are used to determine the emotions expressed by the audience and to provide an avatar representation.

[0172] Figure 15A visual representation of various types of interactions is provided in one implementation for generating word clouds, which are used to generate expressive avatars for rendering on an audience's visual representation. Interaction types include chat or messaging content (including text, emojis, GIFs, emoticons, other graphic content, etc.), live video content (capturing audience facial expressions while watching online gameplay of a video game), audio commentary / content, and emoji responses. Different audiences can choose not to share interactions they generate while watching gameplay of a video game, select multiple types, or all types of interactions. Based on the sharing options selected by the audience, an emotion expression engine 304 collects corresponding types of interactions from different audiences and uses these interactions to generate word clouds, which, along with the audience's facial expressions, are used to identify emotions, and generate and scale avatars for each identified emotion. Figure 9 , Figure 13 and Figure 14 Similarly, in Figure 15 In the exemplary illustration, a graphic overlay is provided on the actual face of the audience already captured in the audience image to protect the audience's privacy, while facial features from the actual face of the audience captured in the live video are used to determine emotions.

[0173] In one implementation, audience interactions are collected during a live video game session. The collected interactions are used to generate a word cloud and an avatar representing the emotions identified from the word cloud, and the avatar is returned to the client device of the audience who has accessed the video game to watch its live gameplay. The generated word cloud and avatar are stored in an emotion collection database 334 for use during video game replays. During the replay, interactions from viewers watching the replay are collected and used to update the word cloud and the avatars representing the emotions identified from the word cloud and other interactions.

[0174] Figure 16 A simplified representation of the timeline of atmosphere factions generated by clustering similar atmospheres in one implementation is illustrated. Figure 16The atmosphere factions depicted in the diagram provide a visual representation of the atmosphere faction coefficients generated for different atmospheres detected by the audience at a specific time in the video game when a change in mood is detected. For example, a timeline can be plotted using a timeline of gameplay along the x-axis and the composition of atmosphere factions (and corresponding confidence scores) along the y-axis. The timeline identifies the specific time when events occur during gameplay that cause a change in the audience's mood. Changes detected during gameplay may or may not lead to changes in the composition of atmosphere factions. In some implementations, when the number of emotions identified from the audience's interactions is too large, the sentiment expression engine can cluster atmospheres of similar nature into a single atmosphere faction. In an alternative implementation, each atmosphere identified from the audience is used to generate a corresponding atmosphere faction, and the sentiment expression engine selects a predefined number of atmosphere factions with the highest probability scores to use as an avatar representation. Figure 16 In the example shown, four atmosphere factions are defined as including four different types of atmosphere, but in reality there may be more than four atmosphere factions defined according to the audience's emotions.

[0175] Figure 17 This illustration depicts a sample view of a representative image of an audience, overlaid with an image of an expressive avatar, in one implementation. The expressive avatar represents different emotions identified from the audience's interactions. The size, color, and other features of the avatar representing each emotion are scaled to provide an appropriate visual representation of the various emotions detected from the audience, the number of audience members expressing each identified emotion, the associated atmosphere rating for each emotion, etc. In one implementation, audience members can be associated with different geographic locations, and audience members at each geographic location can be associated with a specific emotion. For example, an audience member at geographic location 1 might associate with or follow player 1 or team 1, an audience member at geographic location 2 might associate with or follow player 2 or team 2, and so on. Therefore, audience members supporting each player or team might express similar emotions. In this example, avatars representing different emotions (i.e., atmospheres) can be rendered on a map, where each avatar represents a specific atmosphere rendered at the geographic location corresponding to an audience member associated with an atmosphere faction of that specific atmosphere. Expressive avatars provide a visual view of the various emotions expressed by the audience and allow viewers to identify and join audience groups whose emotions resonate with their own (i.e., atmosphere factions), making viewers feel as if they are playing the video game with their friends or like-minded audiences while watching gameplay.

[0176] In one implementation, once an atmosphere faction is formed, the facial expressions of the audience included in each faction are monitored to detect any changes in their expressions. The audience's facial expressions may change based on the current game state of the video game's gameplay, which is influenced by game events and player interactions. When a player scores or loses points, game prizes, or game lives, the audience's emotions may change to reflect their feelings towards the player, the outcome of the gameplay, etc. In addition to changes in the game, the audience's emotions may also be influenced by other viewers' reactions to the gameplay, players' comments or actions during gameplay, and other viewers' reactions / interactions. The emotion expression engine 304 monitors the changes in the emotions of the audience in each atmosphere faction and dynamically adjusts the emotions of the corresponding atmosphere faction's avatar to reflect the current emotions of the audience group included in said atmosphere faction.

[0177] In one implementation, the confidence score for identifying the number of viewers expressing the sentiments of each atmosphere faction may vary over time. This could be due to some viewers leaving the group forming the atmosphere faction or new viewers joining the group. New viewers can join the first group (i.e., the first cluster) from a second group (i.e., the second cluster) representing a different atmosphere faction, or vice versa. Alternatively, new viewers might join to watch gameplay of a video game. In some implementations, some viewers might join the first group simply to experience the emotions and interactions expressed by the viewers in the first group. Options can be provided on the user interface rendered next to the video game content and the image representation of the audience for the first group's viewers to grant permission for viewers from different groups (e.g., the second group, the third group, etc.) to join the first group. Additional options can be provided to viewers from other groups (e.g., viewers from the second group, the third group, etc.) to request to join the first group. Viewers from the second group, the third group, etc., can be allowed to join the group when the request is accepted by the first group or based on the settings of the first group. For example, this might be analogous to spectators watching a live game in a stadium playing with friends who may support different teams. When a spectator from a second or third group chooses to join the first group, that spectator automatically leaves the second or third group, or disassociates themselves from the second or third group and becomes attached to or associated with the first group. Similar association and disassociation can be envisioned when a spectator from the first group chooses or requests to join another group.

[0178] In one implementation, associating a viewer with a group allows the viewer to access the interactions of the group's audience. Similarly, unassociating a viewer from a group prevents the viewer from accessing the interactions of the audience of the group they are unassociated with. Providing this option allows viewers to experience not only the atmosphere of their own atmosphere faction's audience but also the atmosphere of other audiences from different atmosphere factions. In another implementation, viewers within a cluster (i.e., an atmosphere faction) are allowed to interact with other viewers within the cluster and access the interactions of other viewers within the cluster. In this implementation, viewers of the first cluster are not allowed to interact with viewers of other clusters and have no access to the interactions of viewers of other clusters.

[0179] In one implementation, an interactive timeline can be generated and presented to viewers in the audience to indicate the intensity of different emotional responses expressed by or detected by the viewers of the video game. Figure 16 An example of this is illustrated. The intensity of reactions to different emotions can vary based on changes occurring within the gameplay of a video game. In one implementation, the intensity of reactions captured in an interactive timemap is associated with a specific portion of the gameplay of the video game, allowing viewers to visualize the intensity of reactions to specific emotional expressions in the timemap and correlate it with specific changes occurring within the gameplay (e.g., events) of the video game. In some implementations, viewers may be able to click on any of the atmosphere factions included in the timemap at a specific time, and viewers may be able to watch gameplay of the video game corresponding to the intensity of reactions to a specific atmosphere faction represented in the timemap at that specific time. The interactive timemap can supplement or replace an avatar rendered on an image representation of the audience alongside the content of the gameplay of the video game. The timemap can be generated during a live broadcast of the gameplay of the video game, and the timemap can also be stored in the gameplay data storage device 332 for subsequent retrieval and presentation. Alternatively, the timemap can be stored in an emotion database 334 along with word clouds and atmosphere factions identified from viewer interactions. When a video game is replayed (i.e., in a delayed time-streamed format), a stored timemap can be retrieved and presented to viewers watching the delayed replay. As viewers interact with or react to different events or actions occurring within the gameplay of the video game during the replay, a new timemap can be generated to include data from the stored timemap and additional reactions identified from the interactions of viewers watching the delayed replay. The new timemap is stored in the gameplay data storage device 332 or the emotion database 334 and retrieved when the video game replay is rendered to the viewers.

[0180] In one implementation, selecting a specific time on the timemap causes the emotion expression engine 304 to query the buffer storing gameplay data for the video game and retrieve specific gameplay data to determine the audience's overall emotion at that specific time or to determine the different reactions that occurred at that specific time. A timemap can be generated separately for each emotion, or a single timemap can be generated for different emotions identified in the audience. In the case of generating timemaps for each emotion, multiple timemaps can be generated, one for each emotion. The timemap generated for each emotion can be rendered next to the corresponding emotion's avatar, or rendered as a thumbnail at the bottom of the client device's display screen, etc.

[0181] In some implementations, the emotions, reaction tracks, and timelines identified for rendering the avatar are represented probabilistically, using only the dominant selected emotion from among the emotions. The dominant emotion is determined based on probability scores of the emotions expressed by the audience through facial features and confidence scores of the atmosphere faction.

[0182] In some implementations, the timeline can be represented as a line chart. In this implementation, the line chart can include lines representing different emotions, with each emotion represented by a different line. In some implementations, the personification of the emotion represented by the line can be rendered next to or overlaid on the corresponding line to provide a visual indication of the emotion corresponding to that line. Figure 16 The line graphs and time graphs shown are some examples of how emotions can be detected from the audience, and other forms of visual representation of audience emotions can also be envisioned.

[0183] In some implementations, rendering an avatar on a viewer's client device may include providing a segmentation option to the user interface for the viewer to choose which avatar to render. Viewers may want to see the audience's emotions at a specific location on the screen, rather than crowding the screen where the video game content is being rendered or interfering with the rendering of other content. In these implementations, the avatar may be rendered individually or on the viewer's image representation. The display screen may be divided into multiple segments (e.g., lower half, upper half, left side, right side, etc.), and the segmentation option may include these segments for each viewer to choose for rendering the avatar. In addition to identifying segments, each viewer may be provided with options to render the avatar individually or on the viewer's image representation, as well as options to format the avatar. Based on each viewer's choice, an avatar representing different emotions detected in the viewer may be rendered on the viewer's image representation in a designated segment or without a viewer's image representation. The segmentation option provides viewers with a degree of autonomy, allowing them to visualize the audience's emotions while simultaneously viewing the gameplay of the video game.

[0184] In addition to rendering options, one or more formatting options can be provided at the user interface for viewers to choose from in order to render the expressive avatar. Some formatting options that can be included in the user interface for viewers to choose from include transparency, overlay, or rendering formats. Of course, the aforementioned formatting options are provided as examples only and should not be considered limiting. Other formatting options may also be included. The avatar can be rendered based on the formatting and segmentation options selected by each viewer, where the avatar is rendered individually or on top of the viewer's image representation.

[0185] In one implementation, in addition to generating avatars and adjusting their emotions, the emotion expression engine 304 can identify specific viewers within each atmosphere faction's audience, capture their reactions during defined game moments, and render the captured reactions alongside the video game content to provide reaction salience. In this implementation, video of the reactions of specific viewers identified for each atmosphere faction can be presented instead of expressive avatars. In other implementations, in addition to expressive avatars for atmosphere factions, reactions of specific viewers can also be provided. Specific viewers can be identified based on the type and number of comments obtained from other viewers within the group or from other groups, based on their reactions.

[0186] In one implementation, the reactions of a specific audience from a particular atmosphere faction can be captured and presented by first identifying actions planned to occur within the gameplay of a video game. The actions can be identified using the current game state of the video game and according to the game logic. A specific audience from an atmosphere faction can be identified based on the type and quantity of reactions provided by the specific audience to different actions occurring in the current gameplay or during previous gameplay of the video game. Based on this information, an emotion expression engine 304 can predictively send signals to one or more image capture devices used to capture live video of the audience to amplify the specific audience's reactions during the identified actions. The captured video of the specific audience is dynamically analyzed and presented or superseded along with an expressive avatar generated for the specific atmosphere faction. In an alternative implementation, a specific audience from a particular group can be identified based on the type and quantity of comments relating to the expressions of the specific audience obtained from the remaining audience in the particular group (i.e., the atmosphere faction). In some implementations, one or more audience members from a specific atmosphere faction within an atmosphere faction can be selected to present their expressions captured during live video streaming. In an alternative implementation, each atmosphere faction can identify a specific audience member and present the facial expression detected from the corresponding audience member within that faction. In this implementation, a specific audience member can be identified based on the reactions of other audience members within the corresponding audience group to that specific audience member's reaction. Alternatively, a specific audience member can be selected randomly.

[0187] The various implementations discussed in this paper allow viewers to observe a range of audience reactions and to associate with specific audience groups. A specific audience group can be friends with whom a viewer wishes to watch an online video game. These friends may or may not express the same emotions (i.e., they may or may not support the same player or team). Regardless of which player or team each viewer supports, and regardless of the emotions expressed by the viewer and their friends, these implementations provide viewers with a way to socialize with their friends while understanding the overall atmosphere within the online video game's audience. In one implementation, viewers clustered into an atmosphere faction can choose to remain in that faction or leave and join another. Viewers can be offered the option to request to join another atmosphere faction or to select another atmosphere faction to join. In this implementation, the option allows viewers to bypass the clustering provided by the emotion expression engine 304 and join their chosen atmosphere faction. The option can identify different atmosphere factions within the audience and allow viewers to select the atmosphere faction they wish to join. This selection can be made by dragging and dropping the viewer's icon or image from a first atmosphere faction to a second atmosphere faction, or via radio buttons, checkboxes, etc. The option to move from one atmosphere faction to another provides viewers with a way to experience the atmosphere of a second atmosphere faction. It should be noted that the interaction data generated by viewers of each atmosphere faction is shared only with viewers of that faction, and not with viewers of other atmosphere factions. In an alternative implementation, selected interaction data from each atmosphere faction's interaction data can be shared with viewers of other atmosphere factions. In this implementation, interaction data shared with other atmosphere factions may prompt viewers of other atmosphere factions to react in a manner similar to that of fans of opposing teams in a stadium.

[0188] In an alternative implementation, as a way to allow viewers to move from one atmosphere faction to another, the Emotional Expression Engine 304 can collect reaction saliences from a specific atmosphere faction and share these saliences with other atmosphere factions. Reaction saliences can be shared for specific events or actions and can be shared within predefined time periods. The reactions of viewers in other atmosphere factions to the reaction saliences of a specific atmosphere faction can also be used by the Emotional Expression Engine 304 to update the avatars and reaction tracks associated with each atmosphere faction. In some implementations, announcer avatars can be generated to host the reactions of viewers in different atmosphere factions, where hosting can include: providing reaction saliences for a specific atmosphere faction to stimulate viewers in other atmosphere factions to respond to the reaction saliences; providing reaction saliences from other atmosphere factions that respond to the reaction saliences of the specific atmosphere faction; and providing commentary that captures both reaction saliences and inverse reaction saliences to show the back-and-forth debate between viewers of different atmosphere factions. Similar to the announcer avatars, cheerleader avatars can be generated for each atmosphere faction to cheer on the player and the viewers supporting the atmosphere faction. In one implementation, the announcer and cheerleader avatars can be generated as AI robots using artificial intelligence (AI), with machine learning algorithms used to control their actions. These avatars can be used to encourage players and viewers during gameplay, including stimulating viewers, encouraging players, and intervening to moderate reactions from different factions of the audience, especially when the reactions from viewers can be described as inherently bullying or abusive.

[0189] Figure 18This illustration depicts a method for expressing the emotions of an audience watching gameplay of a video game, in one implementation. In this implementation, the audience may be watching a live performance of the video game. In other implementations, the audience may be watching a replay of the gameplay. In one example, operation 1802 may be configured to capture interaction data from the audience watching gameplay of the video game. The interaction data may include facial expressions of the audience while they are watching gameplay captured by one or more image capture devices, or interactions generated by the audience via an interactive interface such as a chat interface, message board, or social media interface, such as audio, text, emoticons, GIFs, emojis, etc., or audio captured via a microphone or other audio detection and / or recording device. The image capture device may be a camera or other image capture device as part of a client device, such as a mobile computing device (e.g., a telephone, laptop computer, tablet computer, etc.), or an external image capture device communicatively connected to the client device. The audience's image is captured as live video by the image capture device and transmitted to an emotion expression engine 304 executing on a game cloud server 300. The image includes facial features used to determine emotions detected from the viewer while they are watching an online video game. Similar to the viewer's live video feed, interactions provided via an interactive interface are also transmitted to the emotion expression engine 304.

[0190] The method proceeds to operation 1804, where the emotion expression engine 304 aggregates emotions identified from interaction data received from the audience and clusters the audience into different groups based on the emotions detected from different audiences. In one implementation, images of the audience from a live video are cropped to retain only facial features, and machine learning algorithms are used to analyze each facial feature and combinations of facial features to determine the emotions detected from the audience. The machine learning algorithm identifies keywords expressing emotions from text content included in chat content and / or audio content. Other interaction data provided via the interactive interface, such as emojis, GIFs, emoticons, etc., are also analyzed in a manner similar to the images of the audience from the live video to identify keywords defining the emotions. Emotions and emotion-related keywords identified from various interactions of each audience (images, text, emojis, emoticons, GIFs, graphics, etc. from the live video) are aggregated. The aggregated emotions and emotion-related keywords are evaluated to define a similarity score for each emotion. Similarity scores are used to determine the dominant emotion, and then, based on the emotion identified from each viewer's interaction, viewers who provide the emotion and emotion-related keywords are clustered into groups (i.e., atmosphere factions), such that each group is associated with a unique emotion. A confidence score is calculated for each group (i.e., atmosphere faction) based on the number of viewers expressing the emotion of said group.

[0191] The method proceeds to operation 1806, where cluster information is used to generate avatars for each group. Each group's avatar expresses a unique emotion associated with that group. Once a group is formed and its avatars are generated, the avatar's expression is adjusted based on changes detected in the group's emotion (i.e., sentiment). The emotion in the group may vary based on changes occurring in the gameplay of the video game. For example, a group supporting the first team might express an emotion based on the first team's gameplay, and so on. Changes in expression and other interaction data are captured and used to adjust the expression of each group's avatar. These changes may cause an avatar that previously exhibited a negative atmosphere to begin exhibiting a positive atmosphere, or vice versa. Therefore, based on changes in emotion detected from viewers within the group and based on the group's confidence score, the characteristics of each avatar are further adjusted to include changes in features such as color and size. In some implementations, the group's confidence score may change based on viewers leaving the group or new viewers joining the group. Therefore, the size of the avatar can dynamically change accordingly to reflect the number of viewers expressing the group's emotion.

[0192] The method concludes at operation 1808, where avatars generated for different emotions exhibited by the audience are presented on an image representation of the audience rendered alongside the video game content. The size of each avatar is dynamically scaled based on the confidence score of the corresponding group associated with the avatar. The scaled avatars provide a visual representation of the audience's emotions, allowing the audience to measure the various emotions and the dominance level of each emotion within the audience. It allows the audience to determine how their reactions compare to those of other viewers and to find audience groups whose emotions align with their own. Avatars also allow gamers to measure feedback on specific interactions during gameplay.

[0193] Various implementations provide a way for remote viewers to experience the atmosphere of a crowd and react to gameplay that can be shared with other users, allowing them to feel as if they are physically watching an online game with a group of viewers. While various implementations have been described with reference to viewers of online games (i.e., live gameplay of video games), these implementations can be extended to include replays of video games, where all interactions from viewers watching the replay can be similarly harvested and used to tailor expressive personas for different groups of viewers watching the replay.

[0194] Figure 19 The illustration depicts an exemplary implementation in which... Figure 4The diagram illustrates variations in the components of the Emotion Expression Engine 304. These components collect interaction data associated with the audience, analyze the data to determine different emotions, cluster emotions into different atmosphere factions, and create avatars for each faction. In addition to avatars, the components of the Emotion Expression Engine 304 also identify the target audience for the response audio tracks rendered to the viewer. Figure 19 The emotional expression engine 304 includes most of its components and Figure 4 The parts identified in the text are similar, and therefore not described in detail, because they are similar to... Figure 4 and Figure 19 The common components function in a similar way. Besides the common components, Figure 19 The emotion expression engine 304 includes a reaction track recognizer 317 within the emotion visualization engine 313. The reaction track recognizer 317 identifies and retrieves reaction tracks corresponding to the emotions expressed in each atmosphere faction, and includes these reaction tracks in conjunction with the video game content returned to the viewer's client device. The reaction track recognizer 317 can use the emotions identified by the emotion aggregation engine 312 to determine the emotions expressed by the audience, and identifies and retrieves appropriate reaction tracks from the reaction track database 318.

[0195] In some implementations, the reaction track recognizer 317 is used to identify reaction tracks only for those emotions represented by the avatar visualizer 316. As described above, the avatar visualizer 316 can select certain emotions identified in the audience to create avatars, and the reaction track recognizer 317 is used to identify reaction tracks for the emotions represented by the avatars created by the emotion visualization engine 313. The reaction track recognizer 317 can query the reaction track database 318 and retrieve appropriate reaction tracks for each emotion represented by the avatar. The reaction track database 318 can include reaction tracks for different content, for different events or actions or activities within each content, and for different emotions. The reaction tracks within the reaction track database 318 can be organized according to content type, events or actions or activities within each content, event context, and emotion. Since the emotions expressed by the audience in each atmosphere faction may change over time based on changes detected in the gameplay of the video game, the expressions on the avatars of each atmosphere faction are dynamically adjusted to correspond to changes in gameplay. In response to the facial expressions of the avatars associated with each atmosphere faction, different reaction tracks are identified for each atmosphere faction in order to relate to the current emotions of the audience in that atmosphere faction.

[0196] In one implementation, the reaction track recognizer 317 can use gameplay data of the video game stored in the gameplay data storage device 332 to determine the background of events, actions, or activities in the gameplay of the video game that cause emotional changes in the audience of each atmosphere faction. The background and event or action or activity data can be used to identify appropriate reaction tracks to return to the audience. The reaction track recognizer 317 can also use interaction data in the audience interaction data storage device 332a to determine changes in the background of reactions identified from the audience, wherein the reactions can be expressions captured from images of facial features, verbal interactions captured using an audio recording device, or interactions captured from interactive interfaces such as chat interfaces, message boards, social media interfaces, etc. Alternatively or additionally, the reaction track recognizer 317 can identify appropriate reaction tracks by: querying the emotion collection database 334 to identify the current emotion of each atmosphere faction identified in the audience's audience; and querying the reaction track database 318 to retrieve the appropriate reaction track for the current emotion of each atmosphere faction. The Emotion Collection Database 334 is a repository for storing the following: various emotions identified in the audience at different times during gameplay of the video game, avatars created for each atmosphere faction identified in the audience of the video game's viewers, and all changes made to the expressions of the corresponding avatars based on changes captured from the audience during gameplay of the video game (through facial expressions, verbal or via interaction with the interface).

[0197] In one implementation, the reaction audio tracks identified for each atmosphere faction detected among the audience of the video game's viewers can be based on the current context of the video game's gameplay. The current context depends on one or more events or activities occurring within the gameplay, which may depend on actions performed by the video game players. Therefore, the reaction audio tracks for each atmosphere faction can be identified as corresponding to the context of the video game's gameplay because it is relevant to the audience of that atmosphere faction.

[0198] Reaction tracks retrieved from the reaction track database 318 for different atmosphere factions are returned to the viewer's client device for rendering along with the generated and updated avatars. In one implementation, the reaction tracks for each atmosphere faction are presented to the audience's viewers along with the corresponding avatar, allowing viewers to visually and aurally experience the various atmospheres expressed by the audience. In another implementation, avatars expressing different emotions and reaction tracks are presented to all viewers of the audience. In this implementation, the reaction tracks for various emotions provide an auditory representation of the actual atmosphere, while the avatars provide a visual representation of the different emotional reactions of viewers to events or actions occurring in the gameplay of the video game, making viewers feel as if they are actually in an arena or stadium, watching and reacting to a live sporting event with their friends and fellow sports fans.

[0199] In an alternative implementation, an avatar and a response track for a specific emotion expressed by the audience of each atmosphere faction can be provided to the audience of that atmosphere faction. In this alternative implementation, the audience of each atmosphere faction sees an avatar expressing the emotion of that atmosphere faction and experiences a response track specific to that atmosphere faction, rather than avatars or response tracks of other atmosphere factions. This implementation can be provided as an option for audiences of different atmosphere factions to reduce the crowding of content rendered on the client devices of each atmosphere faction's audience and / or avoid the audience being overwhelmed by response tracks of other atmospheres in the audience. In yet another implementation, the response track of the most dominant emotion (i.e., feeling) is identified, retrieved, and presented on the respective client device of each audience member. The most dominant emotion is identified based on a confidence score associated with each atmosphere faction, and the response track of the most dominant emotion is identified and presented to the audience. In this implementation, an expressive incarnation of the emotion associated with the corresponding atmosphere faction can be presented to the audience of each atmosphere faction, and a reaction track representing the dominant emotion can be presented to all viewers of the audience to indicate the dominant emotion expressed by the audience. Alternatively, an incarnation representing all emotions identified in the audience can be presented to all viewers, and a reaction track representing the dominant emotion can be presented to all viewers of the audience.

[0200] In one implementation, options can be provided to viewers on the user interface to select how they want to experience the reaction soundtracks identified for an online game audience, whether they are members of that audience or are watching the video game at a later time (i.e., subsequently and not in real-time). These selection options can be provided to allow each viewer to customize the rendering of the reaction soundtracks identified for their audience on their own client device. For example, the selection options could include a first option for rendering reaction soundtracks for all emotions expressed in different atmosphere factions identified for the audience, a second option for rendering reaction soundtracks for the audience's most dominant emotion, a third option for rendering reaction soundtracks for only the emotions expressed in the atmosphere faction associated with the viewer, and so on. Similar selection options can also be provided to viewers to select the avatar to be rendered on the client device. Viewer selections are detected and used by an emotion visualization engine 313 to identify and render appropriate reaction soundtracks alongside the avatars and video game content on the respective client devices. Furthermore, the rendering of avatars and reaction soundtracks can be based on segmentation and formatting options selected by the viewer. Rendering reaction soundtracks alongside video game content allows viewers to experience a sense of camaraderie within the audience and encourages them to engage with the gameplay. In some implementations, reaction soundtracks capture not only viewers' reactions to events occurring in the video game, but also the inverse reactions of other viewers to a particular viewer's response.

[0201] Figure 20 Examples of various reaction audio tracks are illustrated in one implementation for recognizing emotions expressed by audience members. Images of the audience captured by an image capture device are analyzed to identify the emotions of each audience member. Although in Figure 20 The text only shows images from live video footage of the audience to identify emotions, but it should be noted that, compared with previous references... Figures 5 to 7 The description also mentions analyzing audience audio and interactive data provided by the audience via an interface in a similar manner to identify the emotions expressed by the audience. The audience is clustered into atmosphere factions, and an avatar is created for each atmosphere faction to correspond to the emotions expressed by the audience of that particular atmosphere faction. Figure 20 In the exemplary diagram, four distinct atmosphere factions were identified, corresponding to emotions such as laughter, anger, neutrality, and surprise (i.e., atmosphere). Once the atmosphere factions were identified, response tracks corresponding to the emotions of each atmosphere faction were identified. Figure 20 This describes the reaction audio tracks identified for the emotions identified for different atmosphere factions. Because changes are detected in the gameplay of the video game, the audience's emotions within each atmosphere faction change over time; therefore, the reaction audio tracks identified for each atmosphere faction also vary to correspond to the changes identified in the emotions within that atmosphere faction. Figure 9 , Figure 13, Figure 14 and Figure 15 Similarly, in Figure 20 In the exemplary illustration, a graphic overlay is provided on the actual face of the audience already captured in the audience image to protect the audience's privacy, while facial features from the actual face of the audience captured in the live video are used to determine emotions.

[0202] Figure 21 An example of a reaction track in one implementation is illustrated, which is identified and presented alongside an expressive avatar on a representative image of the audience. In one implementation, the volume of the reaction track for each atmosphere faction is adjusted to match the size of the avatar, which corresponds to a confidence score associated with the emotion corresponding to that atmosphere faction. Since the size of the avatar rendered on the representative image of the audience corresponds to the number of audience members in the atmosphere faction expressing the emotion associated with that avatar (i.e., the confidence score of the atmosphere faction), the volume of the reaction track is adjusted to be related to the number of audience members associated with the atmosphere faction. Figure 21 The various ambient factions' reaction tracks shown are represented by different sizes to visually indicate the relative volume of each reaction track rendered at the viewer's client device. For example, the reaction track representing the most dominant emotion (i.e., the ambient faction with the largest size incarnation) is rendered larger than the reaction tracks representing less dominant emotions, in order to indicate that the volume of the dominant emotion's reaction track is louder than that of the less dominant emotion. Changing the volume of the reaction tracks provides the viewer with a realistic representation of the audience's emotions.

[0203] The number of audience members in each atmosphere faction can change based on audience members leaving or new audience members joining. Therefore, the size of the avatar and the volume of the corresponding atmosphere faction's response track are dynamically adjusted to correspond to the size of the atmosphere faction associated with each mood. Figure 21 Visual examples are shown of reaction tracks rendered alongside corresponding avatars from different atmosphere factions. The size of the reaction track rendered next to each avatar corresponds to the volume of the corresponding reaction track, where the volume of the reaction track is related to the size of the corresponding avatar rendered on the audience image. The size of the avatar corresponds to the number of audience members in the atmosphere faction, and the size of the reaction track indicates the volume of the corresponding reaction track rendered on the client device. Figure 21 In its implementation, it provides reaction tracks for each atmosphere faction for rendering, so that the audience can more realistically experience all the atmospheres identified in the audience.

[0204] Figure 22An example is illustrated where, in one implementation, each atmosphere faction identified among the audience is associated with a corresponding reaction track, and each atmosphere faction's reaction track is included with its corresponding avatar for presentation to the audience of that particular atmosphere faction. In this implementation, the audience of each atmosphere faction is presented with the avatar and reaction track of the atmosphere faction to which the audience is a member, rather than with the avatars and reaction tracks of all atmosphere factions identified among the audience. The presentation of atmosphere faction-specific avatars and reaction tracks can be driven by selection options chosen by the audience of the corresponding atmosphere faction, and the presentation can vary from audience to audience and / or from atmosphere faction to atmosphere faction. For example, the audience of a particular atmosphere faction within the atmosphere faction may choose to receive the avatar and corresponding reaction track of their own atmosphere faction, which is associated with or more closely matches their own, while the audience of the other atmosphere factions within the atmosphere faction may choose to receive avatars and reaction tracks for all emotions identified by the audience. In another implementation, specific viewers within a particular atmosphere faction can choose to receive avatar and reaction tracks associated with that particular atmosphere faction, while the remaining viewers within the atmosphere faction can choose to receive avatar and reaction tracks for all emotions identified among the audience. The emotion visualization engine 313 detects the selections made by viewers from different atmosphere factions and identifies and presents avatar and reaction tracks based on these selections.

[0205] Figure 23 An exemplary representation of a reaction track associated with a dominant emotion rendered alongside a corresponding expressive avatar on the audience's image representation is illustrated in one implementation. The avatar visualizer 316 of the emotion visualization engine 313 identifies each emotion expressed by the audience in the audience for which an avatar will be created, and provides the avatar for the identified emotion. The reaction track recognizer 317 uses the emotion associated with the created avatar to identify the appropriate reaction track for rendering alongside the corresponding avatar. The reaction track recognizer 317 also uses selection options chosen by the audience to render the avatar and / or reaction track, and provides the appropriate avatar and / or reaction track based on the selected selection options. Figure 23 In the example illustrated, the audience within an atmosphere faction may have selected options for using a reaction track to render all incarnations identified within the audience and only for rendering a dominant emotion on the viewer's client device. Therefore, the reaction track recognizer 317 examines the confidence score associated with each atmosphere faction to determine the dominant emotion within the audience and retrieves an appropriate reaction track for rendering on the client device. Figure 23An example is shown where a neutral emotion is presented as the dominant emotion in the audience, and a reaction track is identified for the neutral emotion, and the reaction track is rendered together with an expressive avatar associated with the neutral emotion. As previously mentioned, the reaction track for the most dominant avatar can be rendered based on selection options chosen by the audience or audience group of the atmosphere faction from an interactive interface that renders various selection options for presenting the avatar and reaction track. In response to the selection options, Figure 23 The image representation of the audience is presented to the audience or audience group.

[0206] In some implementations, a reaction interface can be provided for viewers to select from, where the reaction interface includes a list of reactions or comments from different viewers for access and viewing. In some implementations, the reactions or comments of a specific viewer or a specific set of viewers in response to a specific viewer's reaction or comment may be more popular than the gameplay of a video game or reaction soundtracks associated with different atmosphere factions. Therefore, the reaction interface provides the option to access and view the reactions or comments of a specific viewer or a specific set of viewers in response to a specific viewer's reaction or comment. This option can be provided only to viewers within an atmosphere faction whose members are specific viewers or a specific set of viewers, or it can be provided to all viewers in the audience.

[0207] Figure 24 This illustration depicts a method for identifying and presenting the emotions and corresponding reaction audio tracks of an audience watching gameplay of a video game, in one implementation. In this implementation, the audience may be watching a live gameplay of the video game. In other implementations, the audience may be watching a replay of gameplay of the video game. In one example, operation 2402 may be configured to aggregate interactive data provided by an audience participating in watching gameplay of the video game. The interactive data may include facial expressions of the audience captured in live video by one or more image capture devices, or audio content expressed by the audience and captured via a microphone or other audio detection and / or recording device, or interactive content generated by the audience via an interactive interface such as a chat interface, message board, social media interface, etc. (e.g., audio, text, emoticons, GIFs, emojis, etc.). The audience's image is captured as live video by an image capture device. The audience's interactive data is transmitted to an emotion expression engine 304 executed on a game cloud server 300 for processing.

[0208] Interaction data is aggregated and processed to identify the emotions expressed by viewers while they watch gameplay of a video game. As part of the aggregation, the emotions expressed by viewers are identified (visually via facial features, verbally, or via the interface), and viewers are clustered into different groups based on the emotions (i.e., sentiments) detected from different viewers. In one implementation, machine learning algorithms are used to analyze interaction data (e.g., facial features, interaction data provided via the interface, audio content, etc.) to determine the emotions detected from viewers and cluster viewers into groups expressing the same or similar emotions.

[0209] The method proceeds to operation 2404, where cluster information is used to identify reaction tracks to correspond to unique emotions associated with each audience group. Reaction tracks are identified based on the group's current emotion, and different reaction tracks are identified to match changes in the group's emotion over time. Reaction tracks can be identified based on the video game's content, background, the emotions of the audience within the group, etc.

[0210] The method concludes at operation 2406, where a reaction audio track identified for each audience group is presented on an image representation of the audience rendered alongside the video game content. The volume of the reaction audio track associated with each group is calibrated to correspond to the number of viewers in the group. The reaction track provides an auditory representation of the audience's emotions, allowing viewers to experience the emotions of other viewers within the audience and identify audience groups that share similar emotions expressed by those viewers.

[0211] In some implementations, the reactions of a specific viewer from a specific group can be highlighted and presented to viewers of that specific group, or to all viewers of the video game. In this implementation, the specific viewer can be selected based on the emotions expressed by the viewer while watching gameplay of the video game (during live or deferred gameplay). Once the specific viewer is identified, video of the specific viewer expressing emotions is captured and presented during key game moments related to events occurring in the gameplay of the video game. The captured video of the specific viewer expressing emotions is presented as a reaction highlight along with the content of the video game. In one implementation, the specific viewer is randomly selected. In another implementation, the specific viewer is selected via predictive analysis of previous facial expressions expressed by the specific viewer during gameplay of the video game. Previous facial expressions of the specific viewer can be identified from the current gameplay session, from a previous gameplay session of the video game, or from a gameplay session of another video game. The emotion expression engine 304 can determine that a specific viewer provides emotions distinct from other viewers in the specific group, and this determination can be made by analyzing interaction data of viewers in the specific group. The Emotional Expression Engine 304 analyzes the gameplay of a video game to determine its current game state and interacts with game logic to determine when to schedule key game events based on the game state. It also responsively signals one or more image capture devices to focus on a specific audience member within a specific group to capture their facial expressions. The captured video of the specific audience member is then streamed live to the specific group or to all viewers watching the gameplay. This live video of the audience provides viewers within the video game's audience with an interesting aspect of watching the gameplay.

[0212] In another implementation, instead of video for the audience, Graphics Interchange Format (GIF) images are identified to provide reaction salience for the video game. GIFs are selected to express specific emotions relevant to a particular group. GIFs can be identified using keywords that identify specific emotions within a particular group. Similar to the audience's video, the identified GIFs can be provided as reaction salience during key game moments relevant to events in the video game.

[0213] The various implementations discussed in this article provide viewers with ways to engage with gameplay in video games and connect with other viewers who are also watching gameplay. Reaction tracks and avatar representations allow viewers to quickly gauge the atmosphere within the audience watching gameplay and identify specific audience groups to engage with.

[0214] Figure 25This illustration depicts some exemplary components of an emotion expression engine 304, implemented in one manner, used to express the emotions of viewers observing gameplay in a video game. Figure 25 The components of the emotional expression engine 304 shown in the illustration are different. Figure 4 and Figure 19 The reason for the components shown in the drawing is that Figure 25 Includes Graphics Interchange Format (GIF) file recognition engine 319. In Figure 4 , Figure 19 and Figure 25 Some components of the CCP are for reference. Figure 4 and Figure 19 The method of discussion works, and therefore there is no reference. Figure 25 Detailed discussion. (See earlier references.) Figure 4 and Figure 19 The emotion expression engine 304 includes components for collecting interaction data associated with an audience and analyzing the interaction data to determine one or more data patterns (i.e., modal data streams) included therein. Modal data streams identifiable from the interaction data may correspond to text data, video data, audio data, chat data, emoticons, emojis, graphic content, etc. The emotion expression engine 304 uses machine learning algorithms to process the one or more modal data streams using a unimodal or multimodal approach. The output of the processed modal data streams is aggregated as needed to identify the emotions expressed by the audience. The identified emotions are used to cluster the audience into different atmosphere factions, where each atmosphere faction is associated with an emotion and includes a group of audience members expressing the emotion of that atmosphere faction. In some implementations, avatars are created to represent the emotions associated with each atmosphere faction. The avatars are scaled according to the number of audience members within the corresponding atmosphere faction. When viewers explicitly (i.e., generate requests to join or associate with different atmosphere factions) or join or leave an atmosphere faction through emotions expressed via interactive data, the generated avatar is dynamically scaled to reflect changes in the number of viewers expressing the atmosphere faction's emotions.

[0215] In addition to the avatar, components of the emotion expression engine 304 are also used to identify reaction tracks for each atmosphere faction, where the reaction tracks are identified as corresponding to the emotions of the respective atmosphere faction. The identified reaction tracks for each atmosphere faction are forwarded to the viewer's client device along with the avatar for rendering. Similar to scaling the avatar, the volume of the reaction tracks associated with each atmosphere faction is calibrated to correspond to the number of viewers in the respective atmosphere faction expressing the stated emotion.

[0216] In addition to generating avatars and recognizing reaction audio tracks, the Emotion Expression Engine 304 is also used to recognize Graphics Interchange Format (GIF) files to visually represent each emotion identified by the Emotion Expression Engine 304. The GIFs identified for each emotion can include still images or animated images. Still images or animated images can include video clips, such as movie or television program clips, or other video clips including promotional content, user-generated content, etc. Figure 25 In the implementation shown in the diagram, the emotion expression engine 304 includes a GIF recognition engine 319 to identify appropriate GIFs for emotions identified in different mood factions. The identified GIFs are forwarded to the viewer's client device for rendering along with the video game content.

[0217] In one implementation, the emotion expression engine 304 returns only GIFs representing different emotions to the client device for rendering alongside the video game's video content. In this implementation, the GIF for each atmosphere faction is presented on the viewer's image representation rather than on their avatar. In another implementation, the GIF can be scaled to correspond to the number of viewers in each atmosphere faction, and the scaled GIF is forwarded to the viewer's client device for rendering. In yet another implementation, the scaled GIF can be configured to overlay the viewer's image representing the gameplay of the video game. Alternatively, the scaled GIF can be presented in a segment defined on a display screen associated with the client device, where the segment used for rendering the GIF can be defined differently by different viewers and included in their preferences. Thus, the GIF can be configured to render in an appropriate segment or portion of the display screen on each viewer's client device based on the respective viewer's preferences. These preferences can be included in the respective viewer's user profile or can be maintained separately and used when forwarding content for rendering on the viewer's client device.

[0218] In an alternative implementation, in addition to the GIF, an appropriate reaction audio track is identified for each emotion and returned to the client device along with the GIF for rendering alongside the video game content. In yet another implementation, in addition to the appropriate GIF, an expressive avatar is generated for each emotion and returned to the client device along with the GIF for rendering alongside the video game content. In this implementation, the reaction audio track may or may not be presented with the GIF and avatar. In one implementation, the GIF may be presented in a portion of the display screen defined by each viewer, while the expressive avatar is provided as an overlay on the viewer's audience's image representation.

[0219] The GIF recognition engine 319 uses machine learning algorithms to identify emotions from interactive data and then identifies appropriate GIFs for the emotions associated with each mood faction. Appropriate GIFs can be identified by querying the GIF database 321 available to the emotion expression engine 304. The GIF database 321 can be maintained within the game cloud server 300 and accessible to the emotion expression engine 304, or it can be located outside the game cloud server 300, where access is granted to the emotion expression engine 304. The GIF database 321 can be a repository containing a variety of GIFs used by different viewers to express different emotions, as well as GIFs not used by viewers but appropriate for different emotions. For example, GIFs used by different viewers could include all GIFs used by viewers in video games, during current and previous gameplay sessions of other video games, in social media, and in other interactive applications and / or interactive user interfaces. GIFs can be organized in the GIF database 321 based on interactive content (e.g., video games, social media content, user-generated content, promotional content, or other interactive content), interactive sessions, viewer preferences, viewer profiles and demographics, GIF popularity, etc. In addition to providing a GIF database 321 that offers a variety of GIFs, one or more links may be provided to access additional GIFs from one or more external GIF libraries (i.e., GIF repositories) 323 on the network 200.

[0220] In one implementation, the GIF recognition engine 319 may query a GIF database 321 and / or use a link 323 to an external GIF library to identify a subset of GIFs suitable for expressing emotions within a specific atmosphere faction. The subset of GIFs may be selected based on previous selections of GIFs by one or more viewers within the atmosphere faction and the frequency with which GIFs are used to express emotions, or based on the popularity of GIFs within a specific viewer set or viewer preferences. In one implementation, viewer preferences may be expressed in a viewer's user profile. In these cases, the GIF recognition engine 319 may query the viewer's user profile to determine if any preferences are specified in the viewer's user profile and use said preferences to identify GIFs targeting emotions associated with the atmosphere faction. In one implementation, after a subset of GIFs has been identified, the GIF recognition engine 319 may automatically select a specific GIF from the identified subset for association with the atmosphere faction. A specific GIF may be selected based on a GIF confidence indicator, where the GIF confidence indicator indicates the number of times a viewer selects a specific GIF to represent an emotion associated with the atmosphere faction. In this extended implementation, the GIF recognition engine 319 may provide options on the user interface to allow viewers of the atmosphere faction to override the automatic selection and customize GIFs for the atmosphere faction. The user interface can be used to render a subset of GIFs identified for the atmosphere faction's mood and includes selection options for each GIF in the subset rendered on the user interface, allowing viewers to customize GIFs by selecting alternative GIFs from the subset, wherein the alternative GIFs selected from the subset are different from the GIFs automatically selected by the mood expression engine for the atmosphere faction. Alternative GIFs selected by one or more viewers are associated with the atmosphere faction and returned to the viewer's client device for rendering alongside the video game content. In one implementation, alternative GIFs with video game content are provided only to viewers of the atmosphere faction who have selected alternative GIFs, while GIFs automatically selected by the GIF recognition engine 319 are presented to the remaining viewers within the atmosphere faction. In another implementation, the option to select alternative GIFs is provided only to viewers of the atmosphere faction.

[0221] In another implementation, instead of the GIF recognition engine 319 automatically selecting GIFs for a specific atmosphere faction, the GIF recognition engine 319 can present a subset of GIFs identified for a particular atmosphere faction on the interactive interface, along with options for the viewer to select one of the GIFs from said subset to associate with that particular atmosphere faction. The viewer's selection is then used to associate the GIF with the specific atmosphere faction. In one implementation, when more than one viewer selects a GIF from the interactive interface and one or more viewers identify more than one GIF for a particular atmosphere faction, the GIF selected by the maximum number of viewers is used to associate with the specific atmosphere faction.

[0222] In one implementation, the GIF for a mood / faction is configured to be rendered on a specific portion of the viewer's client device's display screen. The client device's display screen can be divided into multiple portions (i.e., segments), and specific portions can be identified for rendering the mood / faction GIF. The specific portion of the display screen can be selected based on each viewer's preferences. Each viewer can specify their own preferences for rendering different content (e.g., game content, chat content, GIFs, etc.), and the GIF will be rendered based on the viewer's specified preferences, either for the mood / faction associated with that viewer or for all mood / factions. (See reference...) Figure 26 Details of the individual segments identified on the display screen are provided. In some implementations, a GIF for each identified atmosphere faction is rendered at a specific portion of the display screen specified by each viewer, wherein the specific portion is specified relative to a portion of the display screen in which the viewer's image representation is rendered, wherein the portion of the display screen used to render the GIF can be the bottom, top, right side, or left side of the portion of the viewer's image representation rendered. In alternative implementations, the GIF associated with each atmosphere faction is presented on the viewer's image representation rendered alongside the video game content. For example, GIFs for different atmosphere factions can be presented in a manner similar to... Figure 17 The avatars are presented in a way that appears on the viewer's visual representation, including not only the avatars for each atmosphere faction but also corresponding GIFs. Similar to... Figure 17 In one implementation, the GIF size is scaled to correspond to the size of the atmosphere faction. In one implementation, the GIF size is scaled according to a confidence level associated with each atmosphere faction, where the confidence level of the atmosphere faction is determined by the number of viewers expressing the unique sentiment of that atmosphere faction.

[0223] In one implementation, viewers can express specific emotions differently (i.e., the specific emotion can be expressed by viewers with different intensities). For example, a first viewer might express happiness with a smile, a second viewer might express happiness with a bright smile, and a third viewer might express happiness by jumping up or dancing happily. Similar ranges of the intensity of happy emotions can be expressed via other modal data streams included in the interaction data, such as text data, audio data, GIFs, emojis, or graphic images. Machine learning algorithms analyze the various modal data streams of the interaction data to identify each emotion expressed by viewers with different intensities and cluster viewers into groups based on the identified emotions. A confidence level for each atmosphere faction is determined, where the confidence level indicates the number of viewers expressing the emotion of the atmosphere faction. An emotion recognition GIF is generated for each atmosphere faction. The size of the GIF for each emotion recognition is scaled according to the confidence level determined for the corresponding atmosphere faction, such that the GIF for the emotion recognition of the atmosphere faction with the highest confidence level is scaled to be rendered larger than the size of the GIF with a confidence level lower than the highest confidence level.

[0224] In an alternative implementation, GIFs can be presented alongside avatars associated with a particular mood or faction. These avatars provide a visual representation of the emotions expressed by the viewer by mimicking their facial expressions, while GIFs offer a more intuitive and engaging way to convey those emotions. In this implementation, the avatar can be rendered on, for example... Figure 17 The image of the audience is shown above, and GIFs for different emotions can be displayed on a portion of the screen designated by the audience.

[0225] Because the emotional changes of the audience in each atmosphere faction are correlated with changes occurring in the gameplay of the video game, the GIFs identified and provided for each atmosphere faction are dynamically updated to correspond to the emotional changes of the audience in that atmosphere faction. The updated GIFs are returned to the audience's client devices so that they can be rendered alongside the video game content when an emotional change in the audience is detected in the corresponding atmosphere faction. In some implementations, the GIF selected for each atmosphere faction can be formatted for rendering based on a rendering format specified by the audience. Some formats that can be used to render GIFs include transparent formats, overlay formats, or rendering formats. The scaled, formatted GIFs are forwarded to the audience's client devices for rendering. In one implementation, the GIFs identified for each audience group are presented to the audience of that group, so that the audience only receives the GIFs associated with the group to which each audience member belongs. In an alternative implementation, GIFs for all atmosphere groups are forwarded to the client devices of the audience members in the audience.

[0226] Figure 26This illustration depicts an exemplary screen representation of a representative image of the audience and various segments of a display screen used to render GIFs for different faction-specific identifications in one implementation. The display screen can be segmented according to the audience's preferences to render different video game-related content on their client devices. In the various implementations described herein, only the audience's image representation and various additional content (avatars, GIFs, reaction soundtracks) identified or generated for the video game are shown; in reality, the audience's image representation is rendered in one section, while the remaining sections are used to present the video game content and any other content. As previously mentioned, the sections of the display screen used to render the audience's image representation can be defined according to the audience's preferences, where each audience member provides their own preferences for rendering the audience's image representation. Similarly, GIFs can be rendered in specific sections of the display screen according to the audience's preferences. For example, a GIF can be rendered in a portion defined next to the audience's image representation, wherein the portion of the display screen rendering the audience's image representation can be divided into a central portion 2501-C, a bottom portion 2501-B, a top portion 2501-T, a right portion 2501-R, and a left portion 2501-L, and a GIF for identifying the atmosphere faction of the audience group can be rendered in any of the identified segments. Figure 26 In the example illustrated, the GIFs identified for emotions within the audience are shown rendered at the bottom portion 2501-B of the audience's image representation, which is rendered on a portion of the display screen according to each viewer's preference. The GIFs representing each emotion can be automatically selected by the GIF recognition engine 319, or they can be selected by the viewer from a subset of GIFs.

[0227] Figure 27 This illustration depicts a subset of GIFs representing different emotion recognitions associated with atmosphere factions in one implementation. Figure 27 In the illustration, the type of emotion identified for each atmosphere faction is represented on the left as an avatar or emoji, and the right side shows a subset of GIFs representing emotions typically expressed by the audience, or a subset of GIFs identified by the GIF recognition engine 319. The GIF subset identified for each emotion may include GIFs of people, video clips of people or celebrities, selected movie scenes featuring popular characters, video clips of anime / manga or movie characters, or user-generated content, etc. For example, an audience of a specific atmosphere faction may have previously selected GIFs of specific manga characters to represent one emotion or a different emotion. With the assistance of machine learning algorithms, the GIF recognition engine 319 can determine the audience's preferred use of specific GIFs to express the emotions of that particular atmosphere faction or other emotions. Therefore, the GIF recognition engine 319 can identify and present subsets of atmosphere factions with different emotions based on the audience's preferred usage. Figure 27 An example of this is shown, where GIF recognition engine 319 identifies different subsets of GIFs for each emotion detected in a viewer of a video game. The GIFs in each subset may have been previously used by viewers of the video game or different video games, or may be identified by the GIF recognition engine based on popularity, frequency of use by other users with profiles similar to the viewer's, etc. The subset of GIFs identified for the happy emotion is shown as subset 2501, the subset of GIFs identified for the unhappy emotion (or sadness) is shown as subset 2502, and the subset of GIFs identified for the surprise emotion is shown as subset 2503. Figure 27 The emotions represented in the image are provided as examples, and other emotions can be similarly represented and appropriate subsets of GIFs can be identified.

[0228] Figure 28 This illustration depicts an image representation of the audience in one implementation and a screen reproduction of a subset of GIFs for the audience to choose from, tailored to a specific mood. The subset of GIFs is presented within the interactive interface on the top portion 2501-T of the image representing the audience and includes GIFs identified based on audience preferences, GIF popularity, prior use of the GIFs by audiences within a particular mood faction, or by other audiences from different mood factions. Prior use of the GIFs may include use by one or more audiences within a video game, in other video games, in social media applications, or in other interactive applications or other interactive interfaces. In addition to the subset of GIFs, the interactive interface includes a selection option 2501a presented for each GIF, allowing the audience to select a suitable GIF to represent the mood of their mood faction. The audience uses selection option 2501a to select a GIF from the subset that can be used to represent the mood of their mood faction when presented to an audience within that mood faction or when presented to an audience within the video game.

[0229] Figure 29 This is an exemplary screen reproduction of an image representation of the audience and a set of GIFs identified to represent different emotions expressed by the audience in one implementation. The set of GIFs is automatically identified by a GIF recognition engine 319 and configured to be presented on the bottom portion of the audience's image on an interactive interface 2501, wherein each GIF rendered in the interactive interface 2501 corresponds to a specific emotion identified from the audience's interactive data. In addition to the identified GIFs, the interactive interface 2501 provides an option 2501b for customizing GIFs for each emotion. Figure 29In the illustrated implementation, the custom option 2501b is provided as a checkbox. The implementation is not limited to checkbox options but may also include radio buttons, interactive links, etc. In one implementation, an option to customize the GIF representing a specific emotion for 2501b can be activated, and this option is only provided to viewers who are part of the corresponding atmosphere faction. In this implementation, viewers of each atmosphere faction are allowed to customize GIFs for emotions associated with their own atmosphere faction, but are not allowed to customize GIFs for emotions associated with other atmosphere factions.

[0230] When a viewer selects the option to customize a specific GIF to describe an emotion, a subset of GIFs suitable for that emotion can be provided to the second interactive interface 2901. Figure 29 A second interactive interface 2901 is illustrated, which features a subset of GIFs identified for the happy mood when the option to customize the happy mood (2501b) is selected at the interactive interface 2501. The subset of GIFs provided on the second interactive interface 2901 is randomly selected, or selected based on prior use of such GIFs for expressing emotions by the audience within the atmosphere faction or other audiences or other viewers of the video game or other video game, the popularity of the mood among viewers with profiles similar to those of the audience, automatic selection by the GIF recognition engine 319, and so on. In addition to presenting a subset of GIFs for the audience to choose from, an option may also be provided in the second interactive interface allowing the audience to select their own GIFs. Figure 29 An example of this is illustrated, in which an interactive link 2901a is provided. Interactive link 2901a is configured to provide access to other GIFs on network 200. Selecting an alternative GIF from the second interactive interface 2901 or via interactive link 2901a results in the rendering of the alternative GIF representing the emotion, as well as other GIFs representing other emotions, in user interface 2501.

[0231] In some implementations, the audience's image can be organized based on the emotions expressed by the viewer. In this implementation, selecting a GIF from the interface may cause the GIF to be rendered on a portion of the audience's image corresponding to an atmosphere group associated with the emotion represented by the selected GIF. The selected GIF may be rendered on the audience's image for a predefined period of time and after the predefined period of time has expired; the GIF can be configured to gradually fade away. (Reference) Figures 25 to 29 The various implementations described allow the Emotional Expression Engine 304 to visually render the audience's emotions as GIFs, serving as a substitute for or supplement to expressive personas. Visual representation allows viewers to understand the general atmosphere within a crowd, enabling them to determine the atmosphere faction to which they best fit.

[0232] Figure 30 This illustration depicts the operation of a method for identifying audience emotions and presenting a Graphical Interchange Format (GIF) file representing the identified emotions alongside the content of a video game, in one implementation. The method begins at operation 3002, where interaction data detected from the audience is aggregated. The interaction data can be provided by the audience in different modalities, and an emotion expression engine 304 aggregates and processes the different modal data streams identified from the interaction data. Some of the modal data streams identified from the interaction data may correspond to text content, audio content, emoticons, GIFs, graphic content, etc. Interaction data can also be captured from facial expressions of the audience. Expressions are identified by capturing an image of the audience while they are watching the video game and analyzing the image to identify expressions from different facial features of the audience. Various modal data streams are collected from the audience in real time while they are watching gameplay of the video game, where the gameplay can be live or deferred. Machine learning algorithms are used to aggregate and process the interaction data collected from the audience to cluster the audience into a group. The machine learning algorithms can use unimodal or multimodal methods to process the interaction data. In unimodal approaches, machine learning algorithms generate and train models for each modal data stream identified from the interaction data. Multiple models can be generated, with the number of models corresponding to the number of modal data streams identified from the interaction data. The outputs from the different models are combined. In multimodal approaches, machine learning algorithms use different modal data streams as input to generate and train a single model.

[0233] The method flows to operation 3004, where the output from the trained model is used to cluster viewers into groups based on the emotions expressed by the viewers, where each group corresponds to a unique emotion expressed by the viewers. Once the viewers are identified, the viewers in each group are maintained within the corresponding group unless an explicit request is received from one or more viewers to move from one group to another, or their emotions are more consistent with the other group.

[0234] Once the cluster is complete, the method moves to operation 3006, where a Graphical Exchange Format (GIF) file is identified for each group. A GIF for each emotion is identified by querying a GIF database 321 maintained for one or more video games. GIFs representing different emotions can be identified based on audience preferences, the use of GIFs to represent relevant emotions within the video game, the popularity of GIFs among users, including the audience, etc. In one implementation, when the number of emotions identified from the audience's interaction data exceeds a predefined value (e.g., 4 or 5), the emotion expression engine 304 can select a specific emotion from those emotions to be represented by a GIF. For example, when 8 or 10 emotions are identified from the interaction data, the emotion expression engine 304 can identify the first 4 or 5 emotions for GIF identification. A specific emotion from those emotions can be selected based on the confidence level of each group. The confidence level of a group is determined by the number of audience members expressing the emotion of that group. The GIF identified for each group is associated with the corresponding audience group. In one implementation, GIFs identified for the first four or five emotions are associated with the corresponding groups, while the remaining groups are not represented by GIFs.

[0235] The method terminates at operation 3008, where the identified GIF, along with the video game content, is returned to the viewer's client device for rendering. The identified GIF can be scaled and / or formatted according to the preferences of the viewer group. Formatting can include presentation formatting and rendering formatting. The presentation format can be based on a portion of the display screen where the GIF needs to be presented. The rendering format specifies how the GIF must be rendered, i.e., a transparent format, an overlay format, or a presentation format. In one implementation, the presentation and / or rendering format can be based on each viewer's preference. Therefore, when a GIF is provided for rendering on the viewer's client device, the presentation and / or rendering format specified by each viewer is taken into account, so that the GIF can be presented in the appropriate portion of the display screen in the specified format. The presented GIF provides an overall atmosphere for the viewer group. In some implementations, the GIF presented to each group can correspond to the mood of the group. In other implementations, the GIF for each group is presented on the display screen of the client device.

[0236] The various implementations described herein provide ways to aggregate the emotions of a large audience and use the aggregated emotions to present expressive avatars or GIFs. Expressive avatars or GIFs provide a way to quickly (i.e., almost in real-time) measure the distribution of reactions from a large group of game viewers and allow viewers to compare their reactions with those of peer groups. Avatars and GIFs also allow players to measure feedback on specific gameplay elements of a video game. The output of a model generated from a machine learning algorithm provides new inference outputs based on the current reactions (i.e., fine-tuning the model) according to changes detected in the audience's facial expressions, making it an intuitive way to measure the emotions of an audience. Further advantages will become apparent to those skilled in the art after reviewing the various implementations and methods of this disclosure.

[0237] Figure 31 An exemplary information service provider architecture is illustrated in one implementation of various embodiments that can be used to carry out this disclosure. Information service provider (ISP) 1902 provides numerous information services (also via...) to geographically dispersed users (i.e., players) 1900 connected via network 1950. Figure 1 (Ref. 200). An ISP may offer only one type of service, such as stock price updates, or it may offer multiple services, such as broadcast media, news, sports, games, etc. Furthermore, the services offered by each ISP are dynamic; that is, services can be added or removed at any point in time. Therefore, the ISP providing a particular type of service to a particular individual may change over time. For example, when a user is in her hometown, she may be served by an ISP located near her, and when she travels to different cities, she may be served by different ISPs. The hometown ISP will transfer the necessary information and data to the new ISP, so that the user information “follows” the user to the new city, making the data closer to the user and more easily accessible. In another embodiment, a master-server relationship may be established between a master ISP that manages information for the user and a server ISP that directly interfaces with the user under the master ISP's control. In yet another embodiment, when a client moves worldwide, data is transferred from one ISP to another so that the ISP in a better location to serve the user becomes the ISP providing those services.

[0238] ISP 1902 includes Application Service Provider (ASP) 1906, which provides computer-based services to customers via a network (e.g., including but not limited to any wired or wireless network, LAN, WAN, WiFi, broadband, cable, fiber optic, satellite, cellular (e.g., 4G, 5G, etc.), the Internet, etc.). Software provided using the ASP model is sometimes also referred to as software-on-demand or Software as a Service (SaaS). A simple form of providing access to a specific application (such as customer relationship management) is using a standard protocol (such as HTTP). The application software resides on the vendor's system and is accessed by users using HTML through a web browser, by dedicated client software provided by the vendor, or by other remote interfaces (such as thin clients).

[0239] Services delivered over vast geographical areas often utilize cloud computing. Cloud computing is a type of computing that provides dynamically scalable and often virtualized resources as a service over the internet. Users do not need to be experts in the technical infrastructure supporting their “cloud.” Cloud computing can be divided into different services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). Cloud computing services often provide common business applications online, accessible from a web browser, while software and data are stored on servers. Based on how the internet is depicted in computer network diagrams, the term “cloud” is used as a metaphor for the internet (e.g., using servers, storage devices, and logic) and is an abstract concept of the complex infrastructure it hides.

[0240] Furthermore, the ISP 1902 includes a game processing server (GPS) 1908, which is used by game clients to play single-player and multiplayer video games. Most video games played on the internet operate via a connection to a game server (e.g., a game cloud server). Typically, games use a dedicated server application that collects data from players and distributes that data to other players. This is more efficient and effective than a peer-to-peer arrangement, but it requires a separate server to host the server application. In another implementation, GPS establishes communication between players, and the players' respective gaming devices exchange information without relying on a centralized GPS.

[0241] Dedicated GPS servers operate independently of the client. These servers typically run on dedicated hardware located within a data center, providing greater bandwidth and dedicated processing power. For most PC-based multiplayer games, dedicated servers are the preferred method for hosting game servers. Large-scale multiplayer online games run on dedicated servers, which are often hosted by the software company that owns the game's title, allowing them to control and update the content.

[0242] A Broadcast Processing Server (BPS) distributes audio or video signals to an audience. Broadcasting to a very small audience is sometimes called narrowcasting. The final stage of broadcast distribution is how the signal reaches the listener or observer, and it can travel through the air to antennas and receivers like a radio or television station, or via cable television or wired broadcasting (or “wireless cable”) through workstations or directly from a network. The Internet can also bring radio or television to receivers, especially multicast, which allows for the sharing of signals and bandwidth. Historically, broadcasting has been defined by geographical regions, such as national broadcasting or regional broadcasting. However, with the widespread availability of the fast Internet, broadcasting is no longer defined by geographical conditions, as content can reach almost any country in the world.

[0243] Storage service providers (SSPs)

[1912] offer computer storage space and related management services. SSPs also provide periodic backups and archiving. By offering storage as a service, users can subscribe to more storage as needed. Another major advantage is that SSPs include backup services, so users will not lose all their data in the event of a hard drive failure on their computer. Furthermore, multiple SSPs can have full or partial copies of user data, allowing users to access data efficiently, regardless of their location or the device used to access the data. For example, users can access personal files on their home computer and mobile phone (when the user is on the move).

[0244] Communication providers (1914) offer connectivity to users. One type of communication provider is an Internet Service Provider (ISP), which provides access to the Internet. ISPs use data transmission technologies suitable for delivering Internet Protocol (IP) datagrams (such as dial-up, DSL, cable modems, fiber optics, wireless, or dedicated high-speed interconnects) to connect their customers. Communication providers may also offer messaging services such as email, instant messaging, and SMS. Another type of communication provider is a Network Service Provider (NSP), which sells bandwidth or network access by providing direct backbone access to the Internet. Network service providers can consist of telecommunications companies, data carriers, wireless communication providers, Internet service providers, cable television operators providing high-speed Internet access, etc.

[0245] Data exchange 1904 interconnects several modules within ISP 1902 and via network 1950 ( Figure 1 (Ref. 200) Connect these modules to user 1900 (player, spectator). Data exchange 1904 can cover a small area where all modules of ISP 1902 are close together, or it can cover a large geographic area when different modules are geographically dispersed. For example, data exchange 1904 may include Fast Gigabit Ethernet (or even faster Gigabit Ethernet) in a data center rack, or an intercontinental virtual area network (VLAN).

[0246] User 1900 uses client device 1920 (i.e., Figure 2 The client device (either a player 101 or a viewer 102) accesses the remote service. The client device includes at least a CPU, memory, a display, and I / O. The client device can be a PC, mobile phone, netbook, tablet computer, gaming system, PDA, etc. In one embodiment, the ISP 1902 identifies the type of device used by the client and adjusts the communication method accordingly. In other cases, the client device uses standard communication methods (such as HTML) to access the ISP 1902.

[0247] Figure 32 Components of an exemplary apparatus 2000 that can be used to perform various embodiments of the present disclosure are illustrated. This block diagram illustrates apparatus 2000, which may be combined with or may be a personal computer, video game console, personal digital assistant, server, or other digital device suitable for practicing embodiments of the present disclosure. Apparatus 2000 includes a central processing unit (CPU) 2002 for running software applications and optionally an operating system. CPU 2002 may consist of one or more homogeneous or heterogeneous processing cores. For example, CPU 2002 is one or more general-purpose microprocessors having one or more processing cores. Other embodiments may be implemented using one or more CPUs with a microprocessor architecture particularly suitable for highly parallel and computationally intensive applications, such as processing operations that interpret queries, identify context-dependent resources, and implement and render context-dependent resources in real time in video games. Apparatus 2000 may be local to the player in the game segment (e.g., a game console), or remote to the player (e.g., a back-end server processor), or one of many servers in a game cloud system that uses virtualization to remotely stream gameplay to clients.

[0248] Machine learning algorithm 320 uses analyzer 2040 to analyze the interaction data to identify the different modal data streams contained therein. The identified modal data streams are then processed by machine learning algorithm 320 using a unimodal or multimodal method to generate one or more AI models 320a. The output from AI model 320a is then used to identify the emotions expressed by the audience.

[0249] Memory 2004 stores applications and data for use by CPU 2002. Storage device (e.g., data storage device) 2006 provides non-volatile storage for applications and data and other computer-readable media, and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other optical storage devices, as well as signal transmission and storage media. User input device 2008 transmits user input from one or more users to device 2000. Examples of such devices may include a keyboard, mouse, joystick, touchpad, touchscreen, still or video recorder / camera, tracking device for recognizing gestures, and / or microphone. Network interface 2014 allows device 2000 to communicate with other computer systems via electronic communication networks and may include wired or wireless communication over local area networks and wide area networks (such as the Internet). Audio processor 2012 is adapted to generate analog or digital audio output from instructions and / or data provided by CPU 2002, memory 2004, and / or storage device 2006. The components of the device 2000, including CPU 2002, memory 2004, data storage device 2006, user input device 2008, network interface 2014 and audio processor 2012, are connected via one or more data buses 2022.

[0250] The graphics subsystem 2020 is further connected to the data bus 2022 and components of the device 2000. The graphics subsystem 2020 includes a graphics processing unit (GPU) 2016 and a graphics memory 2018. The graphics memory 2018 includes display memory (e.g., a frame buffer) for storing pixel data for each pixel of an output image. The graphics memory 2018 may be integrated into the same device as the GPU 2016, connected to the GPU 2016 as a separate device, and / or implemented within memory 2004. Pixel data may be provided directly from the CPU 2002 to the graphics memory 2018. Alternatively, the CPU 2002 provides the GPU 2016 with data and / or instructions defining the desired output image, and the GPU 2016 generates pixel data for one or more output images based on the data and / or instructions. The data and / or instructions defining the desired output image may be stored in memory 2004 and / or graphics memory 2018. In one implementation, GPU 2016 includes the ability to generate pixel data of an output image from instructions and data defining the geometry, lighting, shadows, textures, motion, and / or camera parameters of a scene. GPU 2016 may also include one or more programmable execution units capable of executing shader programs.

[0251] The graphics subsystem 2020 periodically outputs pixel data of an image from the graphics memory 2018 for display on the display device 2010. The display device 2010 can be any device capable of displaying visual information in response to signals from the device 2000, including CRT, LCD, plasma, and OLED displays. The device 2000 can provide, for example, analog or digital signals to the display device 2010.

[0252] It should be noted that access services delivered over a wide geographical area (such as providing access to games in current implementations) often utilize cloud computing. Cloud computing is a type of computing in which dynamically scalable and often virtualized resources are provided as a service over the Internet. Users do not need to be experts in the technical infrastructure supporting their “cloud.” Cloud computing can be divided into different services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). Cloud computing services typically provide commonly used applications (such as video games) online and accessible from a web browser, while the software and data are stored on servers in the cloud. Based on how the Internet is depicted in computer network diagrams, the term cloud is used as a metaphor for the Internet and is an abstract concept of the complex infrastructure it conceals.

[0253] In some implementations, a game server can be used to operate a platform for recording video game player duration information. Most video games played over the internet operate via a connection to a game server. Typically, games use a dedicated server application that collects data from players and distributes that data to other players. In other implementations, the video game can be executed by a distributed game engine. In these implementations, the distributed game engine can run on multiple processing entities (PEs), such that each PE executes a functional segment of a given game engine on which the video game runs. The game engine simply treats each processing entity as a computing node. The game engine typically performs a diverse range of operations to execute the video game application and additional services for the user experience. For example, the game engine implements game logic, performs game calculations, physics effects, geometric transformations, rendering, lighting, shading, audio, and additional in-game or game-related services. Additional services may include, for example, messaging, social utilities, audio communication, gameplay replay functionality, help functionality, etc. While game engines may sometimes run on an operating system virtualized by a hypervisor of a specific server, in other implementations, the game engine itself is distributed across multiple processing entities, each of which may reside on a different server unit in a data center.

[0254] According to this implementation, the corresponding processing entity used for execution, depending on the needs of each game engine segment, can be a server unit, a virtual machine, or a container. For example, if a game engine segment is responsible for camera transformations, a virtual machine associated with a graphics processing unit (GPU) can be provided to that particular game engine segment, as it will perform a large number of relatively simple mathematical operations (e.g., matrix transformations). Processing entities associated with one or more higher-powered central processing units (CPUs) may be provided to other game engine segments that require fewer but more complex operations.

[0255] By using a distributed game engine, the game engine possesses elastic computing properties that are not constrained by the capabilities of physical server units. Instead, more or fewer computing nodes are provided to the game engine as needed to meet the demands of the video game. From the perspective of video games and video game players, a game engine distributed across multiple computing nodes is no different from a non-distributed game engine running on a single processing entity, because the game engine manager or supervisor distributes the workload and seamlessly integrates the results to provide the video game output components to the end user.

[0256] Users access remote services using client devices, which include at least a CPU, a display, and I / O. The client device can be a PC, mobile phone, netbook, PDA, etc. In one implementation, the network on the game server identifies the type of device used by the client and adjusts the communication method accordingly. In other cases, the client device uses standard communication methods such as HTML to access applications on the game server over the Internet.

[0257] It should be understood that a given video game or game application can be developed for a specific platform and a specific associated controller device 2024. However, when such a game becomes available via a game cloud system as presented herein, the user can access the video game using different controller devices 2024. For example, a game may have been developed for a game console and its associated controller 2024, while the user may access a cloud-based version of the game from a personal computer using a keyboard and mouse. In this scenario, input parameter configuration can define a mapping from inputs generated by the user's available controller device 2024 (in this case, the keyboard and mouse) to inputs acceptable for the execution of the video game.

[0258] In another example, users can access the cloud gaming system via a tablet, touchscreen smartphone, or other touchscreen-driven device. In this case, the client device and controller device 2024 are integrated into the same device, where input is provided via detected touchscreen input / gestures. For such devices, input parameter configuration can define specific touchscreen inputs corresponding to game inputs in the video game. For example, buttons, directional pads, or other types of input elements may be displayed or overlaid during the gameplay of the video game to indicate locations on the touchscreen that the user can touch to generate game inputs. Gestures (such as swipes in a specific direction or specific touch movements) can also be detected as game inputs. In one implementation, for example, before starting gameplay of the video game, a tutorial instructing the user on how to provide input via the touchscreen for playing the game can be provided to familiarize the user with operating controls on the touchscreen.

[0259] In some implementations, the client device acts as a connection point for the controller device 2024. That is, the controller device 2024 communicates with the client device via a wireless or wired connection to send input from the controller device 2024 to the client device. The client device can then process this input and transmit the input data to the game cloud server via a network (e.g., via a local networked device such as a router). However, in other implementations, the controller device 2024 itself can be a networked device with the ability to transmit input directly to the game cloud server via a network without first transmitting such input through the client device. For example, the controller device 2024 may be connected to a local networked device (such as the router mentioned above) to send data to and receive data from the cloud gaming server. Therefore, while the client device may still need to receive video output from the cloud-based video game and render it on a local display, input latency can be reduced by allowing the controller device (also referred to as the “controller”) 2024 to send input directly to the game cloud server over the network, thus bypassing the client device.

[0260] In one implementation, networked controllers and client devices can be configured to send certain types of input directly from the controller to the game cloud server, and other types of input via the client device. For example, inputs whose detection does not rely on any additional hardware or processing other than the controller itself can be sent directly from the controller to the game cloud server via the network, bypassing the client device. Such inputs may include button inputs, joystick inputs, embedded motion detection inputs (e.g., accelerometers, magnetometers, gyroscopes), etc. However, inputs utilizing additional hardware or requiring processing by the client device can be sent to the game cloud server by the client device. These may include video or audio captured from the game environment, which may be processed by the client device before being sent to the game cloud server. Additionally, input from the controller's motion detection hardware may be processed by the client device in conjunction with captured video to detect the controller's position and movement, which the client device then transmits to the game cloud server. It should be understood that controller devices according to various implementations may also receive data (e.g., feedback data) from the client device or directly from the cloud gaming server.

[0261] It should be understood that the various implementations defined herein can be combined or aggregated into specific implementations using the various features disclosed herein. Therefore, the examples provided are merely some possible examples and are not limited to a greater variety of possible implementations by combining various elements. In some examples, some implementations may include fewer elements without departing from the spirit of the disclosed or equivalent implementations.

[0262] The embodiments of this disclosure can be practiced with various computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and the like. The embodiments of this disclosure can also be practiced in distributed computing environments where tasks are performed via remote processing devices based on wired or wireless network links.

[0263] Although the method operations are described in a specific order, it should be understood that other housekeeping operations may be performed between operations, or the operations may be adjusted so that they occur at slightly different times, or the operations may be distributed throughout the system. As long as the processing of telemetry and game state data for generating the modified game state is performed in the desired manner, the system allows processing operations to occur at various intervals associated with the processing.

[0264] One or more embodiments may also be made into computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device capable of storing data that can subsequently be read by a computer system. Examples of computer-readable media include hard disk drives, network-attached storage devices (NAS), read-only memory, random access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tape, and other optical and non-optical data storage devices. The computer-readable medium may include computer-readable tangible media distributed across network-coupled computer systems, enabling the distributed storage and execution of computer-readable code.

[0265] Although the foregoing embodiments have been described in slightly more detail for the purpose of clarity, it will be apparent that certain variations and modifications may be practiced within the scope of the appended claims. Therefore, the embodiments of the invention are to be considered illustrative rather than restrictive, and are not limited to the details given herein, but may be modified within the scope and equivalents of the appended claims.

Claims

1. A method for identifying a Graphical Interchange Format (GUI) file, the method comprising: Capture interaction data from viewers who participate in watching gameplay of video games; The interaction data captured from the audience in the audience is aggregated, the aggregation including clustering the audience into different groups based on the emotions detected from the audience in the audience, wherein each audience group is associated with a unique emotion and a confidence score, the confidence score corresponding to the number of audience members in the corresponding group expressing the unique emotion; Generate an avatar representing the unique emotion of each group, wherein the expression of the avatar associated with each group is dynamically adjusted to match the changes in the expression of the audience in the corresponding group; as well as Next to the content of the video game, an embodiment representing the unique emotions of different audience groups within the audience is presented.

2. The method of claim 1, wherein aggregating the interactive data includes, Identify one or more modal data streams included in the interaction data. Process the one or more modal data streams to identify the emotions expressed by the viewers watching the video game; and The audience is grouped into groups based on the emotions expressed by the audience, wherein each audience group is associated with a unique emotion, and The one or more modal data streams identified from the interaction data correspond to any or a combination of text data, video data, audio data, chat data, emoticons, emojis, graphic content, or graphics exchange format files collected in real time from the audience watching the gameplay of the video game. The video data captures the facial expressions of different viewers while the audience is watching the video game, and the audio data includes audio content and captures one or more of pitch, amplitude, or duration.

3. The method of claim 2, further comprising: Multiple models are generated and trained using machine learning algorithms, wherein each of the multiple models is trained using data from a specific modal data stream identified from the modal data stream of the interaction data; and The outputs of the multiple models are aggregated to classify the emotions, and the probability of each of the emotions expressed by the audience is determined via the interaction data.

4. The method of claim 2, further comprising: A machine learning algorithm is used to generate and train a model, wherein the model is trained using the modal data stream identified from the interaction data as input, and the output of the model is used to classify the emotions and determine the probability of each of the emotions expressed by the audience via the interaction data.

5. The method of claim 1, wherein the size of the avatar for each unique emotion is scaled based on the confidence score associated with the corresponding audience group, and The confidence score associated with each audience group varies depending on the number of audience members detected in the corresponding audience group.

6. The method of claim 1, wherein aggregating the interaction data comprises: A word cloud is generated and dynamically updated using keywords identified through sentiment analysis of selected interaction data in the interaction data. The keywords in the word cloud capture the emotional state of the audience at each point in time. The keywords from the word cloud are used as input by a machine learning algorithm to generate and train one or more models. The output from the one or more models is used to identify the emotions expressed by the audience and the probability of each of the emotions via the interaction data.

7. The method of claim 1, wherein the expressions of the audience in each group change according to changes occurring in the gameplay of the video game, and the expressions of each avatar associated with the respective group are dynamically adjusted to reflect the changes detected in the expressions of the audience in the respective audience group.

8. The method of claim 1, further comprising: An interactive timeline is generated and presented, representing the intensity of responses to different emotions detected from different audience groups in the video game, wherein the intensity of the response to each of the different emotions in the interactive timeline varies over time according to changes occurring within the gameplay of the video game. The changes in reaction intensity captured in the interactive timeline are linked to specific parts of the gameplay of the video game, which cause changes in reaction intensity of the different emotions detected from the corresponding audience groups, wherein the links allow access to the specific parts of the gameplay of the video game to view the interactions that cause the corresponding changes in reaction intensity.

9. The method of claim 8, wherein the interactive time graph is a line graph with multiple lines, wherein each line corresponds to a specific emotion detected from a specific audience group, and the avatar corresponds to the specific emotion rendered on the corresponding line.

10. The method of claim 1, wherein the gameplay of the video game is streamed in real time, and a record of the gameplay is stored in a gameplay data storage device, and the record is made available for subsequent streaming at a later time. The process involves generating an interactive timeline representing the intensity of emotional responses associated with the audience of the video game, and presenting the interactive timeline along with the content of the video game. The interactive timeline is generated in real-time and stored in the gameplay data storage device along with the gameplay records. When the recording of the gameplay is subsequently streamed at a later time, a new interactive timemap is generated by modifying the interactive timemap to include the intensity of reactions from multiple viewers captured during the streaming of the recording of the gameplay during the later time. The new interactive timemap capturing the intensity of reactions from the multiple viewers is presented along with the content of the video game during the replay of the recording of the gameplay at the later time.

11. The method of claim 1, wherein presenting the avatar includes providing a user interface with a segmentation option and a formatting option for a viewer to select, the segmentation option providing an option to select a segment from a plurality of segments defined on a display screen to render the avatar, and the formatting option providing rendering options to be used when rendering the avatar on the display screen. The formatting options mentioned include any of the following: transparent format, overlay format, or rendering format.

12. The method of claim 1, wherein presenting the avatar comprises, Determine the geographical location of the audience in each audience group; When the audience in each group is associated with a single geographic location and each audience group is associated with a distinctly different geographic location, a map is presented that includes the geographic locations associated with each audience group; as well as The corresponding avatar associated with the corresponding audience group is overlaid on the geographical location identified in the map that is associated with the corresponding audience group.

13. The method of claim 1, further comprising: Identify specific audiences within each group; Capture the reactions of the specific audience in each group during the defined game time; and The reactions of the specific audience captured during a defined game moment are rendered alongside the content of the video game.

14. The method of claim 13, wherein identifying the specific audience within each group comprises, Identify actions planned to occur in the video game, the actions being identified based on the game state of the video game; Machine learning algorithms were used to identify the types of reactions that different viewers exhibited to different actions occurring in the video game. as well as The selection of a specific audience member in each group is based on the specific audience member’s reaction to the different actions, and the selection of the specific audience member includes predictive amplification to capture the reaction of the specific audience member in each group when the action occurs in the video game.

15. The method of claim 13, wherein the specific audience in each group is selected based on the type and number of comments generated by the remaining audience in the respective group relating to the expression of the specific audience in each group, or is selected randomly, or is selected based on the expressive responses provided by the specific audience.

16. The method of claim 1, wherein the interaction data captured from the audience includes reactions to events or actions occurring in the video game, or reverse reactions to the reactions of a particular audience member watching gameplay of the video game.

17. The method of claim 16, wherein the interaction data captured from the audience of a group includes the responses of the particular audience member associated with the group, and wherein aggregating the interaction data captured from the audience member includes aggregating the responses of other audience members in the group who responded to the response of the particular audience member.

18. The method of claim 1, wherein clustering the audience into different groups further comprises providing the audience with the option to move from a first cluster to a second cluster, the option being provided on a user interface rendered alongside the video game content, the movement causing the audience to be dynamically unassociated from the first cluster and dynamically associated with the audience to the second cluster. The dynamic association allows the audience to access the audience's interactions in the second cluster, while the dynamic unassociation prevents the audience from accessing the audience's interactions in the first cluster.

19. A method for identifying a Graphics Interchange Format (GUI) file, the method comprising: Capture interaction data from viewers who participate in watching gameplay of video games; The interaction data captured from the audience is aggregated by performing sentiment analysis on the interaction data, the aggregation including... Identify one or more modal data streams included in the interactive data; Process the one or more modal data streams to identify the emotions expressed by the audience watching gameplay of the video game; The audience is grouped into groups based on the emotions expressed by the audience, wherein each audience group is associated with a unique emotion and a confidence score, the confidence score corresponding to the number of audience members in the corresponding group that expressed the unique emotion; Generate an avatar representing the unique emotion of each group, wherein the expression of the avatar associated with each group is dynamically adjusted to match the changes in the expression of the audience in the corresponding group; as well as Next to the content of the video game, an embodiment representing the unique emotions of different audience groups within the audience is presented.

20. The method of claim 19, wherein the size of the avatar for each unique emotion is scaled according to the confidence score corresponding to the respective audience group, and wherein the identification and processing of the one or more modal data streams included in the interaction data are performed using a machine learning algorithm.

Citation Information

Patent Citations

  • Recognition and feedback of facial and vocal emotions

    CN103514455A