Sensor-based augmentation of live audience and crowd sounds

The system uses sensors and classifiers to generate audio augmentations based on viewer behavior, addressing the lack of sophistication in existing methods by providing personalized and real-time crowd sound enhancements, thereby improving viewer engagement and media content quality.

WO2025171251A1PCT designated stage Publication Date: 2025-08-14DTS INC(US)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/015001
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-08
Filing Date
2025-02-07
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing methods for augmenting audience and crowd sounds in media content lack sophistication in reflecting the diverse and dynamic reactions of individual viewers, failing to provide a personalized and real-time emotional feedback to enhance the viewing experience.

Method used

A method and system that uses sensors to capture audio and video information, classify user behavior through classifiers, and generate procedurally triggered audio augmentations based on these classifications to enhance media content with crowd sounds, thereby simulating a larger audience reaction.

Benefits of technology

The system accurately captures and reflects individual viewer reactions, providing a personalized and real-time audio augmentation that enhances the viewing experience by aligning the user's emotional feedback with a collective audience response, improving engagement and quality of live performances and media content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025015001_14082025_PF_FP_ABST
    Figure US2025015001_14082025_PF_FP_ABST
Patent Text Reader

Abstract

Aspects of the present invention regard a method and system for augmenting audience and crowd sounds related to a streamed or broadcast media content. The method comprises capturing audio and / or video information of a user exposed to the streamed or broadcast media content using one or several sensors; classifying the captured information into at least one label using at least one classifier, wherein each classifier classifies at least one behavioral feature of the user; providing or generating an audio augmentation of the streamed or broadcast media content based on the at least one classified label; and adding the audio augmentation to the media content.
Need to check novelty before this filing date? Find Prior Art

Description

XPD.442 SENSOR-BASED AUGMENTATION OF LIVE AUDIENCE AND CROWD SOUNDS RELATED APPLICATION AND PRIORITY CLAIM

[0001] This application is related to and claims priority to U.S. Provisional Application No.63 / 551,278, filed on February 8, 2024, and entitled “SENSOR-BASED AUGMENTATION OF LIVE AUDIENCE AND CROWD SOUNDS FOR STREAMING AND BROADCAST MEDIA”, which is hereby incorporated by reference in its entirety. FIELD OF THE INVENTION

[0002] This invention describes methods and systems of augmenting crowd or audience sounds broadcast with various audio-visual entertainment scenarios with procedurally generated sounds that are triggered by the viewers’ personal reactions to the content. BACKGROUND

[0003] It is common knowledge that people react more emotionally when surrounded with like-minded people at a live event than when watching the same event alone. For example, people might laugh more at a comedy club, might shout / chant more at a live football match, and might sing along with favorite songs at a concert. Humans are, by default, social animals. As such, they feel heightened levels of physical and emotional arousal when surrounded by other humans in a crowd. The sense of being with one’s tribe at such events tends to create a feedback mechanism that resonates within a crowd and greatly amplifies the emotional impact of the performance. Accordingly, engagement of a viewer with the viewed content is increased when increasing the perception that the viewer is not watching alone, but instead is viewing as part of a larger crowd.

[0004] Prior art approaches to “sweeten” live entertainment broadcasting include augmenting the sounds of football crowds with recordings made in previous games, augmenting sitcoms and live comedy shows using a laugh track, and combining at live concert events the bands track from a front-of-house mixing board with a separate live audience track for maximum emotional impact.

[0005] Despite advances in the field of audio augmentation for media content, there remains a desire for more sophisticated methods and systems that allow to augment audience and crowd sounds.XPD-442 SUMMARY

[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0007] An aspect of the invention provides for a method for augmenting audience and crowd sounds related to a streamed or broadcast media content. The method comprises capturing audio and / or video information of a user exposed to the streamed or broadcast media content using one or several sensors, classifying the captured information into at least one label using at least one classifier, wherein each classifier classifies at least one behavioral feature of the user, providing or generating an audio augmentation of the streamed or broadcast media content based on the at least one classified label, and adding the audio augmentation to the media content.

[0008] Aspects of the invention are thus based on the idea to use sensors to classify individual viewers reactions which are then used to trigger procedurally generated augmentations for streamed or broadcast media content. For classification, the sensor information is ingested by a classification / analyzing process that outputs one or more labels as a result of the classification. The labels trigger audio augmentations of the streamed or broadcast media content, wherein these audio augmentations are added to the video content and provide a feedback to the user of its own behavior and emotion, as the one or more classifiers classify a behavioral feature of the user. The classification may drive a large audience reaction. As an example, if the user is laughing, the audio augmentation may augment the user’s laughing by adding additional laughter. If the user is cheering such as when watching a sports event, the audio augmentation may augment the users cheering by adding additional cheering such that the user has the perception of cheering in a large crowd even though being by himself or in a small group.

[0009] The feature of "capturing audio and / or video information" may refer to the process of recording the user's reactions to the media content through sensors such as microphones and cameras. The feature of "classifying the captured information" may involve analyzing the recorded data to identify specific behavioral features, such as laughter, cheering or jumping up, using classifiers which could be machine learning models. The feature of "providing or generating an audio augmentation" may mean to create additional sound effects that may mimic crowd reactions or amplify the behavior of the user based on the classified user behavior. The feature of "adding the audio augmentation to the media content" may involve integrating these generated sounds with the original media content to enhance the overall experience.XPD-442

[0010] Aspects of the invention thus provide for a sophisticated method that can accurately capture and reflect the diverse and dynamic reactions of individual viewers or user crowds to media content. A real-time, personalized audio augmentation is provided for that is adapted to the specific emotional and behavioral features of each user. Additionally, valuable feedback to performers and content creators may be provided by reflecting audience reactions, thereby improving the quality of live performances and media content.

[0011] A classifier within the meaning of the present invention may be any classification means that classifies or categorizes data into predefined labels (which labels may also be referred to as classes or categories). A classifier may comprise one or several trained neural networks, while this is not necessarily the case and other classification means such as analytic models, classical signal processing or decision trees may be used instead well. According to aspects of the invention, the classification process is implemented using one or several classifiers. Each classifier classifies the information captured by the sensors into one of several predefined labels that refer to a behavioral feature of the user such as a body posture, a sentiment and a laughter of the user. The label or labels are thus the result of the classification and determine the audio augmentation that is generated or provided for.

[0012] The classifiers can depend on the type of event. For example, classifiers for sports crowd sounds may be more influenced by sitting posture, standing, arms raised, team-specific chants, etc. A comedy show classifier may be more focused on audio features such as actual laughter and other laughter-related facial and body pose landmarks. A classifier model used in conjunction with a music concert may use a viewer’s singing voice to trigger a larger crowd sing-along of a specific song or chorus. Sentiment classifiers may be used to trigger procedurally generated crowd sounds which can sweeten the default of pre-existing crowd sounds within the media content.

[0013] Each classifier may be a unimodal classifier or a multimodal classifier. A unimodal classifier uses data from a single modality to make classifications, such as facial data. A multimodal classifier uses data from multiple modalities to make classifications, such as both facial data and voice data.

[0014] The audio augmentation of the streamed or broadcast media content is provided or generated. “Providing” the audio augmentation means that the audio augmentation has been generated or recorded before and is, e.g., stored in a storage space. Depending on the classified label, a specific audio augmentation is then retrieved, e.g., based on a predefined table. “Generating” the audio augmentation means that the audio augmentation is generated in real time depending on the classified label. Both variants may be implemented.XPD-442

[0015] Within the present disclosure, crowd sound refers to noise made by a large group of people, often in an uncontrolled or spontaneous manner. It may be associated with events like sports games or festivals. Audience sound refers to the noise made by a group of people who are gathered to watch or listen to a performance, lecture, or other organized event. The sound may be more controlled and can include subtle audible reactions such as laughter.

[0016] In an embodiment, the audio augmentation consists of a mix of prerecorded sounds triggered by the classified label or labels. Alternatively, a more complex system using triggered granular synthesis may be implemented for a more dynamic and interactive result.

[0017] The sensors that capture the audio and / or video information of a user may include one or several microphones and cameras for traditional screen-based viewing (such as TV, laptop, smartphone, etc.). Also, instead of or in addition to cameras one or several infrared or mmWave sensors may be used to capture the user. The sensors may optionally further include health / biometric sensors which are configured to track features of the user’s health system such as body heat, pupil dilation, heart rate, and electrical activities (EEG), wherein these embodiments of sensors represent examples only.

[0018] Within the meaning of the present invention, the streamed or broadcast media content may be of any type, including live broadcasting or live streaming and broadcasting or streaming of prerecorded or cached content. Any media content may be considered including any live events, sports events, shows, movies, and content of virtual or augmented reality applications. Also, the media content may include audio only such as with a radio broadcast, or may include both audio and video, or may include video only.

[0019] In an embodiment, the at least one sensor captures position related information. In such case, the classification process comprises classifying the captured information into at least one of a number of predefined body gesture labels such as a seating posture, a standing posture, a hands raised gesture, a hands over head gesture, a joining hands gesture, and an idle gesture of the user. Each such different posture / gesture is associated with a specific label.

[0020] In an embodiment, to classify the captured information into a body gesture label, keypoints of the skeleton of the user may be determined. With keypoint determination, specific anatomical points of the human body such as shoulders, hands, knees, etc. are identified. The process is implemented by an algorithm or a trained neural network as is known to the skilled person. A body gesture label is determined in a subsequent classification process that is based on the position and / or movement of such keypoints.XPD-442

[0021] In a further embodiment, the at least one sensor captures facial information of the user. In such case, the classification process comprises classifying the captured information into at least one of a number of predefined facial sentiment labels. In an embodiment, the classification may include extracting a cropped image of the user’s face and analyzing the cropped image to identify at least one of the number of predefined facial sentiment labels. In an embodiment, the cropped image is analyzed to identify which of the number of predefined facial sentiment labels is dominant, wherein the captured information is classified into the dominant label. Trained neural networks may be used both for the cropped image extraction and for the identification of facial sentiments.

[0022] In a further embodiment, the at least one sensor captures voice information of the user. In such case, the classification process comprises classifying the captured information into at least one of a number of predefined voice sound labels such as different kinds of laughter (such as laugh out loud, normal laugh, and giggle). By classifying the user's voice into predefined labels, the system can generate audio augmentations that better reflect the user's actual reactions.

[0023] In an embodiment, the classification may include to derive from the captured information audio feature representations as spectrograms by applying a time-frequency transform algorithm and using the spectrograms to determine the at least one voice sound label. Other audio feature representations may be implemented in other embodiments, e.g., providing raw digital time domain samples.

[0024] In a further embodiment, classifying the captured information into at least one label comprises classifying the captured information into at least one of a number of predefined body gesture labels using a body gesture classifier, and classifying the captured information into at least one of a number of predefined facial sentiment labels using a facial sentiment classifier. In such embodiment, two classifiers (a body gesture classifier and a facial sentiment classifier) are used in parallel. Such embodiment may be implemented in the context of a sports crowd audio augmentation. In other embodiments, other classifiers may be used in parallel.

[0025] In a further embodiment, the at least one sensor captures facial information and voice information of the user. In this embodiment, the classification comprises classifying the captured information into at least one of a number of predefined emotional labels using a multimodal classifier classifying both the facial information and the voice information. Accordingly, a multimodal classifier is used in this embodiment, wherein two different kinds of input (facial and voice) are provided to the classifier.

[0026] In this respect, it is pointed out that within the language used in the present disclosure an “emotional” label is the result of a classification which considers both facial information and voice information. On theXPD-442 other hand, a facial sentiment label is the result of a classification which considers facial information only. A voice sound label is the result of a classification which considers voice information only.

[0027] In embodiments of the multimodal classifier, a cropped image of the user’s face may be extracted from the captured facial information, and an audio feature representations may be extracted from the captured voice information, wherein classifying an emotional label of the user is based on the cropped image and the audio feature representations using the multimodal classifier.

[0028] In such embodiment, it may be provided that the extracted cropped image and the extracted audio feature representations are encoded before the classification process to enable them to be input into the multimodal classifier. The reason for such encoding may the that, due to the multimodal nature of the system, the data from the cropped image and the data from the audio feature representations may not be directly usable for classification as they originate from different data domains (image and audio), each with distinct structural and numerical characteristics. For example, the audio input features could be energy values in different frequency bands. The imaging features may be pixel intensities within a 64x64 face crop. The encoding process acts as a bridge between the audio input features and the imaging features on the one hand and the classification stage on the other hand, ensuring that the classifier receives data in a format that is optimized for processing.

[0029] In a further embodiment, the captured information may be classified into at least one vector of labels, wherein each vector of labels comprises several labels and associated probability values for the labels. This means that instead of merely assigning a single label to the captured information, the method evaluates multiple possible labels and assigns a probability to each, reflecting the likelihood that the captured information corresponds to each label. The rationale of this embodiment lies in that classification may not lead to a clear result in the sense that only one label is output as the result of the classification. Rather, in particular when using trained neural networks as classifiers, the classification may lead to several results or labels, each associated with a specific probability reflecting the likelihood that the label is correct. In such case, the different labels and their probabilities may be included in a vector of labels. This probabilistic approach allows for a more nuanced and accurate classification of user behavior. The audio augmentation may be triggered by the label with the highest probability, or may be triggered by several or all of the labels included in the vector including their respective probabilities. For example, such multilabel classification may be implemented for classifying body movements and for classifying emotions.

[0030] In a still further embodiment, the at least one sensor captures a user’s singing voice. For example, the user is singing while watching a pop concert or a sports event. In such case, the classification includes identifying the song sung by the user, wherein the classified label identifies of the song. In such case, it mayXPD-442 be further provided that the audio augmentation comprises to provides or generate a crowd sound or chorus sound singing the identified song along with the user, thereby triggering a larger crowd sing-along of the specific song sung by the user. In case the actual media content includes the singing of a song, it is safe to assume that the user is singing the same song. In such case, the classifier only needs to detect that the viewer is singing. The audio augmentation in such case could be to simply increase the level of a pre-recorded crowd singalong track.

[0031] In embodiments, the classification process may comprise two steps. In a first step, the captured information is analyzed. For example, input features such as skeleton keypoints or a face crop information are provided for in the first step. In the second step, the classification is carried out based on the input features. For example, a classification may be implemented based on skeleton key points. Both steps may be implemented using trained neural networks. However, in other embodiments, classification may start directly from the raw data captured by the one or several sensors, wherein the captured information is input into the trained neural network, and wherein the trained neural network is configured to classify the user information and output the at least one label. This may depend on the kind of training of the neural network that is used to classify the captured information.

[0032] Classification of the captured information may be implemented at a user side (e.g., on a tablet, TV or set-top box), in particular if user privacy is a concern. At the same time, the results of the classification (the at least one label) may be sent to a remote server, wherein the audio augmentation is performed in the cloud. Further, the audio augmentation (an audio signal) may be added to or mixed with the streamed or broadcast media content by such remote server, wherein the procedurally generated crowd reactions are mixed with the original content source and broadcast or streamed back to the viewer / user. For example, the audio signal may be mixed in the cloud with a default audience or crowd track of the streamed or broadcast media content, wherein the audio signal is included in an audience / crowd track of the streamed or broadcast media content. This approach leverages the computational power and storage capabilities of remote servers to handle the processing and mixing of audio signals, thereby potentially reducing the computational load on local devices.

[0033] Generally, however, adding the audio augmentation to the media content may be implemented in a plurality of manners. In other embodiments, the audio augmentation may be added to the media content at the user side. It may be mixed with other content or added in a separate channel.

[0034] In a further embodiment, the added or mixed audio signal may be provided to a 3D audio renderer of a multichannel or binaural audio rendering system at the user side. This allows for a more natural spatial distribution of the audio content, wherein the 3D audio renderer allows for the precise placement of sounds inXPD-442 a three-dimensional space. For example, a virtual crowd synthesis could make use of such interactive 3D audio renderer.

[0035] In an embodiment, a plurality of audio augmentations are predefined and stored in a storage space, wherein at least one of the stored audio augmentations is provided based on the at least one classified label. For example, the plurality of predefined audio augmentations may be several prerecorded sounds. Depending on the classified label, a specific prerecorded sound or a specific mix of prerecorded sounds may be provided for as augmentation. The corresponding assignment can be stored in a table.

[0036] In a further embodiment, the method further comprises that audio and / or video information is captured for several users. The captured information is then classified into at least one label for each of the several users. These at least one labels for each of the users are then averaged or aggregated. The audio augmentation is provided or generated based on the averaged or aggregated labels.

[0037] This aspect of the invention is thus based on the idea to aggregate labels of different viewers before an aggregated audio augmentation is synthesized for broadcast or streaming. A group label is determined, wherein the group label considers the behavioral features of a plurality of individuals. This allows to align the audio augmentation with the group reaction, thereby bringing each individual in stronger alignment with the group. For example, a group sentiment or group emotion may be determined before audio augmentation. Audio augmentation thus reflects a more representative and collective audience reaction, rather than being based on a single user's behavior.

[0038] In such embodiment, several users may be identified as individual users from the captured information (using person / object detection). For each of the recognized individual users associated captured information is classified into at least one label. The averaging or aggregating of the labels of the different users may include local averaging or aggregating of labels and global averaging or aggregating of labels, wherein with local averaging or aggregating of labels, at least some of the several users are located at the same location, and with global averaging or aggregating of classifiers, at least some of the several users are located remotely at different locations.

[0039] Accordingly, this embodiment considers two variants: local averaging or aggregating of labels and global averaging or aggregating of labels. In the context of local averaging or aggregating, at least some of the several users are situated at the same location, this implying a scenario in which users are in a shared physical space, such as a sports bar or a living room. The local averaging or aggregating process combines these classifications to produce a unified audio augmentation reflective of the group's collective response. Conversely, global averaging or aggregating of labels pertains to users who are located remotely at differentXPD-442 locations. This method involves capturing and classifying the behavioral features of users who are dispersed across various geographical locations.

[0040] In a further embodiment, the audio augmentation may be sent to a content presenter which provides the streamed or broadcast media content, thereby providing a feedback to the content presenter. For example, during the recent pandemic, musicians and comedians noted the difficulty of performing without any kind of audience feedback, which affected their performance. Crowd-sourcing individual viewer sentiment and reflecting the procedurally generated reactions back to the performer can create a symbiotic relationship between the performer and their audience, creating a better experience on each side. The audio augmentation may also help advertisers or programmers to better understand the level of content engagement with an audience. This is often done in film and broadcasting when a small audience is shown a piece of content, and their reactions are monitored. In this case, the program producer may receive feedback on what kind of emotional reaction is garnered for every scene or line of dialogue. In such case, sentiment analysis may also include cues related to boredom or disinterest, such as looking at a mobile phone for an extended period.

[0041] The present method may be implemented in real time such that the audio augmentation is provided to the user as an instantaneous augmentation while the user is exposed to the media content.

[0042] In a further aspect of the invention, a system for augmenting audience and crowd sounds related to a streamed or broadcast media content is provided for. The system comprises: at least one sensor configured to capture audio and / or video information of a user exposed to the streamed or broadcast media content; at least one classifier configured to classify the captured information into at least one label, wherein each classifier classifies at least one behavioral feature of the user; a sound generator configured to provide or generate an audio augmentation of the streamed or broadcast media content based on the at least one classified label; and adding means configured to add the audio augmentation to the media content.

[0043] The term "sensor" refers to devices that capture audio and / or video data from the user, such as microphones or cameras. "Classifier" refers to a system that processes the captured data and assigns it to specific categories or labels based on user behavior, such as laughter, cheering or body gestures. "Sound generator" is a component that produces additional audio effects or sounds to enhance the media content, based on the classified user behavior. "Adding means" refers to any mechanism that combines the generated audio augmentation with the original media content.

[0044] The advantages and embodiments described with respect to the method similarly apply to the system. For example, the system may further comprise a body keypoints extraction unit which is configured toXPD-442 determine keypoints of the skeleton of the user, wherein the at least one classifier is configured to determine a body gesture label based on the position and / or movement of the keypoints.

[0045] In another embodiment, the system may further comprise a facial crop extraction unit configured to extract a cropped image of the user’s face, wherein the at least one classifier is configured to analyze the cropped image to identify which of the number of predefined facial sentiment labels is dominant, and wherein the captured information is classified into the dominant label.

[0046] In still another embodiment, the system may further comprise a time-frequency transform unit configured to derive from the captured information audio feature representations as spectrograms, when the at least one classifier is configured to use the spectrograms to determine the at least one voice sound label.

[0047] In still another embodiment, the system may comprise a multimodal classifier configured to classify the captured information into at least one of a number of predefined emotional labels using both the facial information and the voice information.

[0048] In still another embodiment, the least one sensor is configured to capture audio and / or video information of several users. The at least one classifier is configured to classify the captured information into at least one label for each of the several users. In such embodiment, the system further comprises an aggregation unit configured to average or aggregate the at least one labels of each of the users into averaged or aggregated labels. The sound generator is configured to provide or generate the audio augmentation based on the averaged or aggregated labels.

[0049] In a still further aspect of the invention, a user device for augmenting audience and crowd sounds related to a streamed or broadcast media content is provided for. The user device comprises: a playback device configured to receive the streamed or broadcast media content and further configured to configure the streamed or broadcast media content for playback; and at least one classifier configured to receive audio and / or video information captured by at least one sensor, wherein the audio and / or video information is of a user exposed to the played media content. The at least one classifier is further configured to classify the captured information into at least one label, wherein each classifier classifies at least one behavioral feature of the user. The at least one classifier is further configured to send the at least one classified label to a network based sound generator which is configured to provide or generate an audio augmentation of the streamed or broadcast media content based on the at least one classified label. The playback device is further configured to receive the audio augmentation added to the streamed or broadcast media content.XPD-442

[0050] The user device represents that part of the system which may be implemented at a user location in embodiments. The user device may comprise further components of the system such as a body keypoints extraction unit, a facial crop extraction unit, or a time-frequency transform unit. Embodiments of the user device correspond to the embodiments of the method and system as described above.

[0051] According to a still further aspect of the invention, a computer program product embodied on a non- transitory computer readable medium is provided. The computer program product comprises instructions stored thereon to cause one or more processors to carry out a method in accordance with the present invention. Embodiments of the computer program product correspond to embodiments of the method discussed above.

[0052] According to a still further aspect of the invention, a non-transitory computer-readable medium having executable instructions stored thereon is provided for, wherein, when the instructions are executed by one or more processors, operations in accordance with a method in accordance with the present invention are performed. Embodiments of the non-transitory computer-readable medium correspond to embodiments of the method discussed above. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The drawings are provided to illustrate embodiments of the inventions described herein and not to limit the scope thereof.

[0054] FIG.1A is a flowchart of a method for augmenting audience and crowd sounds;

[0055] FIG.1B shows schematically an exemplary system for augmenting audience and crowd sounds;

[0056] FIG. 2 shows an embodiment of the method of FIG. 1, wherein audio augmentation based on behavioral features of a single user is provided;

[0057] FIG. 3 shows a further embodiment of the method of FIG. 1, wherein audio augmentation based on behavioral features of a plurality of local and remote users is provided;

[0058] FIG. 4 shows an embodiment of a system for sports crowd augmentation, illustrating the input and recognition process involving camera / video input, people recognition, face crop extraction, facial emotion recognition, body pose extraction, and body gesture recognition;

[0059] FIG.5 is an example of real-time extraction of a face crop for sentiment analysis;XPD-442

[0060] FIG.6 is an example of real-time extraction of skeleton keypoints;

[0061] FIG.7 shows an embodiment of a system for audience augmentation, illustrating the input and recognition process involving camera input and video input, people recognition, face crop extraction, body pose extraction, feature encoding and emotion classification;

[0062] FIG. 8 is example flowchart of a method to generate crowd / audience reactions aggregating the reactions of multiple individuals in a crowd;

[0063] FIG.9 illustrate an embodiment of a classification that comprises “Goal”, “Miss”, “Clap”, and “Idle” as a gesture labels;

[0064] FIG.10 illustrate an embodiment of a classification that comprises “Anger”, “Disgust”, “Fear”, “Joy”, “Neutral”, “Sadness”, and “Surprise” as facial sentiment labels;

[0065] FIG. 11 illustrate an embodiment of a laughter classification that comprises “Neutral”, “Laughter”, “Surprised” and “Crying” as different labels; and

[0066] FIG.12 illustrates a gesture label and the corresponding media content. DETAILED DESCRIPTION

[0067] The following description describes various embodiments of methods and systems that augment audience and crowd sounds related to a streamed or broadcast media content. Generally, reference is made to a user which may be any person enjoying media content. The user may be a viewer of audiovisual content or a listener of audio content in embodiments.

[0068] FIG. 1A illustrates the steps of a first method for augmenting audience and crowd sounds that are related to a streamed or broadcast media content. According to step 101, audio and / or video information of a user exposed to the streamed or broadcast media is captured using one or several sensors. The one or several sensors may include one or several audio sensors such as a microphone and one or several video sensors such as a camera, without the sensors being limited to these examples. In other embodiments, sensors measuring health data of the user such as pulse or blood pressure may be further included.

[0069] In step 102, the captured information is classified into at least one label / category. The classification is performed using at least one classifier that classifies at least one behavioral feature of the use. BehavioralXPD-442 features of the users are any features of the user which depend on the reaction of the user on the streamed or broadcast media content. Examples of behavioral features are body pose, facial expression, laughter and singing. The classifier serves to classify the information captured by the at least one sensor into at least one of a number of predefined labels. There may be provided one or several classifier, e.g., a classifier classifying laughter, a classifier classifying facial sentiment, and a classifier classifying body pose. The classifier may be a unimodal classifier or a multimodal classifier.

[0070] It is pointed out that the result of the classification is at least one label. This may be a single label if the classification is ambiguous but also includes that the classification may lead to at least one vector of labels, wherein each vector of labels comprises several labels and associated probability values for the labels if the classification is not unambiguous.

[0071] In step 103, an audio augmentation of the streamed or broadcast media content is provided or generated based on the at least one classified label. Accordingly, the at least one classified label triggers an audio augmentation which is based on the previously determined behavioral feature of the user, thereby enhancing and reinforcing in the user sentiment.

[0072] In step 104, the audio augmentation is added to the media content, which may include mixing the audio augmentation with the original media content.

[0073] FIG.1B shows schematically an embodiment of a system for augmenting audience and crowd sounds related to a streamed or broadcast media content. The system may be configured to carry out the method of FIG.1A. It comprises as main components a user device 10, a remote server 20, a broadcast transmitter system 30, one or several sensors 40, a media playback system 50, and a user 60, wherein instead of a single user a plurality of users may be present.

[0074] The user device 10 comprises a first interface 11, a playback device 13, a second interface 12, and a classifier 14. The first interface 11 is configured to communicate with the remote server 20 and is further configured to receive media content from the broadcast transmitter system 30. The media content may be audiovisual content. Alternative to being broadcast, the media content may be streamed to the user device 10. In embodiments, the first interface 11 may comprise an over-the-air radio broadcast hardware communication module configured to receive an over-the-air radio broadcast signal and a wireless internet protocol hardware communication module configured to receive a wireless internet protocol signal.

[0075] The playback device 13 receives the media content from the first interface 11 and provides the media content in a format suitable to be played back by the media playback system 50, which may include a screenXPD-442 and a speaker. In embodiments, the media playback system 50 may be integrated into the user device 10. The user 60 is exposed to the content provided by media playback system 50. The one or several sensors 40 may include at least a camera and a microphone positioned near the user 60 and configured to capture audio and video information of the user 60 while the user 60 is exposed to the media content provided by media playback system 50. In other embodiments, other user-related information such as pulse and blood pressure may be sensed in addition or instead.

[0076] The information captured by the one or several sensors 40 is provided through the second interface 12 to the classifier 14 of the user device 10. Instead of one classifier 14, there may be several classifiers of a classifying unit. The classifier 14 is configured to receive the sensor captured information and to classify the captured information into at least one label, wherein the classifier 14 classifies at least one behavioral feature of the user 60. The classifier 14 may be an AI-based classifier which is implemented on the user device 10. It is configured to be suitable for real-time processing and low latency and optimized to run on the hardware of user device 10. Alternatively, the classifier 14 may be implemented in the cloud and hosted in cloud servers, wherein in such case classifier 14 sends data to the cloud and receives the results in return.

[0077] The classifier 14 provides the results of the classification, namely, at least one classified label over the first interface 11 to the remote server 20. The remote server 20 may comprise a sound generator 21 which generates an audio augmentation of the original media content based on the at least one classified label. To this end, the classifier labels may be converted into an input format expected by the sound generator's API. The audio augmentation may be mixed in the remote server 20 with the original content and broadcast back to the revised 10. To this end, the remote server 20 may provide the audio augmentation to the broadcast transmitter system 30. Alternatively, the remote server 20 may provide the audio augmentation directly back to the user device 10, wherein the mixing is implemented in the user device 10, e.g., in playback device 13.

[0078] The user device 10 includes one or several processors, a memory, and service applications for execution by the processor(s). One service application may implement the playback device 13. Another service application may implement the classifier 14. Other service applications may implement the interfaces 11, 12.

[0079] Embodiments of the user device 10 of FIG. 1B are shown in FIGs. 4 and 7, wherein the user device may comprise further components such as a body keypoints extraction unit, a facial crop extraction unit, a time-frequency transform unit, and features encoding units, as will be discussed with respect to these figures.

[0080] Accordingly, a system is provided in FIG. 1B which is designed to augment audience and crowd sounds related to streamed or broadcast media content. In embodiments, the classifier 14 utilizes a combinationXPD-442 of imaging and audio input features to classify the user’s behavior into one or more predefined labels. These labels may be derived from ground truth examples of images and audio labeled with the most applicable classes. For instance, in a sports scenario, the classifier may use imaging-derived classifications. Additional biometric features and sensors, such as mmWave presence sensors, can also be included if relevant.

[0081] Alternatively, the system could be designed to produce a continuous output that maps to predefined thresholds between labels. In such case, the thresholds define the labels which are output as the result of the classification. In some embodiments, the sound generator 21 may be integrated with the user device 10, either as a separate unit within user device 10 or by integrating the sound generator with the classification model, eliminating the need for handshaking between applications or APIs. The model could be trained using appropriate crowd or audience sound waveforms as the target output and trained accordingly.

[0082] FIG. 2 shows an embodiment of the method of FIG. 1, wherein additional exemplary details are considered. In step 201, audio and video is captured similar to step 101 of FIG. 1. In step 202 the captured audio and video is converted into input features. Such conversion into input features is an initial step in the classification process. For example, the captured audio and video is converted into skeleton key points of a user or a face crop of the user. Such conversion may include analytic processing or using a trained neural network.

[0083] In step 203, a classification is carried out based on the input features. In the considered example, a sentiment classification is carried out. Accordingly, the captured information is classified into at least one of a number of predefined facial sentiment labels such as neutral, happiness, sadness, surprise, fear, disgust or anger. The sentiment classification may be carried out based on a face crop of the user determined in step 202.

[0084] In step 204, the classification is conversed to a crowd noise engine which may be any sound generator which provides or generates an audio augmentation of media content enjoyed by the user, wherein the audio augmentation is based on the results of the sentiment classification. For example, if the classified sentiment is happiness, the crowd noise may contain laughter. The rendered crowd noise may be mixed with a default audience or crowd tracks in step 206, thereby augmenting already prerecorded audience tracks or crowd tracks. However, this is not necessarily the case. Alternatively, there is no default audience / crowd track and the rendered crowd noise is simply mixed with the regular media content enjoyed by the user.

[0085] FIG. 3 shows a further embodiment of the method of FIG. 1. The method of FIG. 3 is in principle similar to the method of FIG. 2 wherein, however, a plurality of users are considered. More particularly, the situation is considered in which there are a plurality of local viewers 301 and / or a plurality of remote viewers 302. For each of the local viewers 301, a local sentiment classification is carried out in step 303, wherein aXPD-442 local sentiment classifier is determined. To this end, steps similar to steps 201 to 203 of FIG.2 are carried out for each of the local viewers 301. Further, remote sentiment classification 304 is carried out for each of the remote viewers 302. In step 305, the local sentiment classifiers are aggregated in a local sentiment aggregation 305. The local sentiment aggregation 305 is then aggregated together with the remote sentiment classifiers 304 to a global sentiment aggregation 306. Based on the global sentiment aggregation 306, a sentiment to crowd noise conversion is implemented and a crowd noise rendered in step 307. As discussed with respect to FIG.2, the rendered crowd noise may or may not be mixed with default audience / crowd tracks in step 308. In other embodiments, there is only remote sentiment aggregation or only local sentiment aggregation.

[0086] Accordingly, in the embodiment of FIG.3, a distributed rendering system is provided, wherein there sentiments of local viewers at the same location are aggretated. The sentiment of each local viewer may be averaged or aggregated before a group sentiment is sent from the client to a remote server. This sentiment may be further aggregated with other remote clients to a greater global sentiment aggregation which may then be used to procedurally generate a crowd reaction based on the aggregate viewer response.

[0087] In some implementations, data may be aggregated, procedurally generated, and redistributed to one or more groups. For example, a viewer may prefer to be surrounded by a procedurally generated representation of fans of their preferred football team. In another case, a viewer may wish to invite a specific route of remote viewers to a co-watching party and prefers that crowd augmentation is based only from their sub-group. In some other implementations, previously aggregated viewer sentiment may be cached so that the system can still be effective when watching pre-recorded content that is not necessarily watch by a larger crowd in real time, such as with comedy shows.

[0088] FIG. 4 illustrates an example system architecture for sports crowd augmentation. This architecture is designed to enhance the auditory experience of a sports event by analyzing and augmenting crowd sounds based on visual and audio inputs from users.

[0089] The system comprises an input stage which includes a camera / video input 401 that collects visual data from the environment, focusing on individuals within the camera's field of view. The system further employs a people recognition process 402 to identify and isolate individuals within a crowd or scene. Classical object detection routines or a trained neural network may be implemented in this respect.

[0090] Once individuals are identified, the system proceeds to extract in facial crop extraction unit 403 a cropped image of the individual's face, wherein the system may utilize advanced tools to provide for real-time extraction of the face crop. An example are open source machine learning tools such as FaceMesh or OpenCV.XPD-442 This is schematically shown in FIG. 5 which shows a face crop 502 within a section 501 of the audiovisual information which contains the user’s head.

[0091] Face crop extraction unit 403 outputs features which are then used as input in a facial sentiment extraction unit 404 which classifies in real time facial sentiments. The sentiments are categorized into predefined labels. Such sentiment labels may correspond to typical face expressions as shown in Figure 10 and may include anger 1001, disgust 1002, fear 1003, joy 1004, neutral 1005, sadness 1006 and surprise 1007. Facial sentiment extraction unit 404 may be implemented as a trained neural network. For example, an open source machine learning library may be pre-trained on datasets which comprise different facial sentiments. The datasets may include the FERPlus dataset which is an enhanced FER (Facial Expression Recognition) dataset that includes label annotations for images.

[0092] In parallel with facial sentiment recognition, the system also extracts body pose information from the visual input. This involves analyzing the individual's body posture and movements to identify specific gestures. To this end, the captured audio and video is analyzed for skeleton keypoints in body keypoints extraction unit 405 which may be implemented as a trained neural network or by object detection routines. The identification of keypoints 602 of a user’s skeleton 601 is schematically shown in FIG. 6. The keypoint extraction may be implemented using open source tools such as Mediapipe Holistic and is implemented in real time.

[0093] Body keypoint extraction unit 405 outputs features which are then used as input in a body gesture extraction unit 406 which classifies in real time body gestures. The body gestures are classified into predefined labels which are indicated by way of example in FIG. 9 and may include Goal 901 (hand raised), Miss 902 (hands over head), Clap 903 (joining hands), or Idle 904 (not making any meaningful movement). The system may employ a Long Short-Term Memory LSTM neural network model to accurately classify these gestures in real time based on the extracted body pose data.

[0094] The system further comprises a body information decoder 407 which interprets the recognized facial sentiments and body gestures into a format that can be used for further processing or decision-making.

[0095] The exact same procedure may be carried out with respect to other users identified by the video input 401.

[0096] In aggregation unit 408, the classified labels of the different users are averaged or aggregated. If there is a single user, naturally, the classified labels for the individual user only are used. The averaged or aggregated labels are provided to a sound generator 409 which is configured to generate an audio augmentation of the media content to which the user is exposed, wherein the audio augmentation is based on the aggregatedXPD-442 classified labels provided by aggregation unit 408. The audio augmentation is responsible for generating sound effects that correspond to the classified labels. Audio signals are thus synthesized or retrieved that mimic the sounds typically associated with the detected emotions and gestures. For instance, a detected “goal” gesture may trigger the generation of cheering sounds, while a “miss” gesture may result in groaning sounds. The generated audio signals are then mixed with the original audio track of the sports event, creating an enhanced auditory experience for the users.

[0097] Generally, the audio generator 409 serves to generate crowd or audience noises. These noises may be synthesized and mixed with the original program soundtrack. The synthesized track may be used to augment crowd noises already mixed with the content.

[0098] Procedurally generated crowd synthesis may be implemented by triggering and mixing a subset of pre- recorded waveforms corresponding to certain crowd reactions (applause, laughter, booing, etc.). Alternative methods such as granular synthesis can be used. Granular synthesis uses a large set of very short audio samples that can be manipulated and stitched together for more independent and realistic sounding crowd noises. It is also possible to generate these kinds of crowd sounds using a machine learning algorithm that may be trained on real-world crowd recordings, labelled with the associated sentiment class(es).

[0099] In one embodiment, the communication between the aggregation unit 408 and the audio generator 409 may be based on MIDI messages. MIDI messages consists of a status byte which indicates the type of message, followed by up to two data bytes that contain parameters. The parameters are: NOTE STATUS NOTE VELOCITY CHANNEL

[0100] These parameters have the values, data types and purposes as indicated in the following table: MESSAGE Values Data Type Purpose Note Status ‘note_on’ or ‘note_off’ Text To turn the note off / on Note 1-105 Integer numbers Which sound to trigger (each of the notes are assigned with different sounds) Velocity 0-127 Integer numbers How loud the sound will be Channel 1-22 Integer numbers Which MIDI channel should the audio come from Thereby, different kinds of sound that are saved for different purposes can be played. Different kinds of sound are triggered by the classified labels, wherein the respective communication may be through MIDI messages in embodiments. In other embodiments, an API of the sound generator may be directly aligned with theXPD-442 classified labels of the machine learning model. Further, it is pointed out that the MIDI protocol is mentioned as an example only. Other protocols may be used instead to pass control data from the classifier to the sound generator, such as internet protocols including TCP, UDP, etc.

[0101] The system architecture is designed to operate in real-time, ensuring that the audio augmentations are provided instantaneously as the user is exposed to the media content.

[0102] FIG.7 illustrates an example system for audience augmentation. The system of FIG.7 is related to the system of FIG.4, but implements some differences. In particular, it implements a multimodal classifier while the system of FIG.4 implements two unimodal classifiers 404, 406.

[0103] More particularly, the system comprises an input stage which includes a camera / video input 701 and a microphone input 702 that capture visual data and audio data from a scene, focusing on individuals within the camera's field of view. There may be implemented several cameras and several microphones in embodiments. The system further employs a people recognition process 702 to identify and isolate individuals within a crowd or scene. This may involve detecting and focusing on each person within a scene. Classical object detection routines or trained neural networks may be implemented in this respect.

[0104] Once individuals are identified, the system proceeds to extract in facial crop extraction unit 704 a cropped image of the individual's face, wherein the system may utilize advanced tools to provide for real-time extraction of the face crop. Extraction unit 704 is similar to extraction unit 403 of FIG.4. The sentiments are categorized into predefined labels such as happiness, sadness, fear, disgust, anger, surprise and neutral.

[0105] Further, the captured audio is processed in audio feature extraction unit 706 to extract audio features. For example, audio feature representations may be derived as a time-frequency spectrograms applying a classical Short Time Fourier transform (STFT) algorithm.

[0106] The extracted facial features at the extracted audio features are each encoded in a features encoding units 705, 707. As the raw face crops and the audio features cannot be directly used for classification, as they originate from different data domains each with distinct structural and numerical characteristics, both the face crops and the audio features (represented as a time-frequency spectrograms) are encoded into a latent space to enable their aggregation. The features encoding units 705, 707 thus standardize the different data modalities into a unified format, ensuring seamless aggregation.

[0107] The extracted and encoded features are input to an emotion classification unit 708. The emotion classification unit 708 is a multimodal classifier and configured to classify in real time an emotion of the userXPD-442 based on the input of the facial crop extraction unit 704 (i.e., cropped face images) and the audio feature extraction unit 706 (i.e., time-frequency spectrograms). The emotion classification unit 708 may be implemented as a trained neural network. The training procedure for this model may involve the use of synchronized video and audio data extracted from the AMI Corpus dataset. This dataset provides a comprehensive set of training examples consisting of multimodal data, with each modality processed through separate pipelines. The training process encodes and aggregates features to utilize both video and audio cues, enhancing the classification performance. The emotion classification algorithm is structured as a machine learning system, iteratively trained by feeding multiple versions of examples to improve its accuracy and reliability.

[0108] The decode emotions information block 709 provides as output of the classification at least one of a number of predefined emotional labels. Such emotional labels may correspond to typical face expressions as shown in Figure 11 and may include neutral / low reaction 1101, laughter 1102, surprise 1103, and crying 1104, wherein the respective emotions are provided with high probability in view of the fact that audio has also been considered in the classification process. Of course, these emotions are mentioned as an example only. In other embodiments, the emotion classification may be similar to the sentiment classification of Figure 11.

[0109] Subsequently, in decode emotions information block 709, the information from the emotion classification unit 708 is decoded which may involve the interpretation of probabilities for each predefined emotional label which can then be utilized for further processing or decision-making. For example, classifier probabilities may be aggregated using a weighted sum of the individual probabilities to generate one or more audience reactions, as will be discussed in more detail below. In other embodiments, the single classification provided by emotion classification unit 708 is in a format compatible with a sound generator API, in which case block 709 may not be required.

[0110] The exact same procedure may be carried out with respect to other users 703 identified by the input 401, 402.

[0111] In aggregation unit 710, the classified labels of the different users are averaged or aggregated. If there is a single user, naturally, the classified labels for the individual user only are used. The average or aggregated labels are provided to a sound generator 711 which is configured to generate an audio augmentation of the media content to which the user is exposed, wherein the audio augmentation is based on the aggregated classified labels provided by aggregation unit 710. The audio augmentation is responsible for generating sound effects that correspond to the classified labels. Units 710 and 711 correspond to units 408 and 409 of Figure 4.XPD-442

[0112] Based on aggregated reactions (such as provided by aggregation unit 710 of FIG.7), the system may create and encode appropriate control messages for the sound generator. The sound generator processes these messages to generate audio based on the recognized emotions and gestures procedurally. For example, it can adjust the background music dynamically to reflect the crowd’s mood.

[0113] Next, the aggregation of individual sentiments is discussed by way of example with respect to the method of FIG.8. FIG.8 illustrates the process of aggregating individual sentiments from multiple individuals in a crowd. The process contains the steps of data collection 801, data representation / abstraction 802, combining data across individuals 803, normalizing results 804, decoding and interpreting data 805 and generating crowd / audience reactions 806.

[0114] Data collection step 801 involves capturing inputs from a camera or video feed. This step identifies individual faces and bodies in the crowd. For each person, the system extracts features such as facial emotions (e.g., happy, sad, neutral) and body gestures (e.g., clapping, raising hands, standing idle). In this respect, reference is made to the example embodiments of FIGs.4 and 7.

[0115] The second step 802 of data representation / abstraction represents the emotions of each person as a set or vector of probabilities, showing the likelihood of each emotion (e.g., 60 percent happy, 20 percent sad, 20 percent neutral). Additionally, body gestures are represented as labels indicating what action each person is performing (e.g., clapping, raising hands).

[0116] Step 803 of combining data across individuals involves calculating the average emotion across all individuals in the scene. This provides a general sense of how the crowd or audience feels overall. For body gestures, it is counted how many people are performing each gesture. This helps identify the dominant actions in the crowd.

[0117] More particularly, step 803 comprises several substeps. In step 831, it is determined to compute an average across the individual data. If an average shall be computed, an average emotion probability and / or a body gestures count is determined in step 832 as follows. To compute the average facial emotion (^^^^) across all individuals, the equation ^ ^^^^ = 1^ ^ ^^^^^ is used, wherein Eirepresents the facial emotional vector for the ithindividual. Each individual's facial emotional vector Ei may consist of a set of probabilities corresponding to different possible emotions as mentioned before. The vector Eiis a probability distribution over predefined emotional labels / categories, suchXPD-442 as anger, disgust, fear, joy, neutral, sadness and surprise, as shown in FIG.10. An example representation for an individual facial emotional vector is: Ei = [0.10, 0.65, 0.05, 0.08, 0.03, 0.02, 0.07] where each value represents the probability of the individual expressing a specific emotion. E.g., the probability for emotional label “anger” 1001 in FIG.10 is 0.10, for label “disgust” 1002 is 0.65, etc.

[0118] To compute the average body gesture (^^^^) across all individuals, the equation 1 ^ ^^^^ =^ ^ ^^is used, wherein Girepresents the gestural individual. The gestural vector Gi represents an individual's physical movement probabilities, which classify predefined actions such as Goal (both arms raised), Miss (both arms on head), Clap (joining hands), Idle (sitting without gestures), as shown in FIG.9. An example representation for an individual's gestural vector is: Gi= [0.10, 0.65, 0.21, 0.04] where each value represents the probability of the individual performing a specific gesture. E.g., the probability for gesture label “Goal” 901 in FIG.9 is 0.10, for label “Miss” 902 is 0.65, etc.

[0119] The probabilities of emotions are thus directly embedded in the facial emotional vector Ei, and the probabilities of gestures are embedded in the gestural vector Gi. These vectors serve as feature representations for individuals, which are then aggregated across the entire audience to derive an overall crowd emotion and gesture state.

[0120] Referring again to FIG.8, if an average is not computed in step 831, a superposition principle is applied in step 833. Application of the superposition principle involves considering the reactions or emotions of each individual to compute a linear system response that includes all reactions and emotions. For instance, it can be applied in scenarios such as a soccer / football match, where a specific event may trigger contrasting reactions among individuals.XPD-442

[0121] In normalizing results step 804, if needed, the aggregated data are adjusted to show proportions or probabilities. For example, it is calculated what percentage of the crowd is clapping or what percentage is happy. ^ ^ =^^^^^^^ ||^^^^||

[0122] In the decoding and interpretingcommon emotion is determined by identifying the one with the highest average score. Further, the most common gesture is identified by finding the one performed by the majority of individuals: Dominant Emotion=argmax(Enorm) Dominant Gesture=argmax(Gnorm)

[0123] In the step of generating crowd / audience reactions 806, the aggregated data is used to create meaningful outputs. For example: if the majority are clapping, cheering sounds are generated. If most individuals appear sad, a disappointed crowd sound is generated. These outputs can then be sent to other systems or applications for real-time interaction or analysis. ^= ^ "^^^^" ^^ ^^^ ^^^(^^^^^) = "^^^^"" ^!!"" ^^ ^^^ ^^^ = " ^!!""

[0124] The system utilizes these averaged or aggregated probability values to determine the most likely sentiments in the most likely gestures of the crowd. The audio augmentation is generated based on the averaged or aggregated labels.

[0125] In a similar manner, there may also be laughter classification, wherein the average laughter emotion (#^^^) is averaged across all individuals, using the equation: 1 ^ #^^^ =^ ^ #^^^^XPD-442 wherein Li represents the laughter vector for the ithindividual. The laughter vector Li represents an individual's laughter situation, which may include predefined labels such as neutral / no reaction, laughter, surprised, and crying, as shown in FIG.12. An example representation for an individual's laughter vector is: Li = [0.05, 0.75, 0.11, 0.09] where each value represents the probability of the individual performing a specific laughter. E.g., the probability for label “Neutral” 1101 in FIG.11 is 0.045, for label “Laughter” 1102 is 0.75, etc.

[0126] FIG. 12 includes screenshots showing on the left hand side hand a particular situation in a media content 1201 which is watched by the user (here: goal) and showing on the right hand side the respective body gesture 1202 of the user (here: both arms raised indicating a goal) which is to be classified. This illustrates that the behavioural features of the user which are classified refer to the media content to which the user is exposed.

[0127] Generally, it is pointed out that the confidence in recognizing of behavioral features such as gestures, facial sentiments or emotions increases with the number of frames of audio and video that is captured. Feature extraction may be implemented on each frame. Over time, the frame results keep adding to the confidence of the final result. A particular number of frames may be collected before a label is identified. For example, if 15 frames (corresponding, e.g., to 1 second) classify the body gestures feature to be “Goal”, the confidence that the behavior of the user refers to the showing of a goal in the media content is high. In another example, if for consecutive 8 frames the user appears to be happy, the confidence for this label increases and the excitement level sound may be increased.

[0128] Accordingly, in embodiments, a list of history of the events detected in each frames is provided for. For example, a buffer may be provided in which the classified labels are stored. Assuming that the buffer size which determines the number of results that shall be observed for making a decision is 7, the format may be: Frame 1 2 3 4 5 6 7 Face neutral happy Happy happy happy happy happy Classification Body idle idle goal goal goal goal goal ClassificationXPD-442 The final decision for the face classification is thus “happy” in the final decision for the gesture classification is “goal”. Alternate Embodiments and Exemplary Operating Environment

[0129] Many other variations than those described herein will be apparent from this document. For example, depending on the embodiment, certain acts, events, or functions of any of the methods and algorithms described herein can be performed in a different sequence, can be added, merged, or left out altogether such that not all described acts or events are necessary for the practice of the methods and algorithms. Moreover, in certain embodiments, acts or events can be performed concurrently, such as through multi-threaded processing, interrupt processing, or multiple processors or processor cores or on other parallel architectures, rather than sequentially. In addition, different tasks or processes can be performed by different machines and computing systems that can function together.

[0130] The various illustrative logical blocks, modules, methods, and algorithm processes and sequences described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and process actions have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. The described functionality can be implemented in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of this document.

[0131] The various illustrative logical blocks and modules described in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a general purpose processor, a processing device, a computing device having one or more processing devices, a digital signal processor DSP, an application specific integrated circuit ASIC, a field programmable gate array FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor and processing device can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.XPD-442

[0132] Embodiments of the system and method described herein are operational within numerous types of general purpose or special purpose computing system environments or configurations. In general, a computing environment can include any type of computer system, including, but not limited to, a computer system based on one or more microprocessors, a mainframe computer, a digital signal processor, a portable computing device, a personal organizer, a device controller, a computational engine within an appliance, a mobile phone, a desktop computer, a mobile computer, a tablet computer, a smartphone, and appliances with an embedded computer, to name a few.

[0133] Such computing devices can typically be found in devices having at least some minimum computational capability, including, but not limited to, personal computers, server computers, hand-held computing devices, laptop or mobile computers, communications devices such as cell phones and PDA’s, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, audio or video media players, and so forth. In some embodiments the computing devices will include one or more processors. Each processor may be a specialized microprocessor, such as a digital signal processor DSP, a very long instruction word VLIW, or other micro- controller, or can be conventional central processing units CPUs having one or more processing cores, including specialized graphics processing unit GPU-based cores in a multi-core CPU.

[0134] The process actions or operations of a method, process, or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in any combination of the two. The software module can be contained in computer-readable media that can be accessed by a computing device. The computer-readable media includes both volatile and nonvolatile media that is either removable, non-removable, or some combination thereof. The computer- readable media is used to store information such as computer-readable or computer-executable instructions, data structures, program modules, or other data. By way of example, and not limitation, computer readable media may comprise computer storage media and communication media.

[0135] Computer storage media includes, but is not limited to, computer or machine readable media or storage devices such as Bluray discs BD, digital versatile discs DVDs, compact discs CDs, floppy disks, tape drives, hard drives, optical drives, solid state memory devices, RAM memory, ROM memory, EPROM memory, EEPROM memory, flash memory or other memory technology, magnetic cassettes, magnetic tapes, magnetic disk storage, or other magnetic storage devices, or any other device which can be used to store the desired information and which can be accessed by one or more computing devices.XPD-442

[0136] A software module can reside in the RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of non-transitory computer-readable storage medium, media, or physical computer storage known in the art. An exemplary storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an application specific integrated circuit ASIC. The ASIC can reside in a user terminal. Alternatively, the processor and the storage medium can reside as discrete components in a user terminal.

[0137] The phrase “non-transitory” as used in this document means “enduring or long-lived”. The phrase “non-transitory computer-readable media” includes any and all computer-readable media, with the sole exception of a transitory, propagating signal. This includes, by way of example and not limitation, non- transitory computer-readable media such as register memory, processor cache and random-access memory RAM.

[0138] The phrase “audio signal” is a signal that is representative of a physical sound.

[0139] Retention of information such as computer-readable or computer-executable instructions, data structures, program modules, and so forth, can also be accomplished by using a variety of the communication media to encode one or more modulated data signals, electromagnetic waves such as carrier waves, or other transport mechanisms or communications protocols, and includes any wired or wireless information delivery mechanism. In general, these communication media refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information or instructions in the signal. For example, communication media includes wired media such as a wired network or direct-wired connection carrying one or more modulated data signals, and wireless media such as acoustic, radio frequency RF, infrared, laser, and other wireless media for transmitting, receiving, or both, one or more modulated data signals or electromagnetic waves. Combinations of the any of the above should also be included within the scope of communication media.

[0140] Further, one or any combination of software, programs, computer program products that embody some or all of the various embodiments of the system and method described herein, or portions thereof, may be stored, received, transmitted, or read from any desired combination of computer or machine readable media or storage devices and communication media in the form of computer executable instructions or other data structures.XPD-442

[0141] Embodiments of the system and method described herein may be further described in the general context of computer-executable instructions, such as program modules, being executed by a computing device. Generally, program modules include routines, programs, objects, components, data structures, and so forth, which perform particular tasks or implement particular abstract data types. The embodiments described herein may also be practiced in distributed computing environments where tasks are performed by one or more remote processing devices, or within a cloud of one or more devices, which are linked through one or more communications networks. In a distributed computing environment, program modules may be located in both local and remote computer storage media including media storage devices. Still further, the aforementioned instructions may be implemented, in part or in whole, as hardware logic circuits, which may or may not include a processor.

[0142] Conditional language used herein, such as, among others, "can," "might," "may," “e.g.,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or states. Thus, such conditional language is not generally intended to imply that features, elements and / or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or states are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term “or” is used in its inclusive sense and not in its exclusive sense so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list.

[0143] While the above detailed description has shown, described, and pointed out novel features as applied to various embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms illustrated can be made without departing from the scope of the disclosure. As will be recognized, certain embodiments of the inventions described herein can be embodied within a form that does not provide all of the features and benefits set forth herein, as some features can be used or practiced separately from others.

Claims

XPD-442 CLAIMS WHAT IS CLAIMED IS:

1. A method for augmenting audience and crowd sounds related to a streamed or broadcast media content, the method comprising: capturing audio and / or video information of a user exposed to the streamed or broadcast media content using one or several sensors; classifying the captured information into at least one label using at least one classifier, wherein each classifier classifies at least one behavioral feature of the user; providing or generating an audio augmentation of the streamed or broadcast media content based on the at least one classified label; adding the audio augmentation to the media content.

2. The method of claim 1, wherein capturing audio and / or video information of the user comprises capturing body position related information; and classifying the captured information into at least one label comprises classifying the captured information into at least one of a number of predefined body gesture labels.

3. The method of claim 2, wherein classifying the captured information into at least one of a number of predefined body gesture labels comprises determining keypoints of the skeleton of the user, and determining a body gesture label based on the position and / or movement of the keypoints.

4. The method of claim 2 or 3, wherein the number of predefined body gesture labels comprise at least two of a seating posture, a standing posture, a hands raised gesture, a hands over head gesture, a joining hands gesture, and an idle gesture of the user.

5. The method of any of the preceding claims, wherein capturing audio and / or video information of the user comprises capturing facial information of the user; and classifying the captured information into at least one label comprises classifying the captured information into at least one of a number of predefined facial sentiment labels.XPD-442 6. The method of claim 5, wherein classifying the captured information into at least one of a number of predefined facial sentiment labels comprises extracting a cropped image of the user’s face and analyzing the cropped image to identify at least one of the number of predefined facial sentiment labels.

7. The method of any of the preceding claims, wherein capturing audio and / or video information of the user comprises capturing voice information of the user; and classifying the captured information into at least one label comprises classifying the captured information into at least one of a number of predefined voice sound labels.

8. The method of claim 7, wherein classifying the captured information into at least one of a number of predefined voice sound labels comprises deriving from the captured information audio feature representations as spectrograms by applying a time-frequency transform algorithm and using the spectrograms to determine the at least one voice sound label.

9. The method of claim 7 or 8, wherein the voice sound labels include different kinds of laughter.

10. The method of claims 2 and 5, wherein classifying the captured information into at least one label comprises classifying the captured information into at least one of a number of predefined body gesture labels using a body gesture classifier; and classifying the captured information into at least one of a number of predefined facial sentiment labels using a facial sentiment classifier; wherein a sports crowd audio augmentation is provided based on these labels.

11. The method of any of the preceding claims, wherein capturing audio and / or video information about the user comprises capturing facial information and voice information of the user; and classifying the captured information into at least one label comprises classifying the captured information into at least one of a number of predefined emotional labels using a multimodal classifier classifying both the facial information and the voice information.XPD-442 12. The method of claim 11, wherein classifying the captured information into at least one of a number of predefined emotional labels comprises extracting a cropped image of the user’s face from the captured facial information; extracting audio feature representations from the captured voice information; classifying an emotion of the user based on the cropped image and the audio feature representations using the multimodal classifier.

13. The method of claim 12, wherein before classifying the extracted cropped image and the extracted audio feature representations these are encoded to enable them to be input into the multimodal classifier.

14. The method of any of the preceding claims, wherein classifying the captured information into at least one label comprises classifying the captured information into at least one vector of labels, each vector of labels comprising several labels and associated probability values for the labels.

15. The method of any of the preceding claims, wherein capturing audio and / or video information of the user comprises capturing the user’s singing voice, wherein classifying the captured information into at least one label comprises providing a classifier which identifies the song sung by the user, and wherein providing or generating the audio augmentation comprises providing or generating a crowd sound or chorus singing the identified song along with the user.

16. The method of any of the preceding claims, wherein classifying the captured information into at least one label comprises using a trained neural network, wherein the captured information is input into the trained neural network, and wherein the trained neural network is configured to classify the user information and output the at least one label.

17. The method of claim 16, wherein the captured information is first analyzed to determine input features and wherein the input features are input into the trained neural network.

18. The method of any of the preceding claims, wherein providing or generating the audio augmentation comprises providing or generating an audio signal, and wherein adding the audio augmentationXPD-442 to the media content comprises adding or mixing the audio signal with / to the streamed or broadcast media content.

19. The method of claim 18, wherein the audio signal is mixed with a default audience or crowd track of the streamed or broadcast media content.

20. The method of claim 18 or 19, wherein the at least one classified label is classified locally and then provided to a remote server, wherein the remote server provides or generates the audio signal based on the at least one classified label, wherein the audio signal is mixed in the remote server with the original streamed or broadcast media content, and wherein the mix is streamed or broadcast to the user.

21. The method of any of claims 18 to 20, wherein the added or mixed audio signal is provided to a 3D audio renderer of a multichannel or binaural audio rendering system at the user.

22. The method of any of claims 18 to 21, wherein the added or mixed audio signal includes audio related to at least one of singing, laughing, crying, cheering, groaning, and sighing.

23. The method of any of the preceding claims, wherein a plurality of audio augmentations are predefined and stored in a storage space, wherein at least one of the stored audio augmentations is provided based on the at least one classified label.

24. The method of any of the preceding claims, wherein the sensors include at least one of a microphone, a camera and a health sensor.

25. The method of any of the preceding claims, further comprising capturing audio and / or video information of several users; classifying the captured information into at least one label for each of the several users; averaging or aggregating the at least one labels of each of the users; and providing or generating the audio augmentation based on the averaged or aggregated labels.

26. The method of claim 25, wherein capturing audio and / or video information of several users comprises recognizing individual users from the captured information, wherein for each of the recognized individual users associated captured information is classified into at least one label.XPD-442 27. The method of claim 25 or 26, wherein averaging or aggregating the at least one labels of each of the users includes at least one of local averaging or aggregating of labels and global averaging or aggregating of labels, wherein with local averaging or aggregating of labels, at least some of the several users are located at the same location; and with global averaging or aggregating of classifiers, at least some of the several users are located remotely at different locations.

28. The method of any of the preceding claims, further comprising sending the audio augmentation to a content presenter which provides the streamed or broadcast media content.

29. The method of any of the preceding claims, wherein the method is implemented in real time such that the audio augmentation is provided to the user as an instantaneous augmentation while the user is exposed to the media content.

30. A system for augmenting audience and crowd sounds related to a streamed or broadcast media content, the system comprising: at least one sensor configured to capture audio and / or video information of a user exposed to the streamed or broadcast media content; at least one classifier configured to classify the captured information into at least one label, wherein each classifier classifies at least one behavioral feature of the user; a sound generator configured to provide or generate an audio augmentation of the streamed or broadcast media content based on the at least one classified label; adding means configured to add the audio augmentation to the media content.

31. The system of claim 30, wherein the at least one sensor is configured to capture body position related information; and the at least one classifier is configured to classify the captured information into at least one of a number of predefined body gesture labels.XPD-442 32. The system of claim 31, further comprising a body keypoints extraction unit configured to determine keypoints of the skeleton of the user, wherein the at least one classifier is configured to determine a body gesture label based on the position and / or movement of the keypoints.

33. The system of any of claims 30 to 32, wherein the at least one sensor is configured to capture facial information of the user; and the at least one classifier is configured to classify the captured information into at least one of a number of predefined facial sentiment labels.

34. The system of claim 33, further comprising a facial crop extraction unit configured to extract a cropped image of the user’s face, wherein the at least one classifier is configured to analyze the cropped image to identify at least one of the number of predefined facial sentiment labels.

35. The system of any of claims 30 to 34, wherein the at least one sensor is configured to capture voice information of the user; and the at least one classifier is configured to classify the captured information into at least one of a number of predefined voice sound labels.

36. The system of claim 35, further comprising a time-frequency transform unit configured to derive from the captured information audio feature representations as spectrograms, when the at least one classifier is configured to use the spectrograms to determine the at least one voice sound label.

37. The system of claims 31 and 33, wherein the at least one classifier comprises a body gesture classifier configured to classify the captured information into at least one of a number of predefined body gesture labels; and a facial sentiment classifier configured to classify the captured information into at least one of a number of predefined facial sentiment labels; wherein the sound generator is configured to provide or generate a sports crowd audio augmentation based on these labels.

38. The system of any of claims 30 to 37, wherein the at least one sensor is configured to capture facial information and voice information of the user; andXPD-442 the at least one classifier is a multimodal classifier configured to classify the captured information into at least one of a number of predefined emotional labels using both the facial information and the voice information.

39. The system of claim 38, further comprising a facial crop extraction unit configured to extract a cropped image of the user’s face from the captured facial information; and an audio feature representation unit configured to derive from the captured information audio feature representations; wherein the at least one multimodal classifier is configured to classify an emotion of the user based on the input of the facial crop extraction unit and the input of the audio feature representation unit.

40. The system of claim 39, wherein the output of the facial crop extraction unit and the output of the audio feature representation unit are each encoded into a unified format in a respective features encoding unit before being input into the at least one multimodal classifier.

41. The system of any of claims 30 to 40, wherein the at least one classifier is configured to classify the captured information into at least one vector of labels, each vector of labels comprising several labels and associated probability values for the labels.

42. The system of any of claims 30 to 41, wherein the at least one sensor is configured to capture the user’s singing voice, wherein the at least one classifier is configured to identify the song sung by the user, and wherein the sound generator is configured to provide or generate an audio augmentation which comprises a crowd sound or chorus singing the identified song along with the user.

43. The system of any of claims 30 to 42, wherein the at least one classifier comprises a trained neural network, wherein the captured information is input into the trained neural network, and wherein the trained neural network is configured to classify the user information and output the at least one label.

44. The system of claim 43, further comprising an analyzer configured to determine input features of the captured information, wherein the input features are input into the trained neural network.XPD-442 45. The system of any of claims 30 to 44, wherein the sound generator provides or generates an audio signal, and wherein the adding means are configured to add or mix the audio signal with / to the streamed or broadcast media content.

46. The system of claim 45, wherein the adding means are configured to mix the audio signal with a default audience or crowd track of the streamed or broadcast media content.

47. The system of any of claims 30 to 46, wherein the at least one classifier is located at a user’s side, and wherein the sound generator and / or the adding means are implemented at a remote server.

48. The system of any of claims 30 to 47, wherein the at least one sensor is configured to capture audio and / or video information of several users; the at least one classifier is configured to classify the captured information into at least one label for each of the several users; wherein the system further comprises an aggregation unit configured to average or aggregate the at least one labels of each of the users into averaged or aggregated labels; and wherein the sound generator is configured to provide or generate the audio augmentation based on the averaged or aggregated labels.

49. The system of claim 48, wherein the system is configured to recognize individual users from the captured information, wherein for each of the recognized individual users associated captured information is classified into at least one label.

50. The system of claim 48 or 49, wherein at least some of the several users are located at the same location, and wherein the aggregation unit is configured to average or aggregate the local users.

51. The system of any of claims 48 to 50, wherein at least some of the several users are located at different locations, and wherein the aggregation unit is configured to average or aggregate the at least one labels of the users located at different locations.

52. A user device for augmenting audience and crowd sounds related to a streamed or broadcast media content, the device comprising:XPD-442 a playback device configured to receive the streamed or broadcast media content and further configured to configure the streamed or broadcast media content for playback; at least one classifier configured to receive audio and / or video information captured by at least one sensor, wherein the audio and / or video information is of a user exposed to the played media content; wherein the at least one classifier is further configured to classify the captured information into at least one label, wherein each classifier classifies at least one behavioral feature of the user; wherein the at least one classifier is further configured to send the at least one classified label to a network based sound generator which is configured to provide or generate an audio augmentation of the streamed or broadcast media content based on the at least one classified label; and wherein the playback device is further configured to receive the audio augmentation added to the streamed or broadcast media content.

53. A computer program product embodied on a non-transitory computer readable medium comprising instructions stored thereon to cause one or more processors to carry out the method of any of claims 1 to 29.

Citation Information

Patent Citations

  • Determining Audience State or Interest Using Passive Sensor Data

    US20170188079A1

  • Augmenting Content Items

    US20220321944A1

  • Simulating crowd noise for live events through emotional analysis of distributed inputs

    US20220383849A1