Media device and method for providing feedback to a user regarding media content
The media device uses a neural network to process media content and user behavior data, enabling enhanced interaction through personalized feedback, addressing the limitations of existing media devices in user engagement.
Patent Information
- Application Number
- PCT/EP2025/057655
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-20
- Filing Date
- 2025-03-20
- Publication Date
- 2025-09-25
AI Technical Summary
Existing media devices lack advanced interaction capabilities with users, particularly in utilizing Large Language Models (LLMs) to provide personalized and context-aware feedback on media content consumption.
A media device equipped with circuitry that acquires media content and user behavior data, processes it through a neural network, specifically a transformer-based LLM, to generate feedback information, which can be visual, audio, or textual, enhancing user interaction and engagement.
The solution improves user interaction by providing personalized feedback, such as summaries, explanations, additional information, and interactive overlays, thereby enriching the user experience with media content.
Smart Images

Figure EP2025057655_25092025_PF_FP_ABST
Abstract
Description
[0001] MEDIA DEVICE AND METHOD FOR PROVIDING FEEDBACK TO A USER REGARDING MEDIA CONTENT
[0002] TECHNICAL FIELD
[0003] The present disclosure generally pertains to a media device for providing feedback to a user regarding media content output on the media device and a method for providing the feedback. In particular, the media device makes use of artificial intelligence to generate the feedback.
[0004] TECHNICAL BACKGROUND
[0005] User interaction with media devices such as (smart) TVs, tablets, smartphones, etc. has evolved significantly over the years, from passive viewing to active engagement. From the beginnings of black-and-white TVs with manual dials, where viewers were passive consumers of content displayed on the TV screen, to the sleek, technology-enabled smart TV of today that offer a multifaceted and interactive viewing experience and allow new forms of interaction beyond the traditional remote control, such as voice commands, gesture commands, facial expressions commands, etc.
[0006] As technology advanced, so did the expectations of users. With the rise of smart TVs, the Internet, and streaming platforms, the users have gained unprecedented control over what, when, and how they consume media content. Users can now access online media content and services like streaming platforms, digital newspaper, podcasts, social media, e-commerce, and the like. Technology has made media devices more interactive, and the introduction of Large Language models (LLMs) should be no exception.
[0007] LLMs are a type of artificial Intelligence (Al) that can perform multiple natural language processing (NLP) tasks. LLMs are built based on deep learning and are trained on vast datasets containing text from the Internet. LLMs can generate human-like texts, answer questions, translate languages, summarize documents, and engage in contextually relevant conversations. From content generation and recommendation systems to virtual assistants and customer service solutions, LLMs have the potential to impact numerous aspects of daily lives.
[0008] Although there exist digital assistants employed with media devices such as voice assistants like Siri (by Apple Inc.), Alexa (by Amazon.com, Inc.), Google Assistant (by Google LLC), etc., it is generally desirable to further enhance the user interaction with a media device by use of digital assistances taking advantage of the possibilities offered by LLMs. SUMMARY
[0009] According to a first aspect, the present disclosure provides a media device for providing feedback to a user regarding media content output on the media device. The media device comprises circuitry which is configured to: acquire media content data related to media content; acquire user data which is indicative of a user behaviour of the user consuming the media content; input the media content data and the user data to a neural network to generate feedback information; and provide the feedback information to the user.
[0010] According to a second aspect, the present disclosure provides a method for providing feedback to a user regarding media content output on a media device, the method comprising: acquiring media content data related to media content; acquiring user data which is indicative of a user behaviour of the user consuming the media content; inputting the media content data and the user data to a neural network to generate feedback information; and providing the feedback information to the user.
[0011] Further aspects are set forth in the dependent, the drawings and the following description.
[0012] BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Embodiments are explained by way of example with respect to the accompanying drawings, in which:
[0014] Fig. 1 shows a media device including a feedback engine for providing feedback to a user;
[0015] Fig. 2a shows a processing of the feedback engine;
[0016] Fig. 2b is a block diagram of an exemplary configuration of the feedback engine;
[0017] Fig. 3 shows a method for providing feedback to a user;
[0018] Fig. 4 shows a method for data processing for providing the feedback;
[0019] Fig. 5 shows a first use case for the media device;
[0020] Fig. 6 shows a second use case for the media device; Fig. 7 shows a third use case for the media device;
[0021] Fig. 8 shows a fourth use case for the media device;
[0022] Fig. 9 shows a fifth use case for the media device;
[0023] DETAILED DESCRIPTION OF EMBODIMENTS
[0024] Before a detailed description of the embodiments under reference of Fig. 1 is given, general explanations are made.
[0025] In some embodiments, a media device for providing feedback to a user regarding media content output on the media device. The media device comprises circuitry which is configured to: acquire media content data related to media content; acquire user data which is indicative of a user behaviour of the user consuming the media content; input the media content data and the user data to a neural network to generate feedback information; and provide the feedback information to the user.
[0026] The media device can be a smart TV, a tablet, a mobile phone, or any other device configured to output the media content. Further, the media device may comprise an interface to access the Internet. By way of the Internet connection, the media device may stream media content provided on streaming portals or download the media content and store it on a storage of the media device. Further, the media device can be configured to access, via APIs, third party applications.
[0027] The media content can be video media content such as a film / movie, a TV show, a video, or the like. The media content can also be image media content such as a photograph. Further, the media content can be audio media content such as a podcast, an audio book, or the like. Further, the media content can be text media content such as a newspaper, an article, E-Mail, text message, or the like.
[0028] The media content is indicated by the media content data. That is, the media content data represents or is related to the media content.
[0029] The media content data comprises at least one of media content image data, media content text data, and media content audio data. For example, a video media content may usually comprise image data and audio data. A (digital) newspaper may usually comprise text data and (in form of photographs) image data. A podcast may usually comprise audio data.
[0030] “Acquiring the media content data” means that the media content data (including, e.g., media content image data, media content audio data, and media content text data) is either streamed or broadcasted to the media device. Alternatively, the media content data is downloaded to the media device. In other words, the media device is configured to obtain or receive the media content data.
[0031] The user data is indicative of the user behaviour. That is, the user behaviour can be derived from the user data. A user behaviour may comprise an action, a reaction, or an intention of the user. For example, the user behaviour may be a gesture, a gaze, a position, a speech (e.g. a comment, a question, an inquiry, etc.), a mood, and the like of the user.
[0032] The user content data comprises at least one of user image data, user text data, and user audio data. “Acquiring the user data” may include that the media device is configured to capture the user data. For example, the media device may comprise a camera configured to capture an image. The camera then outputs the image user data corresponding to the captured image. The camera may be configured as a RBG camera or event camera (also “event-based camera”). Further, a presence sensor (such as a mmWave sensor which uses short- wavelength electromagnetic waves) may be used to capture user data.
[0033] Further, the media device may comprise a microphone to capture a speech of the user and the microphone then outputs the audio user data corresponding to the speech.
[0034] Furthermore, the media device may comprise input means capturing text data entered by the user in the input means. The input means may be wired to or wirelessly coupled with the media device. The input means can be a remote controller (such as a conventional TV remote controller), gaming pad, touchpad on a screen of the media device, etc.
[0035] In another example, the user data can also be captured beforehand by different device than the media device itself. For example, the user data may be captured by a controller device including a microphone and / or camera for capturing the user data. The controller device stores the captured user data in its storage. Additionally, the controller device may be connected to the Internet and transmit the user data to a server for cloud storage. The media device can be in wired communication or wireless communication with the controller device and receive the user data from the controller device. Additionally, or alternatively, the media device may access the server and retrieve the user data from the server. The media content data and the user data are input to the neural network to generate the feedback information for the user. The neural network is trained to understand how the user is reacting to the media and enables interaction such as an explanation of the media content, providing additional information, statistics, context or even purchasing items presented in the media content.
[0036] The neural network has multimodal capabilities in that it may process text data and image data and generate the feedback information.
[0037] The neural network may be trained to perform NPL, object segmentation, action recognition to enable the interaction between the user and the media device.
[0038] In some examples, the neural network may be an LLM. In further examples, the LLM may be based on a transformer-based neural network architecture as described in A. Vaswani et al., “Attention is all you need,” in 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 2017. Such exemplary neural networks known in the art include BERT, ChatGPT, and the like.
[0039] The feedback information can be provided as visual feedback, audio feedback or text feedback. Further, a combination of different feedback modalities is also possible.
[0040] For example, the feedback information can be output by at least one of the screen of the media device and the speaker of the media device. Further, the feedback information can be output as a graphical overlay over the media content displayed on the media device.
[0041] By using the neural network, the user interaction with the media device is improved as the neural network is particularly suitable to understand the user behaviour with respect to the media content output by the media device.
[0042] In some embodiments, the neural network can be a transformer-based Large Language Model, LLM.
[0043] In some embodiments, the media device is a smart TV.
[0044] In some embodiments, the circuitry of the media device is further configured to input the media content data as media content embeddings and the user data as user embeddings to the neural network.
[0045] The media content embeddings include media content vision embeddings and media content text embeddings. Correspondingly, the user embeddings include user vision embeddings and user text embeddings. Generally, text embeddings are a numerical representation of text, used for measuring semantic similarities and aiding in context-based Al interaction. For example, text embeddings are vectors or arrays of numbers which represents the meaning text data. The text embeddings are then used by the neural network.
[0046] Correspondingly, vision embeddings are numerical representations of images that encode the semantics of in image data. For example, vision embeddings are vectors or arrays of numbers which represents the meaning and the context of the image data.
[0047] The media content embeddings are generated based on the media content data. More specifically, the media content vision embeddings are generated based on the media content image data and the media content text embeddings are generated based on the media content text data and / or on the media content audio data.
[0048] Correspondingly, the user vision embeddings are generated based on the user image data and user text embeddings are generated based on the user text data and / or on the media content text data.
[0049] The media device may comprise a vision encoder which is configured to generate the vision embeddings (i.e. the media content vision embeddings and the user vision embeddings) based on the image data (i.e. the media content image data and the user image data, respectively).
[0050] The media device may comprise a text embedder which is configured to generate the text embeddings (i.e. the media content text embeddings and the user text embeddings) based on the text data (i.e. the media content text data and the user text data, respectively).
[0051] Further, the media device may comprise a speech-to-text encoder which is configured to convert speech included in audio data into text data which can be subsequently fed to the text embedder to generate text embeddings. More specially, speech included in the media content audio data may be converted into media content text data by the speech-to-text encoder. The converted media content text data can then be input to the text embedder which generates media content text embeddings based on the converted media content text data. That is, the media content text embeddings are generated based on the media content audio data.
[0052] Correspondingly, the user text embeddings can also be generated based on the user audio data as described with respect to the processing of the media content audio data.
[0053] In some embodiments, the circuitry of the media device is configured to generate the user embeddings based on the user data. That is, processing of the user data is performed on the media device. Thus, user data which is user specific and therefore corresponds to privacy sensitive data is processed locally on the media device. Thereby, it is possible to comply with data protection regulations such as the General Data Protection Regulation (GDPR) (European Regulation 2016 / 679).
[0054] In some embodiments, the circuitry of the media device is configured to obtain the media content embeddings from on a server external to a network comprising the media device. That is, processing of the media content data can be performed on the server or in the cloud computing environment.
[0055] Further, (pre-)processing such as object segmentation and / or action recognition on the media content data or media content embeddings can be performed on the server or in the cloud computing environment. To this end, the neural network may comprise an interface which is connected to modules configured for object segmentation and / or action recognition which are located on the server or in the cloud computing environment, i.e. external to the media device.
[0056] Since the media content data does not comprise user specific information (such as name, age, etc. of the user, and the like), processing of the media content data and / or media content embeddings can be performed on the server or the cloud computing environment without violating data protection regulations. Thereby, the processing load of the media device may be reduced by outsourcing data processing of the media content data and / or media content embeddings.
[0057] In some embodiments, the media device may comprise the camera and / or the microphone for acquiring the user data.
[0058] In some embodiments, the media content data may comprise at least one of the media content image data, the media content text data, and the media content audio data. Additionally, or alternatively, the user data may comprise at least one of the user image data, the user text data, and the user audio data.
[0059] In some embodiments, the feedback information may comprise at least one of a summary of the media content, the explanation of the media content, supplementary information regarding the media content and the graphical overlay on the media content.
[0060] In some embodiments, the media content data may include meta data regarding the media content. The media content meta data can be part of the media content image data, the media text data, or the media content audio data. The meta data may be provided by the media device or retrieved from the Internet. For example, content text data may include the meta data. In some embodiments, the method for providing feedback to a user regarding media content output on a media device, comprises: acquiring media content data related to media content; acquiring user data which is indicative of a user behaviour of the user (110) consuming the media content; inputting the media content data and the user data to a neural network to generate feedback information; and providing the feedback information to the user.
[0061] The method may be carried out with the media device comprising the circuitry discussed herein.
[0062] In some embodiments, the neural network may be a transformer-based Large Language Model, LLM. In some embodiments, the media device is a smart TV. In some embodiments, the method may further comprise inputting the media content data as media content embeddings and the user data as user embeddings to the neural network. In some embodiments, the method may further comprise generating the user embeddings based on the user data. In some embodiments, the method may further comprise obtaining the media content embeddings from on a server external to a network comprising the media device. In some embodiments, the user data may be generated by a camera and / or a microphone of the media device. In some embodiments, the media content data may comprise at least one of media content image data, media content text data and media content audio data, and / or the user data comprises at least one of user image data, user text data and user audio data. In some embodiments, the feedback information may comprise at least one of a summary of the media content, an explanation of the media content, supplementary information regarding the media content and a graphical overlay on the media content. In some embodiments, the media content data may include meta data regarding the media content.
[0063] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer- readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.
[0064] Returning to Fig. 1, a media device 100 is shown which includes a feedback engine 200, a camera 105 and a media device 100, e.g. a TV. Further, a user 110 is shown which consumes media content output by the media device 100. The media content can comprise videos, images, radio, podcasts, or the like. The media device 100 comprises a screen 101 and / or a speaker 103 to output the media content.
[0065] The camera 105 and the microphone 107 are configured to capture user data 203 regarding the user’s behaviour with respect to the media content. For example, the camera 105 and the microphone 107 are configured to capture actions, reactions, and intentions of the user 110 consuming the media content.
[0066] The camera 105 is configured to capture image data regarding the user 110 (user image data 203a). For example, the camera 105 is configured to capture a user’s gesture and to provide user image data indicative of the user’s gesture. The user’s gesture may a be hand gesture (e.g. pointing to a position or object on the screen of the media device 100), head gesture (e.g. nodding) and the like. The camera 105 is further configured to capture a gaze or a gaze direction of the user 110 and based on the captured image data indicative of the gaze or the gaze direction, it can be determined to which portion of the screen of the media device 100 the user’s gaze is directed.
[0067] The microphone 107 is configured to capture audio data regarding the user 110 (user audio data 203c) and media content audio data 201c output by the media device 100. For example, the microphone 107 may capture, as user audio data, speech.
[0068] Further, the media device 100 may comprise input means 109 configured to receive input by the user 110. The input means 109 can be integrated into the media device 100 such as touch pad. In another example, the input means 109 can be remotely arranged from but coupled to the media device 105 such as a remote controller, keyboard and mouse, gaming controller. The user 101 may enter user text data 203b via the input means 109.
[0069] The media device 100 and the feedback engine 200 are connected or have access to the Internet 120. By this, media content can be streamed on or downloaded to the media device 100. Further, the Internet connection enables the media device 100 and the feedback engine 200 access data bases, use cloud storages, use search engines (e.g. Google, Bing, etc.), etc.
[0070] Fig. 2a schematically shows the processing of input data of the feedback engine 200.
[0071] The feedback engine 200 is a multimodal Al solution based on (not only verbal) language understanding, object segmentation, action recognition and LLM-based language processing that allows the interaction between the user 110 and the media device 100 based on the saved or ongoing media content. In case the media content is TV program content, it is also possible to provide interaction based on other TV programs being broadcast on channels which are not displayed, if the user 110 wants to be aware of what is happening in two, or more different channels at the same time.
[0072] As will be described later in more detail, the feedback engine 200 analyses the comments, questions, and reactions that the user 110 has with respect to the media content, namely user input. The feedback engine then analyzes the context of the user input with the media content and complements the media content with data available in databases, on the Internet 120, etc., to provide accurate human-like responses to the user input, namely feedback.
[0073] In order to provide feedback information 205, the feedback engine 200 receives media content data 201 and the user data 203 as input data and processes the input data to generate feedback information 205 as feedback. That is, the feedback engine 200 analyses the media content (indicated by the media content data 201) and the actions, the reactions, and the intentions of the user 110 (indicated by the user data 203) to generate the feedback information 205.
[0074] The media content data 201 is indicative of the media content to be output on the media device 100. The media content data 201 may comprise at least one of media content image data 201a, media text data 201b, and media content audio data 201c. Further, media content data 201 can further comprise meta data regarding the media content (content meta data).
[0075] With the media content data 201, the feedback engine 200 may be able to understand what is output on the media device 100 and may provide context regarding the media content data 201 for a user’s query. That is, the feedback engine 200 is configured to analyse the content data 201 to obtain descriptive text information regarding the media content.
[0076] The user data 203 is indicative of behaviour of the user 110. The query (or an inquiry, a question, or the like) of the user 110 regarding the media content can be determined based on the user data 203. The query is or corresponds to a textual description of the behaviour of user 110 which is input to a later described LLM 260 for providing the feedback information 205.
[0077] As described above, the user data 203 may comprise the user image data 203a and the user audio data 203c captured with the camera 105 and the microphone 107, respectively. Further, the user data 203 may also comprise the text input by the user 110 (user text data 203b) through the input means 109.
[0078] The feedback engine 200 merges the media content data 201 and the user data 203 to understand how the user 110 is reacting to the media content output on the media device 100. Then, the feedback engine 200 enables interaction such as explanation of what is happening, providing additional information, statistics, context or even purchasing items indicated in the media content. With help of the input (i.e. the media content data 201 and the user data 203), the feedback engine 200 can utilize LLM sources such as ChatGPT or similar to provide the feedback information 205. The feedback information 205 may comprise at least one of audio feedback information, text feedback information and image feedback information.
[0079] For example, the feedback information 205 can be provided as audio feedback, provided by the media device 100 (e.g. the speaker of the media device 100) or by a connected device chosen by the user 110.
[0080] Further, the feedback information 205 can be provided as text displayed on the media device 100 or the connected device chosen by the user 110.
[0081] Furthermore, the feedback information can be provided as a video, displayed on the media device 100 or the connected device. The feedback information 205 can also be provided as overlaying graphics on the media content displayed on the media device 100 (for example as a statistics graph superposed over the displayed media content or displaying arrows pointing at something on the displayed media content).
[0082] The feedback information 205 may be provided as a mixture of at least two of the audio feedback, the text, the video, and overlaying graphics.
[0083] The connected device chosen by the user 110 may be, for example, a speaker, a tablet, a mobile phone, or any other media content output device.
[0084] Fig. 2b shows a block diagram of an exemplary configuration of the feedback engine 200. The feedback engine 200 comprises a speech-to-text encoder 210, a tokenizer 220, a text embedder 230, a vision encoder 240, a conversion component 250 and a LLM 260. The feedback engine 200 is configured to process the media content data 201 and the user data 203. By way of example, processing of the media content data 201 is explained with respect to Fig. 2b. The processing of the user data 203 is performed correspondingly.
[0085] The speech-to-text encoder 210 is configured to convert spoken language into text form. That is, the speech-to-text encoder 210 is configured to receive the media content audio data 201c indicative of spoken language and to convert the media content audio data 201c into text data expressed as a sequence of text. Configurations of the speech-to-text encoder 210 are known to the skilled person and may comprise, for example, an attention-based encoder-decoder framework or a transducer framework. The tokenizer 220 is configured to receive, as input, the media content text data 201b. Further, the input to the tokenizer 220 may also the converted text data output by speech-to-text encoder 210.
[0086] The tokenizer 220 is configured to convert the media content text data 201b into tokens. That is, the tokenizer 220 is configured to break a text sequence comprised in the media content text data 201b into smaller parts, i.e. tokens. For example, the sentence “One ring to rule them all” comprised in the media content text data 201b is tokenized into the individual words “one”, “ring”, “to”, “rule”, “them”, “all”. The output of the tokenizer 220 is a sequence of tokens that represents the text included in the media content text data 201b. In other examples, the tokens can additionally comprise sub words, signs, and other small units, which collectively represents the sentence.
[0087] The text embedder 230 is configured to receive, as input, the tokens generated by the tokenizer 220 based on the media content text data 201b and is to generate (output) text embeddings indicative of the media content text data 201b. The text embeddings are then used by the LLM 260 as described later.
[0088] The vision encoder 240 is configured to receive the media content image data 201a. The vision encoder 240 is configured to generate (output) vision embeddings based on the received media content image data 201a. The vision embeddings are then used by the LLM 260 as described later.
[0089] The conversion component 250 is configured to receive, as input, the vision embeddings generated by the vision encoder 240, and to convert the vision embeddings into a format that the LLM 260 can understand and process. Thereby, the LLM 260 can jointly process the (converted) vision embeddings and the text embeddings. For example, the conversion component 250 is configured to transform the vision embeddings, i.e. convert the output of the vision encoder 240 into the format that the LLM 260 can understand, to align modalities such that the vision embeddings and text embeddings are compatible for joint processing by the LLM 260. In some example, the conversion component 20 is configured to resample the vision embeddings to ensure that vision embeddings have a common fixed length, making them suitable for further interactions with the LLM 260 (ensuring so called “Fixed-Length Representations”).
[0090] The LLM 260 is configured to receive, as input, the text embeddings generated by the text embedder and the output of the vision encoder 240 in order to generate the feedback data 205. In other examples, the LLM 260 can also configured as any other transformer-based neural network. Fig. 3 shows a method 300 for providing feedback to the user 110 based on the media content displayed on the media device 100 and behaviour of the user 110. The method 300 is carried out by the media device 100 and more specifically by the feedback engine 200.
[0091] In step 301, the media device 100 obtains the media content data 201 of the media content which is output on the media device 100.
[0092] In step 302, the media device 100 obtains the user data 203 by using the camera 105, the microphone 107 and / or input means 109.
[0093] In step 303, the feedback information 205 is generated by inputting the media content data 201 and the user data 203 to the LLM 260. To this end, the obtained media content data 201 is processed by the speech-to-text-encoder 210, the tokenizer 220, the text embedder 230, the vision encoder 240 and the conversion component 250. The generated output (the text embeddings and converted image embeddings) are fed to the LLM 260 such that the LLM 260 (and therefore the feedback engine 200) understands the media content output on the media device 100, wherein the media content is indicated or represented by the media content data 201. Likewise, the user data 203 is processed by the speech-to-text-encoder 210, the tokenizer 220, the text embedder 230, the vision encoder 240 and the conversion component 250. The generated output (the text embeddings and converted image embeddings) are fed to the LLM 260. The LLM 260 is configured to understand the behaviour of the user 110 and how this relates to the media content.
[0094] That is, in step 303, the LLM 260 merges the obtained media content data 201 and obtained user data 203 to understand the user’s behaviour with respect to the media content in order to generate the feedback information 205. Here, the query based on the user data 203 is input to the LLM 260 using the context determined based on the media content data 201.
[0095] In step 304, the media device provides the feedback information 205 to the user 110. The feedback information 205 is provided as at least one of an audio feedback, visual feedback, and text feedback.
[0096] Fig. 4 shows a method 400 for processing the media content data 201 and the user data 203 for providing feedback to the user 110. The processing is described with the with respect to the components of the feedback engine 200 as shown in Fig. 2
[0097] In step 401, the media content data 201 included in the media content is obtained. More specifically, the media content data 201 includes at least one of the media content image data 201a, the media content text data 201b and the media content audio data 201c. In step 403, content vision embeddings are generated based on the media content image data 201a. To this end, the media content image data 201a is input to the vision encoder 200 which is configured to generate vision embeddings based on image data.
[0098] In optional step 403a, the content vision embeddings are converted to a format which the LLM 260 can process. To this end, the contend vision embeddings are input into the conversion component 250 which is configured to convert the content vision embeddings to obtain converted content vision embeddings. The presence of step 403a depends on the capability of the LLM 260. If the LLM 260 is able to process text embeddings and vision embeddings without any conversion of the vision embeddings, step 405a is omitted.
[0099] In step 405, content text embeddings are generated based on the media content text data 201b and / or the media content audio data 201c. To generate text embeddings based on the media content text data 201b, the media content text data 201b is input into the tokenizer 203 which generated tokens. The tokens are then input into the text embedder 230 which is configured to generate text embeddings based on the tokens. To generate text embeddings based on the media content audio data 201c, the media content audio data 201c is input to the speech-to-text-encoder 210 which is configured to detect speech included in the media content audio data 201c and to output the detected speech as text in form of text data. The text data output by the speech-to-text encoder 210 is then processed by the tokenizer 220 and the text embedder 230 to generate text embeddings as described with respect to the media content text data 201b. The content embeddings are input to the LLM 260 as described later with step 417.
[0100] In step 411, the user data 203 indicative of a user’s behaviour is obtained. The user data 203 includes at least one of the user image data 203 a, the user text data 203b and the user audio data 203c.
[0101] In step 413, content vision embeddings are generated based on the user image data 203a. To this end, the user image data 203a is input to the vision encoder 200 as described in step 403 with respect to the content vision embeddings. In optional step 413a, the user vision embeddings are converted by the conversion component 205 as described in step 403a with respect to the content vision embeddings. In step 415, content text embeddings are generated based on the user text data 203b and / or the user audio data 203c as described in step 405 with respect to the media content text data 201b and the media content audio data 201c.
[0102] In short, the user data 203 is input to the feedback engine 200 as described in steps 403, 403a, 405 with respect to the media content data 201. Therefore, the above-described processing of the media content data 201 applies correspondingly to the processing of the user data 203. In step 417, the (optionally converted) content vision embeddings and the content text embeddings are input to the LLM 260 to generate the feedback information 205. The LLM 260 may analyse the media content. For example, the LLM 260 may be configured to perform at least one of language understanding, object segmentation, action recognition and LLM-based language processing on the user data 203 for understanding the media content.
[0103] Further, in step 417, the (optionally converted) user vision embeddings and the user text embeddings (which are user input embeddings) are input to the LLM 260. By this, the LLM 260 is configured to analyse the user data 203. To this end, the LLM 260 is configured to perform at least one of language understanding, object segmentation, action recognition and LLM-based language processing on the user data 203.
[0104] The LLM 260 is configured to merge the media content data 201 (more specifically, the content vision embeddings) and the user data 203 (more specifically, the user text embeddings) to understand the user’s behaviour with respect to the media content and to generate the feedback information 205. Therefore, the LLM 260 generates the feedback information 205 based on the media content data 201 and the user data 203.
[0105] In step 419, the feedback information 205 provided to the user 110. To this end, the LLM 260 generates the feedback information 205 as output. The feedback information 205 is then provided as at least one of an audio feedback, visual feedback, and text feedback to the user 110.
[0106] In some examples, in order to comply with data protection regulations, processing of the user data 203 (as in steps 411, 413, 413a, 415 and 419) is performed on the media device 100 which comprises the feedback engine 200. Since the media content data 201 of the media content does not usually comprise user specific information (such as name, age, etc. of the user 110), processing of the media content data 201 (as in steps 401, 403, 403 a and 405) can be performed on a server or a cloud computing environment. That is, (pre-)processing such as object segmentation and / or action recognition on the media content data 201 can be performed on the server or in the cloud computing environment. In this case, the LLM 260 may comprise an interface which is connected to modules configured for object segmentation and / or action recognition which are located on the server or in the cloud computing environment, i.e. external to the media device 100. The resulting information obtained by processing the content data 203 can then be used for all end-users in the same and non-personalised format.
[0107] The above processing of the media content data 201 and the user data 203 can be performed in parallel or sequentially. The outcome (i.e., embeddings) are then input together to the LLM 260, where the user query (derived from the user data 203) is analysed in context of the media content (derived from the media content data 201).
[0108] In the following, several use cases for the media device 100 are described with respect to Figs. 5 to 9.
[0109] In a first example scenario as shown in Fig. 5, the user 110 is watching news 500 (as the media content) on a news channel on the media device 100 but needs to attend to his family for a time window. Therefore, the user 110 misses the media content during the time window. The user 110 asks the media device 100 to summarize the news 500 which was output by the media device 100 during the missed time window. The media device 100 is configured to retrieve the media content data 201 output during the missed time window and to provide, as the feedback information 205, a summary 501 of the missed content to the user 110. The summary 501 can be provided, e.g., via the speaker of the media device 100.
[0110] In a second example scenario as shown in Fig. 6, the user 110 is watching a movie 600 (as the media content) on the media device 100 and sees a character 601 with a yellow shirt in a scene. The user 110 voices a question about the character 601 to the media device 100 such as “Who is this person with the yellow shirt?”. The media device 100 is configured to understand this question and to identify the character 601 with the yellow shirt shown in the scene. The feedback engine 200 is configured to recognize and to identify the face of the character 601 and to provide, as the feedback information 205, an actor information 603 to the user 110 as to who the actor playing character is and, optionally, additional information such as filmography of the actor, i.e. a list of movies and TV shows in which the actor 601 appeared.
[0111] In a third example scenario as shown in Fig. 7, the user 110 is watching a movie or a TV show 700 (as the media content) and something happened that the user 110 does not understand (forgot an important part of the movie or the TV show or did not understand the scene acoustically / visually). In the present example, the character 701 displayed in the scene 700 is waiting to board a train. The user 110 asks about what happened, e.g., by voicing a question (“Why is this person travelling?”) or inputting the question via input means as a text to the media device 100. The feedback engine 200 is configured to understand the question based on the media content data 201 indicative of the media content, and to complement the media content data 201 with the information available online. Then, the media device 200 is configured to provide, as the feedback information 205, an explanation 703 to the user 110 which explains the context of the scene and what happened. In a fourth example scenario as shown in Fig. 8, the user 110 is watching a soccer game 800 on the media device 100 and gets passionate about the game and starts screaming as a goal is scored. The feedback engine 200 is configured to understand the sentiment of the user 110 (e.g. excitement) and to understand what is happening in the game (goal was scored by a player 801) The feedback engine 200 is configured to react to the user’s sentiment and to add immersion and engagement to the user experience. In this example, the feedback engine 200 is configured to provide, as the feedback information 205, additional information 803 regarding the player 801 who scored with respect to the sports game, such as indicating player statistics regarding the player 801.
[0112] In a fifth example scenario as shown in Fig. 9, the user 110 is watching a movie 900 (as the media content) and likes an outfit of a character shown in a scene. The user 110 comments that he wants to purchase the outfit (“I would like to buy this outfit”). The feedback engine 200 is configured to identify the clothes of the outfit in question, to complement the media content data 201 and / or the user data 203 with available online data and to provide, as the feedback information 205, a link 903 to a website which sells the outfit or similar clothing.
[0113] It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding. Changes of the ordering of method steps may be apparent to the skilled person.
[0114] All units and entities described in this specification and defined in the appended claims, if not stated otherwise, can be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.
[0115] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.
[0116] Note that the present technology can also be configured as described below.
[0117] (1) A media device for providing feedback to a user regarding media content output on the media device, the media device comprising circuitry configured to: acquire media content data related to media content; acquire user data which is indicative of a user behaviour of the user consuming the media content; input the media content data and the user data to a neural network to generate feedback information; and provide the feedback information to the user.
[0118] (2) The media device according to (1), wherein the neural network is a transformer-based Large Language Model, LLM.
[0119] (3) The media device according to (1) or (2), wherein the media device is a smart TV.
[0120] (4) The media device according to any one of (1) to (3), wherein the circuitry is further configured to input the media content data as media content embeddings and the user data as user embeddings to the neural network.
[0121] (5) The media device according to (4), wherein the circuitry is configured to generate the user embeddings based on the user data.
[0122] (6) The media device according to (4) or (5), wherein the circuitry is configured to obtain the media content embeddings from on a server external to a network comprising the media device
[0123] (7) The media device according to any one of (1) to (6), wherein the media device comprises a camera and / or a microphone for acquiring the user data.
[0124] (8) The media device according to any one of (1) to (7), wherein the media content data comprises at least one of media content image data, media content text data and media content audio data, and / or the user data comprises at least one of user image data, user text data and user audio data.
[0125] (9) The media device according to any one of (1) to (8), wherein the feedback information comprises at least one of a summary of the media content, an explanation of the media content, supplementary information regarding the media content and a graphical overlay on the media content.
[0126] (10) The media device according to any one of (1) and (9), wherein the media content data includes meta data regarding the media content.
[0127] (11) A method for providing feedback to a user regarding media content output on a media device, comprising: acquiring media content data related to media content; acquiring user data which is indicative of a user behaviour of the user consuming the media content; inputting the media content data and the user data to a neural network to generate feedback information; and providing the feedback information to the user.
[0128] (12) The method according to (11), wherein the neural network is a transformer-based Large Language Model, LLM.
[0129] (13) The method according to (11) or (12), wherein the media device is a smart TV.
[0130] (14) The method according to any one of (l l) to (13) further comprising: inputting the media content data as media content embeddings and the user data as user embeddings to the neural network.
[0131] (15) The method according to (14), further comprising: generating the user embeddings based on the user data.
[0132] (16) The method according to (14) or (15), further comprising: obtaining the media content embeddings from on a server external to a network comprising the media device.
[0133] (17) The method according to any one of (11) to (16), wherein acquiring the user data are generated by a camera and / or a microphone of the media device.
[0134] (18) The method according to any one of (11) to (17) wherein the media content data comprises at least one of media content image data, media content text data and media content audio data, and / or the user data comprises at least one of user image data, user text data and user audio data.
[0135] (19) The method according to any one of (11) to (18), wherein the feedback information comprises at least one of a summary of the media content, an explanation of the media content, supplementary information regarding the media content and a graphical overlay on the media content.
[0136] (20) The method according to any one of (11) and (19), wherein the media content data includes meta data regarding the media content.
[0137] (21) A computer program comprising program code causing a computer to perform the method according to anyone of (11) to (20), when being carried out on a computer. (22) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (11) to (20) to be performed.
Claims
CLAIMS1. A media device for providing feedback to a user regarding media content output on the media device the media device comprising circuitry configured to: acquire media content data related to media content; acquire user data which is indicative of a user behaviour of the user consuming the media content; input the media content data and the user data to a neural network to generate feedback information (205); and provide the feedback information to the user.
2. The media device according to claim 1, wherein the neural network is a transformerbased Large Language Model, LLM.
3. The media device according to claim 1, wherein the media device is a smart TV.
4. The media device according to claim 1, wherein the circuitry is further configured to input the media content data as media content embeddings and the user data as user embeddings to the neural network.
5. The media device according to claim 4, wherein the circuitry is configured to generate the user embeddings based on the user data.
6. The media device according to claim 4, wherein the circuitry is configured to obtain the media content embeddings from on a server external to a network comprising the media device.
7. The media device according to claim 1, wherein the media device comprises a camera and / or a microphone for acquiring the user data.
8. The media device according to claim 1, wherein the media content data comprises at least one of media content image data, media content text data and media content audio data, and / or the user data comprises at least one of user image data, user text data and user audio data.
9. The media device according to claim 1, wherein the feedback information comprises at least one of a summary of the media content, an explanation of the media content, supplementary information regarding the media content and a graphical overlay on the media content.
10. The media device according to claim 1, wherein the media content data includes meta data regarding the media content.
11. A method for providing feedback to a user regarding media content output on a media device, comprising: acquiring media content data related to media content; acquiring user data which is indicative of a user behaviour of the user (110) consuming the media content; inputting the media content data and the user data to a neural network to generate feedback information; and providing the feedback information to the user.
12. The method according to claim 11, wherein the neural network is a transformer-based Large Language Model, LLM.
13. The method according to claim 11, wherein the media device is a smart TV.
14. The method according to claim 11, further comprising: inputting the media content data as media content embeddings and the user data as user embeddings to the neural network.
15. The method according to claim 14, further comprising: generating the user embeddings based on the user data.
16. The method according to claim 14, further comprising: obtaining the media content embeddings from on a server external to a network comprising the media device.
17. The method according to claim 11, wherein the user data are generated by a camera and / or a microphone of the media device.
18. The method according to claim 11, wherein the media content data comprises at least one of media content image data, media content text data and media content audio data, and / or the user data comprises at least one of user image data, user text data and user audio data.
19. The method according to claim 11, wherein the feedback information comprises at least one of a summary of the media content, an explanation of the media content, supplementary information regarding the media content and a graphical overlay on the media content.
20. The method according to any one of claims 1 and 9, wherein the media content data includes meta data regarding the media content.
Citation Information
Patent Citations
Device, method and software for providing supplementary information
US20140111687A1
System and method for recommending media content based on actual viewers
US20210029391A1
Video recommendation method and device, computer device and storage medium
US20210281918A1
Joint embedding content neural networks
US20230091110A1