Information processing apparatus and method
The information processing apparatus uses a neural network to analyze event and user data, enhancing user engagement by providing personalized insights and interactive feedback during sport events, addressing the limitations of existing media devices.
Patent Information
- Application Number
- PCT/EP2025/057451
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-20
- Filing Date
- 2025-03-19
- Publication Date
- 2025-09-25
AI Technical Summary
Existing media devices lack the ability to provide personalized and interactive event insight information during sport events, limiting user engagement and experience.
An information processing apparatus and method that utilizes a neural network, specifically a transformer-based Large Language Model (LLM), to analyze event and user data, generating event insight information by combining visual and audio feedback based on user behavior, such as gestures, speech, and facial expressions, and providing contextual insights like player statistics and match details.
Enhances user engagement by offering personalized and interactive event insights, improving the viewing experience by providing relevant information in real-time through visual and audio feedback.
Smart Images

Figure EP2025057451_25092025_PF_FP_ABST
Abstract
Description
[0001] INFORMATION PROCESSING APPARATUS AND METHOD
[0002] TECHNICAL FIELD
[0003] The present disclosure generally pertains to information processing apparatus and a method for providing event insight information regarding a sport event.
[0004] TECHNICAL BACKGROUND
[0005] User interaction with media devices such as (smart) TVs, tablets, smartphones, etc. has evolved significantly over the years, from passive viewing to active engagement. From the beginnings of black-and-white TVs with manual dials, where viewers were passive consumers of content displayed on the TV screen, to the sleek, technology-enabled smart TV of today that offer a multifaceted and interactive viewing experience and allow new forms of interaction beyond the traditional remote control, such as voice commands, gesture commands, facial expressions commands, etc.
[0006] Technology has made media devices more interactive, and the introduction of Large Language models (LLMs) should be no exception. LLMs are a type of artificial Intelligence (Al) that can perform multiple natural language processing (NLP) tasks. LLMs are built based on deep learning and are trained on vast datasets containing text from the Internet. LLMs can generate human-like texts, answer questions, translate languages, summarize documents, and engage in contextually relevant conversations. From content generation and recommendation systems to virtual assistants and customer service solutions, LLMs have the potential to impact numerous aspects of daily lives.
[0007] It is generally desirable to use the above technologies to improve user experience when watching a sport event.
[0008] SUMMARY
[0009] According to a first aspect, the disclosure provides an information processing apparatus for providing event insight information regarding a sport event observed by a user, wherein the information processing apparatus comprises circuitry configured to: acquire event data related to the sport event; acquire user data which is indicative of a user behaviour; input the event data and the user data to a neural network to generate the event insight information; and output a feedback based on the event insight information to the user.
[0010] According to a second aspect, the disclosure provides a method for providing event insight information regarding a sport event observed by a user, comprising: acquiring event data related to the sport event; acquiring user data which is indicative of a user behaviour; inputting the event data and the user data to a neural network to generate the event insight information; and outputting a feedback based on the event insight information to the user.
[0011] Further aspects are set forth in the dependent claims, the drawings, and the following description.
[0012] BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Embodiments are explained by way of example with respect to the accompanying drawings, in which:
[0014] Fig. 1 shows a user and a computer vision system observing a sport event;
[0015] Fig. 2 shows a block diagram of an exemplary configuration of a terminal device;
[0016] Fig. 3 shows a block diagram of an exemplary configuration of an event insight engine of the terminal device;
[0017] Fig. 4 shows a block diagram of an exemplary configuration of a multimodal encoder of the insight event engine;
[0018] Fig. 5 shows a method for providing event context information;
[0019] Fig. 6 shows a method for providing user information;
[0020] Fig. 7 shows a method for providing event insight information;
[0021] Fig. 8 shows a first example scenario for the terminal device;
[0022] Fig. 9 shows a second example scenario for the terminal device;
[0023] Fig. 10 shows a third example scenario for the terminal device;
[0024] Fig. 11 shows a plurality of terminal devices; and
[0025] Fig. 12 shows a server and an exemplary configuration of the plurality of terminal devices. DETAILED DESCRIPTION OF EMBODIMENTS
[0026] Before a detailed description of the embodiments under reference of Fig. 1 is given, general explanations are made.
[0027] In some embodiments, there is provided an information processing apparatus for providing event insight information regarding a sport event observed by a user, wherein the information processing apparatus comprises circuitry configured to: acquire event data related to the sport event; acquire user data which is indicative of a user behaviour; input the event data and the user data to a neural network to generate the event insight information; and output a feedback based on the event insight information to the user.
[0028] The information processing apparatus can be a tablet, a smart phone, smart glasses, a smart watch, a smart contact lens, a virtual retinal display, or any other device wearable by the user. In some examples, the information processing apparatus can be a smart TV with a transparent display screen (e.g., a transparent OLED display). Further, the information processing apparatus may comprise an interface (communication module) to access the Internet. By way of the Internet connection, the information processing apparatus may retrieve information from the Internet and may be connected with other electronic devices. Further, the Internet connection allows the information processing apparatus to stream media content provided on streaming portals or download the media content and store it on a storage of the information processing apparatus for later consumption. Further, the information processing apparatus can be configured to access, via APIs, third party applications.
[0029] The user may be present in a stadium or any other location where he / she can observe the sport event live. The user may also watch the sport event in case it is displayed on the display screen of the information processing apparatus.
[0030] The sport event is indicated by the event data. That is, the event data represents or is related to the sport event.
[0031] The event data comprises at least one of event image data, event text data, and event audio data. For example, the event image data may correspond to a video recording of the sport event. Further, the even audio data may correspond to noise, speech, and the like recorded during the sport event. The text event data may be text data related to the sport event. Such text data can be obtained by a speech-to-text encoder of speech recorded of the sport event or may be retrieved as additional information from the Internet.
[0032] “Acquiring the event data” means that the event data captured in real-time by a camera and / or microphone of the information processing apparatus or that the event data is streamed (or broadcasted) to the information processing apparatus. Alternatively, the event data is downloaded to the information processing apparatus.
[0033] The user data is indicative of the user behaviour. That is, the user behaviour of the user observing the sport event can be derived from the user data. The user behaviour may be a gesture, a gaze, a position, a speech (e.g. a comment, a question, an inquiry, etc.), a mood, and the like.
[0034] The user content data comprises at least one of user image data, user text data, and user audio data. “Acquiring the user data” may include that the information processing apparatus is configured to capture the user data. For example, the information processing apparatus comprises the camera to capture an image. The camera then outputs the user image data corresponding to the captured image. The camera may be configured as a RBG camera or event camera (also “event-based camera”). Further, a presence sensor (such as a mmWave sensor which uses short-wavelength electromagnetic waves) may be used to capture user data.
[0035] Further, the information processing apparatus comprises the microphone configured to capture a speech of the user as user audio data.
[0036] Furthermore, the information processing apparatus may comprise input means configured to capture text data entered by the user in the input means. The input means may be wired to or wirelessly coupled with the information processing apparatus. The input means can be a remote controller (such as a conventional TV remote controller), gaming pad, touchpad of the information processing apparatus, etc.
[0037] The information processing apparatus observes the actions of the user, such as to where the user is paying attention, which objects or actions are being observed, gestures, and voice inputs. With this, the context and interests of the user can be determined and mapped to the actions happening in the observed sport event, i.e. the event data.
[0038] Therefore, the event data and the user data are input to the neural network to generate the event insight information for the user. The neural network is trained to understand how the user is reacting to the sport event and to provide explanations of the sport event, additional information, statistics, and the like with respect to the sport event.
[0039] The neural network may be configured to perform NPL, object segmentation, action recognition to capture the behaviour of the user and the sport event.
[0040] In some examples, the neural network may be an LLM. In further examples, the LLM may be based on a transformer-based neural network architecture as described in A. Vaswani et al., “Attention is all you need,” in 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 2017. Such exemplary neural networks known in the art include BERT, ChatGPT, and the like.
[0041] The neural network generates the event insight information. The event insight information is textual information generated in response to the user behaviour as indicated in the user data.
[0042] The information processing apparatus is configured to output the feedback based on the event insight information to the user.
[0043] That is, the event insight information is used to provide a visual feedback, an audio feedback or a text feedback representing the event insight information. Further, a combination of different feedback modalities is also possible.
[0044] When providing visual feedback, the text-based insight information is converted into visual information and output via the display unit of the information processing apparatus. Similarly, when providing text feedback, the textual information is displayed on the display unit. When providing audio feedback, the text-based insight information is converted into speech information and output via a speaker of the information processing apparatus.
[0045] With the information processing apparatus, it is possible to observe the state, actions and intentions of the user watching the sport event (broadcasted or live in a stadium) and provide relevant event information given the context. The information processing apparatus is driven by the neural network, e.g., an LLM, which has access to relevant information (retrievable from the Internet) about the sport event, such as teams involved, weather information, current league status, and the like.
[0046] Further, as the user data is captured with the camera and / or microphone, the user can interact and exchange information with the information processing apparatus in free language format by utilising the neural network (LLM) and combined with eye / hand tracking to understand whether the user is pointing to a specific object or player. Thus, it is possible to improve the user engagement by providing fully personalized experience and information to the user observing the sport event.
[0047] The information processing apparatus combines visual analysis, speech-to-text encoding, text-to- speech encoding and LLM technology such that the context of the sport event and an interest of the user can be analysed to provide insights into the sport event (event insight information). Such might include specific details of the match, such as who is controlling the ball, any penalties, movement analysis of the players, player statistics, and match statistics. The event insight information can be provided as feedback using voice outputs and / or graphical overlays displayed on the display unit of the information processing apparatus or over the view of smart glasses (in case the information processing apparatus is of a smart glasses type).
[0048] In some embodiments, the neural network is a transformer-based LLM as described above.
[0049] In some embodiments, the user data is indicative of at least one of a speech, a gesture, a text input, and a facial expression. A gesture may be a user’s finger pointing towards an object or in a direction, a head nodding, and the like. A facial expression may be a smile, a frown, a wink, an eyebrow movement, and the like. A speech may include a question, an exclamation, a comment, and the like.
[0050] In some embodiments, the information processing apparatus is in data communication with a computer vision system. The computer vision system is configured to capture the sport event. An example for a known computer vision is Hawkeye. The data communication may be established directly or via the Internet. The direct connection may be wireless or wired. Thereby, the information processing apparatus can retrieve the data (metadata), e.g. audio data and / or image data, captured by the computer vision system.
[0051] In some embodiments, the event data comprises the metadata provided by the computer vision system. That is, the information processing apparatus may collect metadata provided by the computer vision system.
[0052] In some embodiments, the information processing apparatus comprises the camera and / or the microphone for acquiring the user data.
[0053] In some embodiments, the circuitry is further configured to input the event data as event embeddings and the user data as user embeddings to the neural network. The event embeddings include event vision embeddings and event text embeddings. Correspondingly, the user embeddings include user vision embeddings and user text embeddings.
[0054] Generally, text embeddings are a numerical representation of text, used for measuring semantic similarities and aiding in context-based Al interaction. For example, text embeddings are vectors or arrays of numbers which represents the meaning text data. The text embeddings are then used by the neural network.
[0055] Correspondingly, vision embeddings are numerical representations of images that encode the semantics of in image data. For example, vision embeddings are vectors or arrays of numbers which represents the meaning and the context of the image data.
[0056] The events embeddings are generated based on the event data. More specifically, the event vision embeddings are generated based on the event image data and the event text embeddings are generated based on the event text data and / or the event audio data.
[0057] Correspondingly, the user vision embeddings are generated based on the user image data and user text embeddings are generated based on the user text data and / or on the user audio data.
[0058] The information processing apparatus may comprise a vision encoder which is configured to generate the vision embeddings (i.e. the event vision embeddings and the user vision embeddings) based on the image data (i.e. the event image data and the user image data, respectively).
[0059] The information processing apparatus may comprise a text embedder which is configured to generate the text embeddings (i.e. the event text embeddings and the user text embeddings) based on the text data (i.e. the event text data and the user text data, respectively).
[0060] Further, the information processing apparatus may comprise a speech-to-text encoder which is configured to convert speech included in audio data into text data which can be subsequently fed to the text embedder to generate text embeddings. More specially, speech included in the event audio data may be converted into event text data by the speech-to-text encoder. The converted event text data can then be input to the text embedder which generates the event text embeddings based on the converted event text data.
[0061] Correspondingly, the user text embeddings can also be generated based on the user audio data as described with respect to the processing of the media content audio data. The text embeddings and the vision embeddings are multimodal inputs for the neural network. Therefore, in some examples, the information processing apparatus comprises a multimodal aligner which is configured to process the text embeddings and the vision embeddings into a format which can be processed by the neural network.
[0062] In some embodiments, the media content data may comprise at least one of the media content image data, the media content text data, and the media content audio data. Additionally, or alternatively, the user data may comprise at least one of the user image data, the user text data, and the user audio data.
[0063] In some embodiments, the circuitry is configured to acquire additional event data regarding the sport event from the Internet. The additional information can be acquired from sport websites, search engines, and other resources from the Internet. The additional information from the Internet may include background information of teams and players participating in the sport event. Based on the additional information, it is possible to provide general statistics and historical information the teams and / or players. For example, statistics over many years, what contracts players had, personal and professional information about them, contracts with the clubs, and the like may be acquired form the Internet.
[0064] In some embodiments, the observed sport event is a live sport event.
[0065] In some embodiments, the information processing apparatus is a mobile terminal device.
[0066] In some embodiments, the method for providing feedback to a user regarding media content output on a media device, comprises: acquiring event data related to the sport event; acquiring user data which is indicative of a user behaviour; inputting the event data and the user data to a neural network to generate the event insight information; and outputting a feedback based on the event insight information to the user.
[0067] The method may be carried out with the information processing apparatus comprising the circuitry discussed herein.
[0068] In some embodiments, the neural network is a transformer-based Large Language Model, LLM. In some embodiments, the user data is indicative of at least one of a speech, a gesture, a text input, and a facial expression. In some embodiments, the method further establishing a data communication with a computer vision system. In some embodiments, the event data comprises metadata provided by a computer vision system. In some embodiments, the user data is acquired by a camera and / or a microphone of a terminal device. In some embodiments, the event data comprises at least one of event image data, event text data and event audio data, and / or the user data comprises at least one of user image data, user text data and user audio data. In some embodiments, the method further comprises acquiring additional event data regarding the sport event from the Internet. In some embodiments, the observed sport event is a live sport event. In some embodiments, the method is carried out by a mobile terminal device.
[0069] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer- readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.
[0070] Returning to Fig. 1, it is shown a user 110 using a terminal device 100 (information processing apparatus) such as smart glasses observing a sport event 1000. In the present example, the sport event 1000 is a football game. A first player 1001, a second player 1003 and a third player 1005 and a ball 1007 are part of the sport event 1000. Further, the sport event 1000 is observed by a computer vision system 400 configured for detecting situation during the sport event 1000 such as Hawkeye.
[0071] Fig. 2 shows an exemplary configuration of the terminal device 100. The terminal device 100 is includes a display screen 101, a speaker 103, a camera 105, a microphone 107, input means 109, a communication module 111, and an event insight engine 200.
[0072] Generally, the camera 105 and the microphone 107 are configured to capture event data 201 regarding the sport event 1000 and user data 203 regarding the behaviour of the user 110.
[0073] More specifically, the camera 105 is configured to capture image data regarding the user 110 (user image data 203a). For example, the camera 105 is configured to capture a user’s gesture and to provide user image data indicative of the user’s gesture. The user’s gesture may a be hand gesture (e.g. pointing to a position or object on the screen of the media device 100), head gesture (e.g. nodding) and the like. The camera 105 is further configured to capture a gaze or a gaze direction of the user 110 and based on the captured image data indicative of the gaze or the gaze direction.
[0074] Further, the camera 105 is configured capture image data regarding the sport event 1000. The microphone 107 is configured to capture audio data regarding the user 110 (user audio data 203c) and event audio data regarding the sport event 1000. For example, the microphone 107 may capture speech of the user or uttered by the players 1001, 1003, 1005 participating in the sport event 1000.
[0075] Further, the terminal device 100 may comprise input means 109 configured to receive input by the user 110. The input means 109 can be integrated into the terminal device 100 such as touch input means. In another example, the input means 109 can be remotely arranged from but coupled to the terminal device 100 such as a remote controller, keyboard and mouse, gaming controller. The user 101 may enter user text data 203b via the input means 109.
[0076] By way of the communication module 111, the terminal device 100 is connected or has access to the Internet. Further, the Internet connection enables the terminal device 100 to access data bases, use cloud storages, use search engines (e.g. Google, Bing, etc.), etc. Further, the terminal device 100 is connected to the computer vision system 400 to acquire event data 201 captured by the computer vision system 400.
[0077] Fig. 3 shows a configuration of the event insight engine 200 which is configured to generate event insight information 311 based on the captured event data 201 and the captured user data 203. The event insight engine 200 comprises a multimodal encoder 300, a context database 201, a retriever 205 and an LLM 207.
[0078] In order to provide event insight information 311, the event insight engine 300 is configured to receive the event data 201 and the user data 203 as input data and processes the input data to generate the event insight information 3311. Thus, the event insight engine 200 analyses the sport event (indicated by the event data 201) and the behaviour of the user 110 (indicated by the user data 203) to generate the event insight information 311.
[0079] The event data 201 is indicative of the observed sport event 1000. The event data 201 may comprise at least one of event image data 201a, event text data 201b, and event audio data 201c.
[0080] The user data 203 is indicative of a behaviour of the user 110. As described above, the user data 203 may comprise the user image data 203a and the user audio data 203c captured with the camera 105 and the microphone 107, respectively. Further, the user data 203 may also comprise the text input by the user 110 (user text data 203b) through the input means 109.
[0081] The event data 201 and the user data 203 are input to the multimodal encoder 200 which is configured to generate event context information 301 and user information 303 for further processing by the LLM 207. That is, the multimodal encoder 200 is configured to encode the event data 201 and the user data 203 into a data format which can be processed by the LLM 207.
[0082] With respect to Fig. 4, the configuration of the multimodal encoder 300 is explained.
[0083] The multimodal encoder 300 comprises a speech-to-text encoder 310, a tokenizer 320, a text embedder 330, a vision encoder 340, and a multimodal aligner 250. By way of example, processing of the user data 203 is explained now with respect to Fig. 4. The processing of the event data 201 is performed correspondingly.
[0084] The speech-to-text encoder 310 is configured to convert spoken language into text form. That is, the speech-to-text encoder 310 is configured to receive the user audio data 203c indicative of spoken language and to convert the user audio data 203c into text data expressed as a sequence of text. Configurations of the speech-to-text encoder 310 are known to the skilled person and may comprise, for example, an attention-based encoder-decoder framework or a transducer framework.
[0085] The tokenizer 320 is configured to receive, as input, the user text data 203b. Further, the input to the tokenizer 320 may also be the converted text data output by speech-to-text encoder 310.
[0086] The tokenizer 320 is configured to convert the user text data 203b into tokens. That is, the tokenizer 320 is configured to break a text sequence comprised in the user text data 203b into smaller parts, i.e. tokens. For example, a sentence “One ring to rule them all” comprised in the user text data 203b is tokenized into the individual words “one”, “ring”, “to”, “rule”, “them”, “all”. The output of the tokenizer 320 is a sequence of tokens that represents the text included in the user text data 303b. In other examples, the tokens can additionally comprise sub words, signs, and other small units, which collectively represents the sentence.
[0087] The text embedder 330 is configured to receive, as input, the tokens generated by the tokenizer 320 based on the user text data 203b and to generate (output) text embeddings indicative of the user text data 203b.
[0088] The vision encoder 340 is configured to receive the user image data 203a. The vision encoder 340 is configured to generate (output) vision embeddings based on the received user image data 203a.
[0089] The multimodal aligner 350 is configured to align the vision embeddings generated by the vision encoder 340, and the text embeddings generated by the text embedder 330 such that the LLM 207 can understand and process the vision embeddings and text embeddings. The multimodal aligner 350 is configured to output the user information 303 which is indicative of a textual description of the user data 203 suitable for processing by the LLM 207.
[0090] Correspondingly, multimodal aligner 350 is configured to generate the event context information 301 which is indicative of a textual description of the event data 201 suitable for processing by the LLM 207. That is, the processing of the user data 203 by the multimodal encoder 300 as described above applies correspondingly to the event data 201.
[0091] Returning to Fig. 3, the configuration of the event insight engine 200 is further explained. The event insight engine 200 is configured to perform a Retrieval Augmented Generation (RAG) based LLM application.
[0092] Therefore, the context database 201 is configured to store the event context information 301. The context database 201 can be a vector storage.
[0093] Further, the retriever 205 is configured to retrieve query context information 305 from the context database 201 based on a user query which corresponds or is the user information 303. As the user information 303 is related to the sport event 1000, the event context information 301 indicative of the sport event 1000 are retrieved from the context database 201 as the query context information 305.
[0094] Then, the LLM 207 is configured to receive the user information 303 and the query context information 305 to generate the event insight information 311.
[0095] Fig. 5 shows a method 500 for processing the event data 201 for generating the event information 303. The method 500 is carried out by the multimodal encoder 300.
[0096] In step 501, the event data 201 as captured by the camera 105 and / or microphone 107 is obtained. Additionally, the event data 201 may comprise the metadata as provided by the computer vision system 400.
[0097] In step 503, event vision embeddings are generated based on the event image data 201a. To this end, the event image data 201a is input to the vision encoder 340.
[0098] In step 505, event text embeddings are generated based on the event text data 201b and / or the event audio data 201c. To generate text embeddings based on the event text data 201b, the event text data 201b is input into the tokenizer 320 which generates tokens. The tokens are then input into the text embedder 330. To generate text embeddings based on the event audio data 201c, the event audio data 201c is input to the speech-to-text-encoder 310 which detects speech included in the event audio data 201c and outputs the detected speech as text data. The text data output by the speech-to-text encoder 310 is then processed by the tokenizer 320 and the text embedder 330 to generate text embeddings as described with respect to the event text data 201b.
[0099] In step 507, the event vision embeddings and the event text embeddings are then aligned by the multimodal aligner 350 to generate the event context information 301 which can later be processed by the LLM 207.
[0100] In step 509, the multimodal aligner 350 outputs the event context information 301 for further processing.
[0101] Fig. 6 shows a method 600 for processing the user data 203 for generating the user information (user query) 303. The method 600 is carried out by the multimodal encoder 300 and corresponds to the processing of the event data 201 as shown in method 500.
[0102] In step 601, the user data 203 as captured by the camera 105 and / or microphone 107 is obtained.
[0103] In step 603, user vision embeddings are generated based on the user image data 203a. To this end, the user image data 203a is input to the vision encoder 340.
[0104] In step 605, user text embeddings are generated based on the user text data 203b and / or the user audio data 203c. To generate text embeddings based on the user text data 203b, the user text data 203b is input into the tokenizer 320 which generates tokens. The tokens are then input into the text embedder 330. To generate text embeddings based on the user audio data 203c, the user audio data 203c is input to the speech-to-text-encoder 310 which detects speech included in the user audio data 203c and outputs the detected speech as text data. The text data output by the speech-to-text encoder 310 is then processed by the tokenizer 320 and the text embedder 330 to generate text embeddings as described with respect to the user text data 203b.
[0105] In step 607, the user vision embeddings and the user text embeddings are then aligned by the multimodal aligner 360 to generate the user information 303 which can later be processed by the LLM 207.
[0106] In step 609, the multimodal aligner 350 outputs the user information 303 for further processing.
[0107] Fig. 7 shows a method 700 for providing the event insight information 311. The method 700 is performed by the terminal device 100.
[0108] In step 701, the event context information 301 is stored in the context database 201.
[0109] In step 703, the user information (user query) 303 is input to the retriever 205. In step 705, the retriever 205 retrieves the event context information 301 based on the user information 303. The retriever 205 outputs the event context information 301 as query context information 305. In some examples, the query context information 305 aggregates the retrieved event context information 305 in a single data file.
[0110] In step 707, the user information 303 and the query context information 305 are input to the LLM 207 to generate the event insight information 311.
[0111] In step 709, the event insight information 311 is output by the LLM 207.
[0112] In step 711, the terminal device 100 outputs a feedback based on the event insight information 311 is output by the LLM 207.
[0113] Fig. 8 shows first example scenario for using the terminal device 100 for providing the event insight information 311. The user 110 voices the question “Which player has the ball?”. The microphone 107 captures the question as the user data 203 (more specifically, as the audio user data 203c). As described above, the user data 203 is processed by the event insight engine 200 to generate the user information 303 based on the user data 203. Further, the terminal device 100 captures the event data 201 regarding the sport event 1000. The event data 201 is processed by the event insight engine 200 to generate the event context information 301. Then, the LLM 207 of the event insight engine 200 generates the event insight information 311 which indicates that the first player 1001 is in possession of the ball 1007. The terminal device 100 outputs via its speaker 103 the feedback to the user information 303 (which is indicative of the question “Which player has the ball?”) an audible feedback 311a based on the event insight information 311. The audible feedback 311a includes the indication of a player name of the first player 1001.
[0114] Fig. 9 shows second example scenario for using the terminal device 100 for providing the event insight information 311. In this example, the terminal device 100 is configured as smart glasses and the sport event 1000 is depicted as seen by the user 110 through the smart glasses. The user voices the question “Who is this?” and, at the same time, points with his finger or gaze to the second player 1005. The terminal device 100 captures the question and the pointing as the user data 203. Further, the current situation of the sport event 1000 is captured by the terminal device 100 as the event data 201. As described above, the event insight engine 200 generates the event insight information 311 based on the event data 201 and the user data 203. A visual feedback 31 lb is output based on the event insight information 311 which indicates the name and player stats of the second player 1003 in response to the question and pointing of the user 110. In this example, the visual feedback 31 lb is an overlay 311 projected on the glasses of the terminal device 100.
[0115] Fig. 10 shows third example scenario for using the terminal device 100 for providing the event insight information 311. In this example, the terminal device 100 is configured as a smart TV comprising a transparent display screen lOland the user 110 observes the sport event 1000 through the transparent display screen. In the current situation of the sport event 1000, the third player 1005 scored a goal with an assist of the first player 1001. The current situation is captured as the event data 201. The user 100 voices “What a nice goal!” which is captured as the user data 203. As described above, the event insight information 311 is generated based on the captured event data 201 and the user data 203. In this example, the LLM 207 is trained to output the event insight information 311 indicating a movement analysis of the involved players, i.e. the first and third player 1001, 1005, and of the ball 1007. The terminal device 100 outputs a visual feedback in form of an overlay displayed on the transparent display screen 101. The overlay includes a movement path 311c (dotted line) of the ball 1007 and a movement path 311 (dashed line) of the involved first and third player 1001, 1005.
[0116] Fig. 11 shows a plurality of users 110a, 110b, 110c who together observe the sport event 1000. Each of the plurality of users 110a, 110b, 110c has a corresponding terminal device of a plurality of terminal devices 100a, 100b, 100c. The plurality of terminal devices 100a, 100b, 100c are connected to each other by way of a “watch together” feature. The plurality of terminal devices 100a, 100b, 100c are connected via the Internet 800 with a server 300. The configuration of Fig. 11 differs from the configuration of Fig. 1 in that the server 900 is configured to perform the processing of the event data 201 and the user data 203 to generate the event insight information 311 and then sends the event insight information 311 to the plurality of terminal devices 100a, 100b, 100c, where the feedback based on the event insight information 311 is output by the plurality of terminal devices 100a, 100b, 100c.
[0117] Therefore, as shown with respect to Fig. 12, each of the plurality of terminal devices 100a, 100b, 100c comprises the components of the terminal device 100 as shown in Fig. 2 except for the event insight engine 200.
[0118] The server 900 comprises the event insight engine 200 (as described with respect to Fig. 3) and a communication module 911.
[0119] Processing of the event data 201 and the user data 203 differs from the above methods 500, 600 in that the step 501 of method 500 and step 601 of method 600 are performed by the server 900. That is, the server 900 obtains or receives from the plurality of terminal devices 100a, 100b, 100c the event data 201 and the user data 203, wherein each of the plurality of terminal devices 100a, 100b, 100c individually gathers the data. The event data 201 and the user data 203 are then processed as described in the remaining steps of method 500, 600 and the method 700.
[0120] However, step 709 of method 700 further includes to transmit the event insight information 311 to the plurality of terminal devices 100a, 100b, 100c, each of which then carry out step 711, i.e. outputting the feedback based on the event insight information 311.
[0121] By way of the “watch together” feature, more event data regarding sport event 1000 may be collected and therefore more event insight information 311 may be provided by the server 900.
[0122] It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding.
[0123] It is noted that the division of the processing of the event data 201 and the user data 203 by into the event insight engine 200 and the multimodal encoder 300 is only made for illustration purposes and that the present disclosure is not limited to any specific division of functions in specific units. For instance, the processing of the event data 201 and the user data 203 for generating the event insight information 311 could be implemented by a respective programmed processor, field programmable gate array (FPGA) and the like.
[0124] All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.
[0125] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.
[0126] Note that the present technology can also be configured as described below.
[0127] (1) An information processing apparatus for providing event insight information regarding a sport event observed by a user, wherein the information processing apparatus comprises circuitry configured to: acquire event data related to the sport event; acquire user data which is indicative of a user behaviour; input the event data and the user data to a neural network to generate the event insight information; and output a feedback based on the event insight information to the user.
[0128] (2) The information processing apparatus according to (1), wherein the neural network is a transformer-based Large Language Model, LLM.
[0129] (3) The information processing apparatus according to (1) or (2), wherein the user data is indicative of at least one of a speech, a gesture, a text input, and a facial expression.
[0130] (4) The information processing apparatus according to anyone of (1) to (3), wherein the information processing apparatus is in data communication with a computer vision system.
[0131] (5) The information processing apparatus according to anyone of (1) to (4), wherein the event data comprises metadata provided by the computer vision system.
[0132] (6) The information processing apparatus according to anyone of (1) to (5), further comprising a camera and / or a microphone for acquiring the user data.
[0133] (7) The information processing apparatus according to anyone of (1) to (6), wherein the event data comprises at least one of event image data, event text data and event audio data, and / or the user data comprises at least one of user image data, user text data and user audio data.
[0134] (8) The information processing apparatus according to anyone of (1) to (7), wherein the circuitry is configured to acquire additional event data regarding the sport event from the Internet.
[0135] (9) The information processing apparatus according to anyone of (1) to (8), wherein the observed sport event is a live sport event.
[0136] (10) The information processing apparatus according to anyone of (1) to (9), wherein the information processing apparatus is a mobile terminal device.
[0137] (11) A method for providing event insight information regarding a sport event observed by a user, comprising: acquiring event data related to the sport event; acquiring user data which is indicative of a user behaviour; inputting the event data and the user data to a neural network to generate the event insight information; and outputting a feedback based on the event insight information to the user.
[0138] (12) The method according to (11), wherein the neural network is a transformer-based Large Language Model, LLM.
[0139] (13) The method according to (11) or (12), wherein the user data is indicative of at least one of a speech, a gesture, a text input, and a facial expression.
[0140] (14) The method according to anyone of (11) to (13), further comprising establishing a data communication with a computer vision system.
[0141] (15) The method according to anyone of (11) to (14), wherein the event data comprises metadata provided by a computer vision system.
[0142] (16) The method according to anyone of (11) to (15), wherein the user data is acquired by a camera and / or a microphone of a terminal device.
[0143] (17) The method according to anyone of (11) to (16), wherein the event data comprises at least one of event image data, event text data and event audio data, and / or the user data comprises at least one of user image data, user text data and user audio data.
[0144] (18) The method according to anyone of (11) to (17), further comprising acquiring additional event data regarding the sport event from the Internet.
[0145] (19) The method according to anyone of (11) to (18), wherein the observed sport event is a live sport event.
[0146] (20) The method according to anyone of (11) to (19), wherein the method is carried out by a mobile terminal device.
[0147] (21) A computer program comprising program code causing a computer to perform the method according to anyone of (11) to (20), when being carried out on a computer.
[0148] (22) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (11) to (20) to be performed.
Claims
CLAIMS1. An information processing apparatus for providing event insight information regarding a sport event observed by a user, wherein the information processing apparatus comprises circuitry configured to: acquire event data related to the sport event; acquire user data which is indicative of a user behaviour; input the event data and the user data to a neural network to generate the event insight information; and output a feedback based on the event insight information to the user.
2. The information processing apparatus according to claim 1, wherein the neural network is a transformer-based Large Language Model, LLM.
3. The information processing apparatus according to claim 1, wherein the user data is indicative of at least one of a speech, a gesture, a text input, and a facial expression.
4. The information processing apparatus according to claim 1, wherein the information processing apparatus is in data communication with a computer vision system.
5. The information processing apparatus according to claim 1, wherein the event data comprises metadata provided by a computer vision system.
6. The information processing apparatus according to claim 1, further comprising a camera and / or a microphone for acquiring the user data.
7. The information processing apparatus according to claim 1, wherein the event data comprises at least one of event image data, event text data and event audio data, and / or the user data comprises at least one of user image data, user text data and user audio data.
8. The information processing apparatus according to claim 1, wherein the circuitry is configured to acquire additional event data regarding the sport event from the Internet.
9. The information processing apparatus according to claim 1, wherein the observed sport event is a live sport event.
10. The information processing apparatus according to claim 1, wherein the information processing apparatus is a mobile terminal device.
11. A method for providing event insight information regarding a sport event observed by a user, comprising: acquiring event data related to the sport event; acquiring user data which is indicative of a user; inputting the event data and the user data to a neural network to generate the event insight information; and outputting a feedback based on the event insight information to the user.
12. The method according to claim 11, wherein the neural network is a transformer-based Large Language Model, LLM.
13. The method according to claim 11, wherein the user data is indicative of at least one of a speech, a gesture, a text input, and a facial expression.
14. The method according to claim 11, further comprising establishing a data communication with a computer vision system.
15. The method according to claim 11, wherein the event data comprises metadata provided by a computer vision system.
16. The method according to claim 11, wherein the user data is acquired by a camera and / or a microphone of a terminal device.
17. The method according to claim 11, wherein the event data comprises at least one of event image data, event text data and event audio data, and / or the user data comprises at least one of user image data, user text data and user audio data.
18. The method according to claim 11, further comprising acquiring additional event data regarding the sport event from the Internet.
19. The method according to claim 11, wherein the observed sport event is a live sport event.
20. The method according to claim 11, wherein the method is carried out by a mobile terminal device.
Citation Information
Patent Citations
Content recommendation based on a system prediction and user behavior
US11490163B1
Augmented experience of media presentation events
US20170006356A1
Live event context based crowd noise generation
US20220114390A1