Apparatus and method for multi-modal video analysis
Through multimodal video analysis, multiple technologies are used to process the frames, audio and metadata of the video stream, and generate video-level and time-stamped tags, solving the problem that traditional solutions cannot analyze complex actions and implementing convenient video search and recommendation functions.
Patent Information
- Application Number
- CN202280102281.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to effectively analyze complex actions and provide corresponding video tags, especially on short video platforms. Traditional solutions cannot meet the users' needs to initiate searches from video content quickly and simply.
By receiving multimodal information of the video stream, including frames, audio and metadata, using automatic speech recognition, optical character recognition, computer vision and natural language processing technologies, video-level tags and timestamped tags are determined and stored in a database, providing user equipment with convenient tags and recommendations.
It realizes the analysis of complex actions in the video stream, provides convenient tags and recommendations, supports users to quickly and simply initiate searches from video content, and improves the efficiency and accuracy of searches.
Smart Images

Figure CN120303653A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video analysis, and in particular to marking videos based on inference technology. Therefore, the present invention provides an apparatus for multimodal video analysis for extracting tags based on several attributes of a video. In addition, corresponding methods and computer programs are also provided. Background Art
[0002] Short video platforms are becoming increasingly popular. People, especially the younger generation, use short videos as their main source of information and learn about topics they are interested in by watching short videos rather than reading text. Online videos are usually paired with a series of tags (for example, a football video can be paired with "football", "Germany", "Italy", "World Championship", etc.). These tags synthesize their content and act as "containers" to group similar videos together. However, this kind of container only works within the website and is customized to facilitate the retrieval of similar video content within the same platform.
[0003] Generally, when browsing the Internet, it usually starts from a page and then navigates from that page, link by link, until the exact content needed is found. However, it is not convenient to initiate a search from a video that has already been watched. On the one hand, this is because the video tags defined by the uploader usually do not cover all or most of the aspects and topics that the user may want to search for. On the other hand, video search has a closed loop on short video platforms. That is to say, there is no simple way to search in another search engine with a simple click.
[0004] To solve and alleviate this problem, existing work uses machine learning to understand the content of videos and automatically generate a list of relevant tags. Nevertheless, the tagging system is mainly used for indexing purposes and recommends other videos that can be searched, rather than directly initiating a search from the video content. For such interactions, there is currently a lack of a solution that can achieve the required level of simplicity.
[0005] In fact, if a user wants to search for a certain topic related to the video content (for example, using a more suitable external search engine), there are two possibilities: if the tag has already been paired with the video, the user can manually search by copying and pasting the content of the tag into the relevant web page. This process is very time-consuming and requires unreasonable manual operations by the user. This is especially obvious when the user's behavior tends to interact with web content faster and more simply, and the required information should always be just a click away. If the video is not tagged with the topic of interest required by the user, the situation will be even worse because the user needs to open a search engine and formulate the correct search query by manual input.
[0006] Traditional solutions focus on retrieving relevant content from a video source that can be used to tag the video. This is typically achieved by using machine learning and RGB frames (i.e., the individual images that make up the video). Some solutions include identifying objects in a scene and displaying interactive content that overlays the video content based on a user request; enriching the video content by retrieving additional video sources that can be displayed along with the main video; or identifying people of interest (e.g., actors) and displaying pop-up labels at different moments during the video.
[0007] Although traditional solutions may seem to cover a wide range of use cases, recent changes in video content, mainly brought about by social media platforms, have presented new challenges and problems that have not been addressed: the focus of the video is no longer on specific people / objects, but on actions (e.g., a trendy dance or a sports performance). Additionally, the information payload of recent videos may not be solely present in the images.
[0008] Therefore, traditional solutions are unable to analyze complex actions and provide corresponding tags in recent videos. Summary of the Invention
[0009] In view of the above problems, an object of embodiments of the present invention is to provide a method for tagging a video based on multimodal video analysis.
[0010] This object or other objects can be achieved by embodiments of the present invention described in the appended independent claims. Advantageous implementations of embodiments of the present invention are further defined in the dependent claims.
[0011] A first aspect of the present invention provides a device for multimodal video analysis, wherein the device is configured to: receive a video stream; determine video-level tags and timestamped tags based on at least two frames of the video stream, audio information of the video stream, and inference techniques.
[0012] This ensures that complex actions in the video stream can be analyzed and tags can be determined accordingly.
[0013] Specifically, the tags indicate specific situations in the video stream. Specifically, the tags indicate a single data modality or multiple data modalities (e.g., data modalities include the presence of a specific object or text in the video, or the presence of sound or speech in the video).
[0014] In one implementation of the first aspect, the video stream includes metadata, and the device is further configured to determine the video-level tags and the timestamped tags based on the metadata.
[0015] This ensures that various information sources in the video can be combined for analysis and tagging.
[0016] Specifically, the metadata includes at least one of the following: the title of the video stream, the duration of the video stream, a text description inserted by a user or website, comments, the number of likes and / or views, and geographical information about the upload of the video stream. In other words, the metadata of the video stream includes any information source other than audio information and the video stream.
[0017] Specifically, when the data is text, natural language processing (NLP) techniques can be used to process the metadata, and / or when the data includes structured data (e.g., GPS location or numbers), general machine learning techniques can be used to process the metadata.
[0018] In another implementation of the first aspect, the inference techniques include at least one of the following: automatic speech recognition (ASR), optical character recognition (OCR), computer vision (CV), natural language processing (NLP).
[0019] This is beneficial because multiple methods for determining tags can be used.
[0020] In another implementation of the first aspect, the device is further configured to store the video-level tags and the timestamped tags in a tag database of the device.
[0021] This is beneficial because the tags only need to be generated once and can be loaded when needed.
[0022] Specifically, the tag database is a database that stores tags associated with a video stream (video-level tags or timestamped tags with associated time ranges) for each video stream.
[0023] In another implementation of the first aspect, the device is further configured to provide the video-level tags to a user device that plays the video stream.
[0024] This is beneficial because the tags obtained by the device can be used on a user device (e.g., a mobile phone, a tablet, a laptop, or a desktop computer).
[0025] In another implementation of the first aspect, the device is further configured to receive a user input including playback time information from a user device that plays the video stream, and provide the timestamped tags to the user device based on the playback time information.
[0026] This ensures that only those timestamped tags corresponding to the playback time of the video playback on the user device (i.e., the time point of the currently displayed video) are provided to the user device.
[0027] In another implementation of the first aspect, the device is further configured to determine video-level recommendations based on the video-level tags and the resource index stored in the device; and / or determine timestamped recommendations based on the timestamped tags and the resource index.
[0028] This ensures that user-related recommendations can also be displayed based on the tags.
[0029] Specifically, the resource index includes a database of recommended resources (e.g., an index of videos or a list of advertisements, e.g., having metadata for recommendation pairing).
[0030] Specifically, the video-level recommendations and / or the timestamped recommendations include at least one of the following: URL (e.g., for initiating a search), related videos, related search queries, map locations, shopping items.
[0031] In another implementation of the first aspect, the device is further configured to store the video-level recommendations and the timestamped recommendations in a recommendation database of the device.
[0032] This ensures that the recommendations only need to be generated once and can be loaded from the database when needed.
[0033] Specifically, the recommendation database includes a database that stores recommended resources (e.g., related videos or suggested advertisements) for each video stream.
[0034] In another implementation of the first aspect, the device is further configured to provide the video-level recommendations to a user device that is playing the video stream.
[0035] This ensures that only relevant recommendations are provided to the user device.
[0036] In another implementation of the first aspect, the device is further configured to, in response to receiving the user input, provide the timestamped recommendations to the user device based on the playback time information.
[0037] This ensures that the specific time point of the currently displayed video is taken into account when providing recommendations to the user device.
[0038] In another implementation of the first aspect, the device is further configured to receive a request from a user device that is playing the video stream; based on the request, receive the video stream, determine the video-level tags and the timestamped tags, and directly provide the video-level tags and the timestamped tags to the user device.
[0039] This ensures that the video stream can be analyzed upon request.
[0040] In another implementation of the first aspect, the video-level tag includes information indicating the video stream as a whole.
[0041] This is beneficial because such a tag can indicate relevant information about the entire video stream.
[0042] Specifically, the video-level tag covers the content of the entire video stream.
[0043] In another implementation of the first aspect, the timestamped tag includes information indicating a specific time range of the video stream.
[0044] This is beneficial because such a tag can indicate information relevant at a specific time point in the video stream.
[0045] Specifically, the timestamped tag is related to a topic that appears only in intervals of the video. Specifically, the timestamped information consists of three elements: (1) start time, (2) end time, and (3) tag content related to the information that can be found between the start time and the end time.
[0046] A second aspect of the present invention provides a method for multimodal video analysis, wherein the method includes the following steps: a device receives a video stream; the device determines a video-level tag and a timestamped tag based on at least two frames of the video stream, audio information of the video stream, and inference techniques.
[0047] In one implementation of the second aspect, the video stream includes metadata, and the method includes: the device determines the video-level tag and the timestamped tag based on the metadata.
[0048] In another implementation of the second aspect, the inference techniques include at least one of the following: automatic speech recognition (ASR), optical character recognition (OCR), computer vision (CV), natural language processing (NLP).
[0049] In another implementation of the second aspect, the method further includes: the device stores the video-level tag and the timestamped tag in a tag database of the device.
[0050] In another implementation of the second aspect, the method further includes: the device providing the video-level tag to a user device that plays the video stream.
[0051] In another implementation of the second aspect, the method further includes: the device receiving a user input including playback time information from a user device that plays the video stream, and providing the timestamped tag to the user device based on the playback time information.
[0052] In another implementation of the second aspect, the method further includes: the device determining a video-level recommendation based on the video-level tag and a resource index stored in the device; and / or the device determining a timestamped recommendation based on the timestamped tag and the resource index.
[0053] In another implementation of the second aspect, the method further includes: the device storing the video-level recommendation and the timestamped recommendation in a recommendation database of the device.
[0054] In another implementation of the second aspect, the method further includes: the device providing the video-level recommendation to a user device that plays the video stream.
[0055] In another implementation of the second aspect, the method further includes: in response to receiving the user input, providing the timestamped recommendation to the user device based on the playback time information.
[0056] In another implementation of the second aspect, the method further includes: the device receiving a request from a user device that plays the video stream; the device receiving the video stream based on the request, the device determining the video-level tag and the timestamped tag, and the device directly providing the video-level tag and the timestamped tag to the user device.
[0057] In another implementation of the second aspect, the video-level tag includes information indicating the video stream as a whole.
[0058] In another implementation of the second aspect, the timestamped tag includes information indicating a specific time range of the video stream.
[0059] The second aspect and its implementations include the same advantages as the first aspect and its corresponding implementations.
[0060] A third aspect of the present invention provides a computer program including instructions that, when executed by a computer, cause the computer to perform the method according to the second aspect or any of its implementations.
[0061] The third aspect includes the same advantages as the first aspect and its corresponding implementation manners.
[0062] It should be noted that all devices, elements, units, and modules described in this application can be implemented in software or hardware elements or any combination thereof. All steps performed by various entities described in this application and the functions described to be performed by various entities are intended to indicate that the corresponding entities are suitable or used to perform the corresponding steps and functions. Even if, in the description of the following specific embodiments, the specific functions or steps to be performed by external entities are not reflected in the description of the specific detailed elements of the entity performing the specific steps or functions, those skilled in the art should clearly understand that these methods and functions can be implemented in the corresponding software or hardware elements, or any combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In conjunction with the accompanying drawings, the following description of specific embodiments will elaborate on the above aspects of the present invention and their implementation manners.
[0064] Figure 1 A schematic diagram of a device provided by an embodiment of the present invention is shown;
[0065] Figure 2 A schematic diagram of a device provided by an embodiment of the present invention is shown in detail;
[0066] Figure 3 Another schematic diagram of a device provided by an embodiment of the present invention is shown;
[0067] Figure 4 A detailed schematic diagram of creating tags based on reasoning technology provided by the present invention is shown;
[0068] Figure 5 A detailed schematic diagram of creating recommendations based on resource indexing provided by the present invention is shown;
[0069] Figure 6 A schematic diagram of an operation scenario provided by an embodiment of the present invention is shown;
[0070] Figure 7 A detailed schematic diagram of displaying tags and recommendations provided by the present invention is shown;
[0071] Figure 8 A schematic diagram of an operation scenario involving a server and a user device provided by an embodiment of the present invention is shown;
[0072] Figure 9 Another schematic diagram of an operation scenario involving a server and a user device provided by an embodiment of the present invention is shown;
[0073] Figure 10 Another schematic diagram of displaying tags and recommendations provided by the present invention is shown;
[0074] Figure 11 Shows another schematic diagram of the display label and recommendation provided by the present invention;
[0075] Figure 12 Shows a schematic diagram of the method provided by an embodiment of the present invention. Detailed implementation manners
[0076] Figure 1 Shows a schematic diagram of device 100. Device 100 is used for multimodal video analysis. For this purpose, device 100 is used to receive video stream 101. Device 100 is also used to determine video-level label 102 and timestamped label 103 based on at least two frames 104, 105 of video stream 101, audio information 106 of video stream 101, and inference technique 107.
[0077] Therefore, device 100 helps user navigation starting from the video using multimodal information. For this purpose, device 100 combines multimodal video understanding and analysis and video-based interactive search methods.
[0078] Device 100 facilitates predicting labels for a given video (i.e., video stream 101). These labels can indicate the video as a whole (i.e., video-level label 102) or indicate a specific time range (i.e., timestamped label 103). In other words, a label associated with a specific time range can be timestamped.
[0079] Specifically, video-level label 102 can be regarded as a label (e.g., a word) representing information about the content of video stream 101 as a whole. Specifically, timestamped label 103 can be regarded as a label (e.g., a word) representing information about a part of video stream 101 (e.g., a segment or an interval).
[0080] For example, for a user's vlog (i.e., video blog) video stream 101, video-level label 102 can just be "vlog" or "a day at school". In this case, the timestamped label can be "math class", which occurs, for example, in the morning, i.e., the first half of the video, or "PE class", which occurs after the math class, i.e., the second half of the video.
[0081] Labels 102, 103 can be inferred using multimodal methods, which means that all information channels in video stream 101 can be used, such as image frames 104, 105 constituting video stream 101, the playing audio information 106, possible voices, text metadata paired with the video, etc.
[0082] Specifically, the audio information 106 includes the audio information of the video stream 101. That is to say, the audio information 106 can be regarded as the audio stream of the video stream 101. For example, the audio information 106 can include at least one of sound, voice, music, and noise.
[0083] The tags 102, 103 can indicate a specific data modality (for example, for video tagging of the appearance of a specific object, for voice tagging of the topic under discussion) or a combination of multiple modalities (for example, tagging a specific dance based on dance-based movements and music). In these cases, information fusion can be performed.
[0084] The video-level tag 102 can include information indicating the video stream 102 as a whole. The timestamped tag 103 can include information indicating a specific time range of the video stream 102.
[0085] The output of the device 100 can be regarded as a list of tags 102, 103 that describe the video content (relative to the whole video or only a part of the video), and is determined by considering one or more information channels in the video stream 101.
[0086] Therefore, the device 100 also helps to implement an interface for simple and fast search starting from the viewed video stream 101. When the user is watching the video stream 101, they can press a button at any time to pop up the interface. The interface can include a set of tags 102, 103 and additional recommended content based on these tags (such as video recommendations or advertisements). The displayed tags can correspond to the most recently viewed part of the video stream 101. In other words, based on the moment when the button is pressed, the interface will change by using the timestamped tag 103. Intuitively, when the user sees something in the video stream 101 that they want to know more about, they will immediately press the button, and a list of tags 102, 103 that may include the target will appear. Then, the user can click to search in the selected search engine or view the recommended content suggestions.
[0087] Now, it will be combined with Figure 2 Device 100 will be described in more detail. Figure 2 The device 100 includes all the functions and features of the device 100 as described in combination with Figure 1 as described.
[0088] As Figure 2 shown, optionally, the video stream 101 can include metadata 201. The device 100 can determine the video-level tag 102 and the timestamped tag 103 based on the metadata 201. That is to say, compared with only the video information and audio information in the stream 101, the video-level tag 102 and the timestamped tag 103 can be determined based on more information.
[0089] As further shown in Figure 2 and optionally, as further shown in Figure 2 , device 100 may store video-level tag 102 and timestamped tag 103 in tag database 202 of device 100. From this tag database, the tags may be provided to user device 203, if necessary. That is, device 100 may provide video-level tag 102 to user device 203 that plays video stream 101.
[0090] To instruct device 100 to provide tags, device 100 may receive user input 204 including play time information 205 from user device 203 that plays video stream 101. After receiving user input 204, device 100 may provide timestamped tag 103 to user device 203 based on play time information 205. In other words, device 100 may provide those timestamped tags 103 that are relevant at the time point when video stream 101 is currently being played on user device 203.
[0091] In addition, optionally, and as also shown in Figure 2 Device 100 may determine video-level recommendation 206 based on video-level tag 102 and resource index 207 stored in device 100. Additionally or alternatively, device 100 may determine timestamped recommendation 208 based on timestamped tag 103 and resource index 207. This is also described in more detail in Figure 5 and is also described in more detail in Figure 5 .
[0092] Returning to Figure 2 and optionally, device 100 may store video-level recommendation 206 and timestamped recommendation 208 in recommendation database 209 of device 100, similar to the case of tags 102, 103 and tag database 202. From database 209, the recommendations may be provided to user device 203, if necessary.
[0093] That is, device 100 may provide video-level recommendation 206 to user device 203 that plays video stream 101. Additionally or alternatively, device 100 may provide timestamped recommendation 208 to user device 203 based on play time information 205 in response to receiving user input 204. That is, the device may precisely provide those timestamped recommendations 208 that are relevant to the time point when video stream 101 is displayed on user device 203.
[0094] As Figure 2As shown, device 100 may receive a request 210 from user device 203 that is playing video stream 101. Triggered by this request, based on request 210, the device may receive video stream 101, determine video-level label 102 and timestamped label 103, and directly provide video-level label 102 and timestamped label 103 to user device 203. In other words, device 100 may operate in an online mode, in which, once request 210 is received from user device 203, labels 102, 103 are determined for the first time.
[0095] Figure 3 An exemplary operation scenario of device 100 is shown, in which device 100 receives video stream 101 and metadata 201 from a network resource such as the Internet. An inference module employing inference technique 107 determines video-level label 102 and timestamped label 103 based on at least two frames 104, 105 of video stream 101, audio information 106 of video stream 101, and metadata 201. Then, video-level label 102 and timestamped label 103 are stored in label database 202 of device 100. This is done, for example, in an offline mode of device 100, in which the device receives video streams 101 from the Internet that are not simultaneously played by user device 203. Device 100 determines labels 102, 103 in its inventory and, so to speak, has them ready once a request to provide the labels is received. Device 100 also provides video-level label 102 and timestamped label 103 to a recommendation module, which determines video-level recommendation 206 based on video-level label 102 and resource index 207 stored in device 100, and additionally or alternatively, determines timestamped recommendation 208 based on timestamped label 103 and resource index 207.
[0096] In Figure 4 this scenario is described in detail, Figure 4 several inference techniques 107 are shown, such as automatic speech recognition, optical character recognition, computer vision, and natural language processing (NLP). The results of these techniques may be fused to determine labels 102, 103.
[0097] Returning to Figure 3 , then, device 100 may store video-level recommendation 206 and timestamped recommendation 208 in recommendation database 209 of device 100. Once device 100 receives a corresponding request, it may provide video-level recommendation 206 and timestamped recommendation 208 to user device 203.
[0098] Figure 6Shows another operating scenario of device 100, where user device 203 is playing video stream 101 that it has received from a network resource.
[0099] At a specific time point t, the user of user device 203 presses a button on user device 203, thereby providing user input 204 to device 100 that includes playback time information 205. Playback time information 205 includes the specific time point at which video stream 101 is currently being played on user device 203.
[0100] Once user input 204 that includes playback time information 205 is received from user device 203, timestamped label 103 is provided to user device 203 based on playback time information 205. Additionally or alternatively, video-level label 102 is also provided.
[0101] In response to receiving user input 204, timestamped recommendation 208 can also be provided to user device 203 based on playback time information 205. Additionally or alternatively, video-level recommendation 206 is also provided.
[0102] In the example shown, labels 102, 103 and recommendations 206, 208 are already in databases 202, 209 because they are generated, for example, in the offline mode of the device. That is, labels 102, 103 and recommendations 206, 208 can be generated in advance before user device 203 actually needs them. However, labels 102, 103 and recommendations 206, 208 can also be generated in the online mode, that is, exactly when they are needed by user device 203 that is currently playing video stream 101.
[0103] Figure 7 Shows an exemplary user interface that can be used by user device 101 to display labels 102, 103 and recommendations 206, 208 when playing video stream 101. In the left part of the figure, a video screen for playing video stream 101 is shown. The screen includes an activation button 701 and a tool 702 for indicating the current time and total duration of video stream 101. Activation button 701 can be used to generate user input 204, and then, user input 204 includes playback time information 205, and the playback time information corresponds to tool 702 that indicates the current time of video stream 101. In other words, activation button 701 can be regarded as an interface trigger that the user can activate to use device 100. It will send a request for labels 102, 103 and recommended content 206, 208 to the server (i.e., device 100) by sending necessary information (such as in the form of a video ID and the current viewing time). Then, the interface will display labels 102, 103 and content 206, 208 sent by the server, and display an interface with which the user can interact.
[0104] Once the device 100 receives the user input 204 and the tags 102, 103 and recommendations 206, 208 are provided to the user device 203, a pop-up interface is presented which, for example, lists the tags 103 and recommendations 208 related to the content of the video at the current time (i.e., the play time information 205). The tags and recommendations are clickable and can be opened based on the tag or recommendation, for example, to launch a search engine or a web browser for a search.
[0105] As has been described in connection with the above figures, particularly in Figure 3 and Figure 6 it is described that the device 100 can operate in an offline mode and / or an online mode.
[0106] For example, as Figure 3 shown, in the offline mode, the server (i.e., the device 100) repeatedly fetches popular videos (i.e., the video stream 101) from the network, performs analysis using the inference module to extract the tags 102, 103, and then uses those tags 102, 103 to retrieve the recommended content 206, 208 of the video stream 101. Then, the tags 102, 103 are stored in the tag database 202, and the recommended content 206, 208 is stored in the recommendation database 209. The server continuously repeats this process in order to pre-compute all the necessary information when the user requests.
[0107] For example, as Figure 6 shown, in the offline mode, on the user side, the user presses the activation button 701, and the video ID (e.g., URL) of the video and the current viewing timestamp are sent to the server. The server will check whether the video has been processed. If so, it will obtain the tags and the recommended content (compatible with the current timestamp) and send them to the user for display in the interface. If the video has not been processed, an error message can be sent to the user, or the online mode of the device 100 can be launched.
[0108] Figure 8 and Figure 9 show the online mode, the combined offline mode and online mode of the device 100 respectively.
[0109] As Figure 8 shown, the server (i.e., the device 100) only takes action when requested by the user (e.g., when receiving the request 210 from the user device 203). Similar to the offline mode, when the user is watching a video, the user presses the activation button 701 and sends the video ID and the current viewing timestamp to the server. Different from before, these will not be checked in the database, but the tags and recommendations will be computed instantaneously in the same process as in the offline version.
[0110] The online mode and the offline mode can also be used in combination, as Figure 9As shown. Similar to the offline mode, the server repeatedly processes the video stream 101 and stores the output in databases 202, 209. In contrast, if the user requests an unprocessed video stream 101, it will not throw an error but will dynamically calculate tags 102, 103 and recommendations 206, 208. The results will be sent directly to the user device 203 (as in the online version), but will also be stored in databases 202, 209 for future requests for the same video 101.
[0111] Figure 9 and Figure 10 The changes in the user interface of the user device 203 are shown by examples demonstrating the possible interactions of the user with the user device 203 and the device 100.
[0112] When watching the video stream 101, the user can press the activation button 701 (left). A pop-up interface (right) may appear, showing the timestamped tag 103 and the timestamped recommendation 208. The tag 103 is related to the most recently viewed content of the video stream 101 (e.g., since a free kick is being shown, the tag "free kick" will appear). Other tags 103 can include many different types of recent content, such as logos that appear on the screen or what the commentary in the video audio is talking about. The recommendation 208 is also related to the recent content.
[0113] Based on the moment when the activation button 701 is pressed, the tag 103 and the recommendation 208 will change, as Figure 11 shown, thus giving the user control over what the tag 103 should focus on. Now that the video stream 101 is focused on a player close-up instead of a football action, the tag 103 will show the player's name, and the recommended content will change to reflect this. Additionally, by clicking on the tag 103, the user can display search results on an external search engine, making navigation faster and easier. The search engine can be selected by the user or suggested by the interface (e.g., clicking on the name of a town can suggest opening with a map application or a travel search engine).
[0114] Figure 12 A schematic diagram of the method 1200 is shown. The method 1200 is for multimodal video analysis. The method 1200 includes a first step: the device 100 receives (1201) the video stream 101. The method 1200 includes a second step 1202: the device 100 determines the video-level tag 102 and the timestamped tag 103 based on at least two frames 104, 105 of the video stream 101, the audio information 106 of the video stream 101, and the inference technique 107.
[0115] The present invention has been described by way of example and in terms of implementations in conjunction with various embodiments. However, those skilled in the art will be able to understand and achieve other variations when practicing the claimed invention based on a study of the drawings, the present invention, and the independent claims. In the claims as well as the specification, the word "comprising" does not exclude other elements or steps, and "a" does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items described in the claims. Stating certain measures in mutually different dependent claims does not indicate that a combination of these measures cannot be used effectively.
Claims
1. An apparatus (100) for multimodal video analysis, characterized in that, The device (100) is used for: - Receiving a video stream (101); - Determining a video-level tag (102) and a timestamped tag (103) based on at least two frames (104, 105) of the video stream (101), audio information (106) of the video stream (101), and an inference technique (107).
2. The device (100) according to claim 1, characterized in that, The video stream (101) includes metadata (201), and the device (100) is further used for determining the video-level tag (102) and the timestamped tag (103) based on the metadata (201).
3. The device (100) according to claim 1 or 2, characterized in that, The inference technique (107) includes at least one of the following: automatic speech recognition (ASR), optical character recognition (OCR), computer vision (CV), natural language processing (NLP).
4. The device (100) according to any one of the above claims, characterized in that, It is also used for storing the video-level tag (102) and the timestamped tag (103) in a tag database (202) of the device (100).
5. The device (100) according to any one of the above claims, characterized in that, It is also used for providing the video-level tag (102) to a user device (203) that plays the video stream (101).
6. The device (100) according to any one of the preceding claims, characterized in that, It is also used for receiving a user input (204) including playback time information (205) from the user device (203) that plays the video stream (101), and providing the timestamped tag (103) to the user device (203) based on the playback time information (205).
7. The device (100) according to any one of the preceding claims, characterized in that, It is also used for determining a video-level recommendation (206) based on the video-level tag (102) and a resource index (207) stored in the device (100); and / or determining a timestamped recommendation (208) based on the timestamped tag (103) and the resource index (207).
8. The device (100) according to claim 7, characterized in that, It is also used for storing the video-level recommendation (206) and the timestamped recommendation (208) in a recommendation database (209) of the device (100).
9. The device (100) according to any one of the preceding claims, characterized in that, It is also used for providing the video-level recommendation (206) to the user device (203) that plays the video stream (101).
10. The device (100) according to any one of the preceding claims, characterized in that, It is also used for providing the timestamped recommendation (208) to the user device (203) based on the playback time information (205) in response to receiving the user input (204).
11. The device (100) according to any one of the above claims, characterized in that, It is also used for receiving a request (210) from the user device (203) that plays the video stream (101), Based on the request (210), receiving the video stream (101), determining the video-level tag (102) and the timestamped tag (103), and directly providing the video-level tag (102) and the timestamped tag (103) to the user device (203).
12. The device (100) according to any one of the preceding claims, characterized in that, The video-level tag (102) includes information indicating the video stream (102) as a whole.
13. The device (100) according to any one of the preceding claims, characterized in that, The timestamped tag (103) includes information indicating a specific time range of the video stream (102).
14. A method (1200) for multimodal video analysis, characterized in that, The method (1200) includes the following steps: - A device (100) receives (1201) a video stream (101); - The device (100) determines (1202) a video-level tag (102) and a timestamped tag (103) based on at least two frames (104, 105) of the video stream (101), audio information (106) of the video stream (101), and an inference technique (107).
15. A computer program comprising instructions, characterized in that, When the program is executed by a computer, the instructions cause the computer to perform the method (1200) according to claim 14.