Method and apparatus for annotating emotion-related metadata to multimedia files

By performing data analysis and sentiment labeling on multimedia objects, knowledge graphs and importance scores are generated, solving the problem of time-consuming and labor-intensive traditional methods, and achieving efficient multimedia content selection and summary generation.

CN114503100BActive Publication Date: 2025-10-24HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080071054.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-01-30
Publication Date
2025-10-24
Estimated Expiration
2040-01-30

AI Technical Summary

Technical Problem

Traditional video search tools cannot fully reflect the content, and traditional movie compression methods are time-consuming, labor-intensive, and do not meet user needs, making it difficult for users to find content that matches their personal preferences from a large number of online videos.

Method used

By performing data analysis on multimedia objects, identifying scene-related emotions, generating knowledge graphs and calculating scene importance scores, selecting key scene subsets, and generating summaries of multimedia objects.

Benefits of technology

It enables users to select multimedia content based on their interests, improving content search efficiency, reducing manpower and time costs, and generating meaningful movie scene descriptions and summaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114503100B_ABST
    Figure CN114503100B_ABST
Patent Text Reader

Abstract

Apparatuses, systems, architectures, and methods for summarizing multimedia objects with a primary narrative are described. Methods and processes can include performing data analysis to identify scene-relevant sentiments indicated in scenes of the multimedia object; generating a knowledge graph that associates each scene in the scenes with a respective one or more scene-relevant sentiments; computing a plurality of scores using the knowledge graph, each score indicating a relative importance of one of the plurality of scenes to conveying the primary narrative; selecting a subset of the plurality of scenes according to the plurality of scores; and / or generating a summary of the multimedia object according to the subset.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to the field of media technology and more specifically, but not exclusively, to movie annotation software. BACKGROUND

[0002] And The advent of streaming media and video on demand platforms has provided the movie industry with new tools to deliver products to targeted audiences. As a result, the amount of online video content available to users has increased dramatically. One of the consequences of this is that it is increasingly difficult for viewers to find content that is relevant to their personal preferences from the wealth of content. Conventional video search tools are only able to find movie content by generic indicators that do not comprehensively reflect the content. Additionally, subjectively, conventional methods of producing compressed versions of movies typically require significant human effort. These methods also typically take a significant amount of time and do not necessarily translate well between users. SUMMARY

[0003] It is an object of the present invention to provide apparatus, systems and methods for enriching movie metadata. It is an object of the present invention to provide apparatus, systems and methods of creating meaningful, semantically rich movie scene descriptions. It is an object of the present invention to provide apparatus, systems and methods for efficiently generating compressed versions of multimedia objects. It is an object of the present invention to provide apparatus, systems and methods for creating abridged versions of movies that preserve the overall tone and essential narrative of the movie. It is an object of the present invention to provide apparatus, systems and methods for searching and categorizing multimedia content according to indicators that more comprehensively represent the content. It is an object of the present invention to facilitate the selection of preferred multimedia object types by viewers according to specific interests, tones, patterns and / or moods.

[0004] The foregoing and other objects are achieved by the features of the independent claims. Further implementation forms are evident from the dependent claims, the description and the figures.

[0005] According to a first aspect of the present invention, there is provided a system for summarizing a multimedia object having a main narrative, comprising: a processor executing readable instructions to: perform data analysis to identify one or more scene-related moods indicative in each of a plurality of scenes of the multimedia object; generate a knowledge graph associating each of the plurality of scenes with a respective one or more scene-related moods; compute a plurality of scores using the knowledge graph, each score indicative of a relative importance of one of the plurality of scenes to convey the main narrative; select a subset of the plurality of scenes according to the plurality of scores; generate a summary of the multimedia object according to the subset.

[0006] According to a second aspect of the present application, there is provided a method for summarizing a multimedia object having a primary narrative, comprising: performing data analysis to identify one or more scene-related sentiments in each of a plurality of scenes of the multimedia object indicative of the scene; generating a knowledge graph associating each of the plurality of scenes with a respective one or more scene-related sentiments; computing a plurality of scores using the knowledge graph, each score indicative of a relative importance of one of the plurality of scenes to convey the primary narrative; selecting a subset of the plurality of scenes according to the plurality of scores; generating a summary of the multimedia object according to the subset.

[0007] In implementations of the various aspects of the present application, the data analysis comprises pre-processing of the multimedia object, the pre-processing comprising extracting from the multimedia object at least one of: a video file; a subtitle file; a chapter text file detailing start times of chapters of the multimedia object; audio file segments of actor speech and non-speech portions.

[0008] In possible implementations of the various aspects of the present application, the data analysis comprises: scrubbing associated metadata describing the multimedia object; analyzing the associated metadata to indicate the one or more scene-related sentiments.

[0009] In possible implementations of the various aspects of the present application, the data analysis comprises implementing semantic lifting according to a scene ontology to capture raw multimedia information of the multimedia object.

[0010] In possible implementations of the various aspects of the present application, the data analysis comprises interlinking with external sources describing features of the multimedia object.

[0011] In possible implementations of the various aspects of the present application, the features comprise at least one of: a scene of the multimedia object; an activity in the multimedia object scene; an actor performing in the multimedia object; a character depicted in the multimedia object.

[0012] In possible implementations of the various aspects of the present application, the data analysis comprises analyzing a descriptive audio soundtrack of the multimedia object to indicate the one or more scene-related sentiments.

[0013] In possible implementations of the various aspects of the present application, the data analysis comprises extracting the sentiments from visual sentiment indicators, the visual sentiment indicators comprising at least one of: a facial expression image; a body posture image; a video sequence of sentiment-indicative behavior.

[0014] In a possible implementation form of the various aspects of the application, the data analysis comprises extracting the emotion from an auditory emotional indicator, the auditory emotional indicator comprising at least one of: an emotional representation of a music soundtrack; an emotionally suggestive vocal indicator.

[0015] In a possible implementation form of the various aspects of the application, the data analysis comprises extracting the emotion from a textual emotional indicator, the textual emotional indicator comprising at least one of: an explicit emotional descriptor; an implicit emotional indicator.

[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the application, exemplary methods and / or materials are described below. In case of conflict, the patent specification will control. In addition, the materials, methods, and examples are illustrative only and are not necessarily intended to be limiting. BRIEF DESCRIPTION OF DRAWINGS

[0017] Some embodiments of the application are herein described, by way of example only, with reference to the accompanying drawings. With specific reference now to the drawings in

[0018] In the drawings:

[0019] Figure 1A schematic flow chart of an optional operational flow provided for some embodiments of the application;

[0020] Figure 1B schematic diagram of an exemplary system provided for some embodiments of the application;

[0021] Figure 1C schematic diagram of an exemplary system provided for some embodiments of the application;

[0022] Figure 2 schematic diagram of an exemplary system architecture provided for some embodiments of the application;

[0023] Figure 3 schematic diagram of an exemplary system architecture provided for some embodiments of the application;

[0024] Figure 4 schematic diagram of an exemplary system architecture provided for some embodiments of the application;

[0025] Figure 5A schematic diagram of an exemplary system architecture provided for some embodiments of the application; Figure 4 ​

[0026] Figure 5B To express Figure 4 Schematic diagrams of various aspects of an exemplary system architecture;

[0027] Figure 6 A schematic diagram of an exemplary system architecture provided for some embodiments of the present invention;

[0028] Figure 7 A schematic diagram of an exemplary system architecture provided for some embodiments of the present invention;

[0029] Figure 8 A schematic diagram of an exemplary system architecture provided for some embodiments of the present invention;

[0030] Figure 9 A schematic diagram of an exemplary system architecture provided for some embodiments of the present invention;

[0031] Figure 10 A schematic diagram of an exemplary system architecture provided for some embodiments of the present invention;

[0032] Figure 11 A schematic diagram of an exemplary system architecture provided for some embodiments of the present invention;

[0033] Figure 12A A schematic diagram of an exemplary system architecture provided for some embodiments of the present invention;

[0034] Figure 12B for Figure 12A A schematic diagram of the scenario indicated by the architecture. DETAILED DESCRIPTION

[0035] Before explaining at least one embodiment of the present invention in detail, it should be understood that the present invention is not necessarily limited to the details of construction and arrangement of components and / or methods described in the following description and / or drawings and / or examples. The present invention has other embodiments or can be practiced or carried out in various ways.

[0036] Embodiments of the present invention include one or more devices, one or more systems, one or more methods, one or more architectures, and / or one or more computer program products. The computer program product may include a computer-readable storage medium having computer-readable program instructions that cause a processor to perform various aspects of the present invention.

[0037] The computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. The computer-readable storage medium may be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above devices.

[0038] The computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network.

[0039] The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to customize the electronic circuitry, in order to perform aspects of the present application.

[0040] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0041] The flow diagrams and / or block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams and / or block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical functions (‘instructions’). In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0042] Each media segment in a multimedia object, such as a movie or a video, can have a specific emotion that describes the feeling of the movie at the current time. Various aspects of the invention can include a metadata model, such as a Resource Description Framework (RDF) schema, defined and developed as a basis for a video content knowledge graph.

[0043] Various aspects of the invention can include annotation toolkits, algorithms, processes, tools, interfaces, and / or operations. In some embodiments of the invention, one or more annotators can be responsible for filling in observable actions and emotions in a semantic model. Annotators can be familiar with characters and plots. Annotators can attempt to label portions of a movie according to activities and emotions.

[0044] Emotions are an important aspect of a movie, which are influenced by various factors. Annotation of emotions can be based on different data sources, such as background music. Movie scores best reflect the intended emotions of the production team, rather than the emotions of the characters or the viewer.

[0045] Deciding at what granularity to show details is not an easy task. For example, in a fight scene, each punch in the scene can be annotated as a separate activity, or, the movie can be annotated with a large portion of the segment as “characters x and y fight.”

[0046] Methods and apparatus of the invention can include or involve: tagging of multimedia files with visual and / or audio emotional cues metadata; apparatus for annotating a movie according to emotions; features that utilize emotion-based annotations; rating of importance of scenes and / or sub-scenes according to diversity of elements between main characters, emotions, and / or adjacent scenes; and / or storing emotion-based annotations as structured information of a knowledge graph, such as in Resource Description Framework (RDF) triples.

[0047] Embodiments of the invention can be used to enrich movie metadata with the richness of semantic technology and / or to create meaningful and semantically rich movie scene descriptions using various video and / or audio processing techniques. The invention can include a scene ontology that semantically defines concepts that capture moments in a movie, such as emotion-based moments, in order to capture the raw multimedia information of a video using meaningful and semantically rich representations. For example, semantic information can be extracted from descriptive audio soundtracks, such as soundtracks for visually impaired viewers, and / or from movie subtitles, such as lines spoken by actors. The invention can include creating a knowledge graph that can be queried to select video summaries.

[0048] In some embodiments, the emotion-based annotation method can utilize structured information of a knowledge graph, which can be stored in RDF triples that can be queried using the W3C standardized SPARQL Protocol and RDF Query Language (SPARQL). A semantic model or ontology can represent data in the knowledge graph, facilitating various features, properties, and resource identifiers available in the knowledge graph to movie-related scenes.

[0049] Embodiments of the present invention include systems, architectures, apparatuses, and methods for movie and / or video annotation. The systems and architectures can relate to one or more steps of the methods and / or can include one or more features of the apparatuses.

[0050] The systems include a system to summarize a multimedia object. The methods include a method to summarize a multimedia object. The multimedia object can include a movie or any other suitable audio and / or video object, such as a video game or a music file. The multimedia object can include a main narrative. The multimedia object can include video data. The multimedia object can include audio data. The term "multimedia" as used herein refers to one or more expressive media, such as audio and / or video. Summarizing the multimedia object can include generating a knowledge graph of the multimedia object.

[0051] The systems can include a processor to execute machine-readable instructions for implementing the methods. The systems can include a memory to store the machine-readable instructions. Alternatively and / or additionally, the methods can be manual, automatic, and / or partially automatic.

[0052] The methods can include performing one or more data analyses to identify one or more scene-related emotions and / or actions indicated in one or more movie scenes of the multimedia object. The methods can include generating a knowledge graph. The knowledge graph can associate one or more scenes and a corresponding one or more scene-related emotions.

[0053] The methods can include calculating a score that indicates a relative importance of one or more scenes for conveying the main narrative. For example, a low score can indicate a lack of main characters and / or thematic elements; can indicate a narrative that is not strongly related to a theme; and / or can indicate superfluous thematic elements, while a high score can indicate a presence of main characters and / or thematic elements; can indicate a main narrative; and / or can indicate a transitional thematic element, where, for example, a narrative theme can undergo a significant shift before and after a thematic climax. The calculation can be made in accordance with the use of the knowledge graph.

[0054] The method can include selecting a subset of the scenes according to the scores. The subset can include higher scored scenes. The method can include generating a summary of the multimedia object according to the subset.

[0055] The method can include preprocessing of the multimedia object. The analysis can include preprocessing of the multimedia object. The preprocessing can include extracting stored data from the multimedia object. The stored data can include one or more of video files, subtitle files, chapter text files detailing start times of chapters of the multimedia object, audio file segments of actor speech and / or non-speech portions, and / or any other suitable data source. The preprocessing can provide some or all of the information for the analysis. The preprocessing can provide technical improvements, richer scene feature indicators that facilitate detection of key emotions, and other key features related to the object summary, such as actions and / or activities. The term "activities" as used herein refers to things that a viewer can observe happening in the multimedia object.

[0056] The analysis can include erasing associated metadata. The associated metadata can describe the multimedia object. The stored data can include the associated metadata. The associated metadata can provide some or all of the information for the analysis. Including the associated metadata can provide improvements by generating emotion-related information that is not recorded or easily extracted from other stored data files, and does not require subjective input from individual users.

[0057] The analysis can include using a semantic web system, using existing and / or novel vocabularies, domain-specific ontologies, and / or World Wide Web Consortium recommended RDF, RDFS, OWL, R2RML, and / or SPARQL, among other techniques, to convert structured and / or semi-structured information into linked data, such as linked open data. The analysis can include implementing semantic enrichment. The semantic enrichment can capture raw multimedia information of the multimedia object. The semantic enrichment can follow a scene ontology. The semantic enrichment can be used to semantically represent video content and / or support user analysis. The semantic enrichment can provide improvements by capturing raw multimedia information of the multimedia object in a meaningful and semantically rich manner that can not be convertible for different people and / or machine readers, for example, in terms of emotion detection and / or object summarization. The scene ontology can facilitate a knowledge graph that has semantically rich and / or machine interoperable properties. The scene ontology can facilitate a knowledge graph that comprehensively describes content of the multimedia object.

[0058] The method can include correlating data. The method can include publishing data, such as publishing data on the Internet. The method can include interconnecting. The method can include interconnecting data between data sources. Correlating data can be configured to facilitate access by a semantic browser. Correlating data can facilitate navigation between data sources through a resource description framework (RDF) correlation. Data sources can include external sources that describe characteristics of multimedia objects. Data sources can include stored data. Interconnecting can improve a knowledge graph with richer semantics of external information.

[0059] Described multimedia object characteristics can include one or more characteristics or attributes of at least one of: one or more scenes of the multimedia object; activity in the multimedia object scene; actors performing in the multimedia object; and / or characters depicted in the multimedia object. Including characteristics in the analysis can improve a knowledge graph that is semantically rich, can comprehensively describe content of the multimedia object, and / or can generate reliable summaries of the multimedia object, as well as for key narratives.

[0060] Stored data can include a descriptive audio soundtrack of the multimedia object. Stored data can include elements of the descriptive audio soundtrack. The soundtrack can improve in more directly representing emotions that the creator of the multimedia object intended to express, rather than relying solely on subjective user input.

[0061] The analysis can include extracting emotions from visual emotional indicators. Visual indicators can include facial expression images, body pose images, and / or video sequences of emotional indicating behaviors. The visual emotional indicators can improve by generating rich social and / or emotional information that is not available from other sources, thereby generating a semantically rich knowledge graph that can comprehensively describe content of the multimedia object, and / or can generate reliable summaries of the multimedia object, as well as for key narratives.

[0062] The analysis can include extracting emotions from auditory emotional indicators, such as musical soundtracks that indicate emotions and / or sound indicators that indicate emotions. Auditory emotional indicators can improve by generating rich social and / or emotional information that is not available from other sources, thereby generating a semantically rich knowledge graph that comprehensively describes content of the multimedia object, and generates reliable summaries of the multimedia object, as well as for key narratives.

[0063] The analysis can include extracting emotions from text emotion indicators, such as explicit emotion descriptors and / or implicit emotion indicators. These text emotion indicators can be improved by generating emotion information that is otherwise implicit, thereby generating a semantically rich knowledge graph that comprehensively describes the content of a multimedia object and can generate a reliable summary that is applicable to a key narrative.

[0064] The method can include a method of dividing a movie into component scenes. The scenes can include one or more actions and / or activities. The method can include tagging one or more scenes in the component scenes. The tagging can be performed according to a detected emotion that is fixed in the one or more scenes.

[0065] As described above, the system and architecture of the present invention can involve one or more steps of the method and / or can include one or more features of the apparatus. The architecture can include one or more features of the system. The method of the present invention can include and the system of the present invention can involve a method of annotating and / or summarizing a multimedia object. The term "multimedia" as used herein refers to one or more expressive media, such as audio and / or video. One or more steps of the method can be manual, automatic, and / or partially automatic. The method can include one or more machine learning processes.

[0066] Reference is made to Figure 1A , Figure 1A An exemplary multimedia object summarization process 100A is shown. The method can include the summarization process 100A.

[0067] A multimedia object can include one or more movies or any other suitable audio and / or video object, such as a video game or a music file. A multimedia object can include a primary narrative. A multimedia object can include video data. A multimedia object can include audio data. Summarizing a multimedia object can include generating one or more knowledge graphs of the multimedia object.

[0068] The process 100A can begin at step 101. At step 101, one or more data analyses can be performed. The method can include performing a data analysis, such as described in step 101. The data analysis can identify one or more scene-related emotions and / or actions indicated in one or more movie scenes of the multimedia object. The first column of Table 1 below shows exemplary emotions.

[0069] The analysis can involve one or more sets of stored data. The stored data can include information of a multimedia object. The multimedia object can include the stored data. The stored data can include one or more video files, caption files, chapter text files detailing start times of chapters of the multimedia object, audio file segments of actor speech and / or non-speech portions, and / or any other suitable data source.

[0070] The stored data can include a descriptive audio soundtrack of the multimedia object. The stored data can include elements of the descriptive audio soundtrack. The soundtrack can improve in representing more directly the mood that the creator of the multimedia object intended to express, rather than relying solely on subjective user input.

[0071] The method can include pre-processing of the multimedia object. The analysis can include pre-processing of the multimedia object. The pre-processing can include extracting the stored data from the multimedia object. The pre-processing can provide some or all of the information for the data analysis. The pre-processing can provide a technical improvement, richer indicators of characteristics of scenes can help detect key moods, and other key features related to the object summary, such as actions and / or activities. The term "activities" as used herein refers to what a viewer can observe as happening in the multimedia object.

[0072] The data analysis can include erasing associated metadata. The associated metadata can describe the multimedia object. The stored data can include the associated metadata. The associated metadata can provide some or all of the information for the data analysis. Including the associated metadata can provide an improvement by generating mood-related information that is not recorded or easily extracted from other stored data files, and does not require subjective input from individual users.

[0073] The data analysis can include utilizing semantic web systems, utilizing existing and / or novel vocabularies, domain-specific ontologies, and / or language technologies such as World Wide Web Consortium recommended RDF, RDF Schema (RDFS), Web Ontology Language (OWL), R2RML, and / or SPARQL to convert structured and / or semi-structured information into linked data, such as linked open data.

[0074] The data analysis can include implementing semantic enrichment. The semantic enrichment can capture the raw multimedia information of the multimedia objects. The semantic enrichment can follow a scene ontology. The semantic enrichment can be used to represent the video content in a semantic way and / or support user analysis. The semantic enrichment can provide improvements by capturing the raw multimedia information of the multimedia objects in a meaningful and semantically rich way, which can not be convertible for different people and / or machine readers, e.g., in terms of emotion detection and / or object summarization. The scene ontology can help the knowledge graph to have semantically rich and / or machine interoperable properties. The scene ontology can help the knowledge graph to comprehensively describe the content of the multimedia objects.

[0075] The method can include linking data. The method can include publishing data, e.g., publishing data on the Internet. The method can include interlinking. The method can include interlinking data between data sources. The linked data can be configured to facilitate access through a semantic browser. The linked data can help navigation between data sources through a resource description framework (RDF). The data sources can include external sources that describe features of the multimedia objects. The data sources can include stored data. The interlinking can improve the knowledge graph to have semantically richer external information.

[0076] The interlinking can help the knowledge graph to have semantically richer external information, which can not be available from the movie content or the movie metadata. For example, the metadata of a movie can only provide the association between the actors and the fictional characters depicted. Interlinking the actor entities with the corresponding entities in external sources can provide more information about the actors, such as gender, birth date, etc.

[0077] The external sources can be interlinked with data nets, etc. The external sources can include public cross-domain knowledge graphs, such as DBPEDIA, WIKIDATA, and / or YAGO. For example, WIKIDATA can be used for movie data and personal data, such as actors and movie crews. Data association of other entities, such as the city and country of a certain scene or segment, can be determined according to the resources of DBPEDIA.

[0078] The apparatus and method of the present disclosure can include a mechanism that helps interlink entities with the most suitable external sources. The interlinking can be expressed in SPARQL.

[0079] The described multimedia object features can include one or more features or attributes of at least one of: one or more scenes of the multimedia object; activity in the multimedia object scenes; actors performing in the multimedia object; and / or characters depicted in the multimedia object. Including features in the data analysis can be improved by providing a semantically rich knowledge graph that can comprehensively describe the content of the multimedia object and / or generate a reliable summary of the multimedia object, also for key narratives.

[0080] The data analysis can include extracting emotions from visual emotion indicators. The stored data can include visual emotion indicators. The visual indicators can include facial expression images, body posture images, and / or video sequences of emotion-indicative behavior. The visual emotion indicators can be improved by generating rich social and / or emotional information that can not be available from other sources, thereby generating a semantically rich knowledge graph that can comprehensively describe the content of the multimedia object and / or generate a reliable summary of the multimedia object, also for key narratives.

[0081] The data analysis can include extracting emotions from auditory emotion indicators, such as emotion-indicative music soundtracks and / or sound indicators. The stored data can include auditory emotion indicators. The auditory emotion indicators can be improved by generating rich social and / or emotional information that can not be available from other sources, thereby generating a semantically rich knowledge graph that comprehensively describes the content of the multimedia object and generates a reliable summary of the multimedia object, also for key narratives.

[0082] The data analysis can include extracting emotions from textual emotion indicators, such as explicit emotion descriptors and / or implicit emotion indicators. These textual emotion indicators can be improved by generating otherwise implicit emotional information, thereby generating a semantically rich knowledge graph that comprehensively describes the content of the multimedia object and can generate a reliable summary suitable for key narratives.

[0083] Various aspects of the described method can be implemented through one or more user interfaces. The user interface can include a graphic user interface (GUI). The user interface can include one or more interface features. The interface features can include widgets and / or virtual buttons. The interface features can be dedicated to one or more steps of the method. The interface features can be used to avoid subjective individual user input.

[0084] Interface features can facilitate generating an emotion-based knowledge graph that is queryable for video summaries. The emotion-based knowledge graph can be used to create a cinematic summary form that preserves the overall tone and story. The method can include creating an emotion-based knowledge graph with annotation information to select the most informative scenes in a movie summary.

[0085] The method can involve a knowledge-based system that includes a knowledge base representing information of a multimedia object. The method can include representing and / or defining categories, attributes, and relationships between concepts, data, and / or entities that validate the multimedia object. The knowledge-based system can include one or more inference engines to derive new information and / or discover inconsistencies. The method can include generating a knowledge graph in step 103. The knowledge graph can associate one or more scenes with a corresponding one or more scene-related emotions. Various aspects of the invention can avoid the need for user sentiment or emotion to create a knowledge graph.

[0086] The method can include calculating scores in step 105 that indicate a relative importance of one or more scenes for conveying a primary narrative. For example, a lower score can indicate a lack of a primary character and / or a thematic element; can indicate a narrative that is not strongly related to a theme; and / or can indicate a redundant thematic element. A higher score can indicate a presence of a primary character and / or a thematic element; can indicate a primary narrative; and / or can indicate a transitional thematic element, where, for example, a narrative theme can undergo a significant shift before and after a thematic climax. The scores can be calculated from the knowledge graph.

[0087] The method can include selecting a subset of scenes in step 107 according to the scores. The subset can include scenes with higher scores. The selection can include setting one or more thresholds to include in the subset.

[0088] The method can include generating a summary of the multimedia object from the subset. Generating the summary can include merging only scenes with scores that satisfy a threshold. Generating the summary can include filtering out scenes with scores that do not satisfy a threshold. Generating the summary can include deleting scenes with scores that do not satisfy a threshold.

[0089] Systems and architectures of the invention can include, and methods and processes of the invention can include, a system for annotating and / or summarizing a multimedia object. The system can include a processor for executing machine-readable instructions for implementing the method. The system can include a memory for storing the machine-readable instructions.

[0090] Reference is made to Figure 1B FIG. 10B illustrates a software system 100B. The system can include any or all of the features of system 100B. The system can be used to perform any or all of the methods described herein. Figure 1AAny or all of the steps of the illustrated method 100A. The system can include one or more modules for performing one, any or all of the steps of the method 100A. The term "module" as used herein refers to one or more software components and / or one or more portions of one or more programs, and can also include and / or be related to hardware for executing the software components and / or program portions. The hardware can include a processor that executes instructions and / or a memory that stores the instructions. The program can contain one or more routines. The program can include one or more modules. As explained in the following paragraphs, representations of the modules can be used to illustrate functional features of the system architecture of the system embodying the application and / or the method implementing the application. The modules can be incorporated into the program and / or software through one or more interfaces. The instructions executed by the processor can include one or more modules.

[0091] Figure 1B The illustrated system 100B is used for annotating multimedia files based on visual and / or audio emotional cues therein. The system 100B contains two main parts, namely a semantic enrichment part 102 and a video summarization part 104. The semantic enrichment part 102 creates a knowledge graph based on emotions and / or activities. The system 100B can be used to implement one or more steps of the process 100A illustrated in Figure 1A

[0092] The semantic enrichment part 102 can include an automatic media processing module 106. The media processing module 106 can be used to automatically pre-process a user-selected multimedia object. The module 106 can be used to perform any or all of the pre-processing steps described with respect to the process 100A. The module 106 can extract video, audio and / or subtitles from a movie. The module 106 can extract chapter text files, audio files, actor sound clips, etc. that detail the start times of the chapters of the movie.

[0093] The semantic enrichment part 102 can include a natural language processing (NLP) module 108. The natural language processing module 108 can be used to locate and classify named entity data processed by the module 106. The module 108 can be used to perform named entity recognition on unstructured multimedia data processed by the module 106.

[0094] The semantic enrichment part 102 includes an annotation tool module 110. The module 110 can include the graphical user interface described above with respect to step 101 of the process 100A. The module 110 facilitates user selection of a movie for processing. The module 110 can facilitate user interaction with the modules 106 and 108. The module 110 can include any or all of the features of the graphical user interface described with respect to step 101 of the process 100A.

[0095] ​The semantic enrichment portion 102 includes an ontology module 112. The ontology module 112 facilitates the data analysis performed by the modules 108 and 110. The ontology module 112 can facilitate the classification of the named entity data processed by the module 106. The module 112 can facilitate the named entity recognition of the unstructured multimedia data processed by the module 106. The ontology module 112 can facilitate the semantic definition of concepts that capture moments in the user-selected movie, such as emotion-based moments, to obtain the raw multimedia information of the movie using meaningful and semantically rich representations. For example, semantic information can be extracted from the descriptive audio soundtrack of the movie (e.g., a soundtrack for visually impaired viewers) and / or from the movie subtitles (e.g., the lines spoken by actors and / or visual cues for hearing impaired viewers). The semantic information can be extracted from one or more transcripts of one or more audio portions of the multimedia object. The transcripts can include descriptions of speech and / or non-speech elements. The transcripts can include one or more languages. The semantic information can be extracted from closed captions, open captions, and / or subtitles. The semantic information can be extracted from translations of dialog, sound effects, related musical cues, and / or any other suitable related audio data.

[0096] The semantic enrichment portion 102 includes a semantic enrichment module 114 and an automatic metadata module 116. The automatic metadata module 116 is used to automatically erase metadata from the movie file of the movie selected by the user for processing by the semantic enrichment module 114. The automatic metadata module 116 is used to erase movie metadata without the need for pre-processing by the modules 106 and 108.

[0097] The semantic enrichment module 114 semantically enriches the data and facilitates the semantic representation of the movie content. The semantic enrichment module 114 is used to process the raw multimedia information of the multimedia object and / or the processed multimedia information directly through the annotation tool module 110, through the ontology module 112, and / or through the automatic metadata module 116. The semantic enrichment module 114 supports user analysis through the annotation tool module 110.

[0098] The semantic enrichment portion 102 includes an external source interlinking identification module 118. The external source interlinking identification module 118 can facilitate the interlinking with external sources of movie metadata. For example, the external source interlinking identification module 118 can facilitate the erasure of movie metadata from one or more movie database application programming interfaces (e.g., IMDB, DBpedia1, Wikidata2, and / or other open source data). The interlinking with external sources can provide the knowledge graph with semantically richer external information that can not be available from the movie content or the movie metadata.

[0099] The semantic enrichment portion 102 includes a movie knowledge graph generation module 120. The knowledge graph generation module 120 implements the generation of an ontology-driven movie knowledge graph. The knowledge graph generation module 120 generates a semantically rich and machine-interoperable knowledge graph that comprehensively describes a movie in terms of scenes, moods, activities, actors, etc. The knowledge graph generation module 120 can store the knowledge graph as resource description framework (RDF) triples that are queryable using SPARQL (e.g., through a SPARQL query endpoint module 128), which can be linked to external resources. A semantic model or ontology can represent the data in the knowledge graph, facilitating various features, properties, and resources in the knowledge graph that can be used to identify movie scenes that are useful for movie-related tasks, such as movie summarization.

[0100] The video summarization portion 104 includes a summary user interface module 130, a summary application programming interface (API) module 132, and a summary main component module 136. The summary main component module 136 can include modules for scene template selection and general purpose selection by a user to customize preferred features of a movie summary, as well as a scene template processor and a scene ranker for scoring movie scenes according to user selections. When a user requests a movie summary through the interface module 130, the summary API module 132 communicates with the summary main component module 136 to select a subset of the highest ranked scenes according to the user-selected template for inclusion in the movie summary 138. The movie summary 138 can be generated by the module 136 to include only scenes that are highly important to convey the main narrative of the movie, extracted from the ranked movie data from the movie knowledge graph module 120 and / or external movie metadata sources through the module 128.

[0101] Reference Figure 1C FIG. 1OC shows an illustrative block diagram of an illustrative system 100C. The system 100C can include any or all of the features of the system 100B. The system 100C can be used to perform the method 200C. Figure 1AAny or all of the steps of the illustrated method 100A. The system 100C can include one or more modules for performing one, any or all of the steps of the method 100A. The system 100C is computer 141 based. The computer 141 has a processor 143 that controls the operation of the computer 141 and its related components. The computer 141 includes a RAM 145, a ROM 147, the input / output module 149, and a storage 155. The processor 143 executes software routines on the computer 141, such as an operating system 157 and software including the steps of the process 100A. Other components typically found in a computer, such as an EEPROM or flash memory or any other suitable component, can also be part of the computer 141.

[0102] The storage 155 includes any suitable storage technology, such as a hard disk. The storage 155 stores software including the operating system 157 and applications 159 and data 151 required for the operation of the system 100C. For example, the storage 155 can also store video, text, and / or audio ancillary files including multimedia objects. The video, text, and / or audio ancillary files can also be stored in a cache memory or any other suitable storage. Alternatively, for example, some or all of the computer executable instructions including those of the process 100A can be embodied in hardware or firmware (not shown). The computer 141 executes the instructions embodied by the software to perform various functions, such as the steps of the process 100A.

[0103] The input / output (I / O) module 149 can include connections to a microphone, a keyboard, a touchscreen, a mouse, and / or a stylus through which a user of the computer 141 can input. The input can be through cursor movement. The input can include in a transfer event and / or an escape event. The input / output module 149 can also include a speaker for providing audio output and / or one or more video display devices for providing text, audio, audiovisual, and / or graphical output. The input and output can be related to computer application functionality, such as facilitating one or more steps of the process 100A.

[0104] Figure 1CThe network connections depicted include a local area network (LAN) 153 and a wide area network (WAN) 169, but can also include other networks. For example, system 100C is connected to a LAN through LAN interface 153. System 100C can operate in a networked environment using logical connections to one or more remote computers, such as a system 181 and a system 191. The systems 181 and 191 can be a personal computer or a server, including many or all of the elements described above with respect to system 100C. When used in a LAN networking environment, computer 141 is connected to the LAN 153 through a network interface adapter 151. When used in a WAN networking environment, computer 141 can include a modem 151 or other means for establishing communications over the WAN 169, such as the Internet 171.

[0105] It will be appreciated that the network connections shown are illustrative and other means of establishing a communications link between the computers 141, 181, and 191 can be used. It will be appreciated that various public switched

[0106] Further, the application(s) 159 used by the computer 141 can include computer- executable instructions which, when executed, implement steps of the process 100A.

[0107] The computer 141 and / or the systems 181 and 191 can also include various other components, such as a battery, a speaker, and an antenna (not shown).

[0108] Systems 181 and 191 can be portable devices, such as a laptop, a smartphone, or any other suitable device for storing, transmitting, and / or communicating relevant information. Systems 181 and 191 can include other devices. These devices can be the same as or different from the devices of system 100C. These differences can relate to hardware components and / or software components.

[0109] The method can include and the system can involve one or more steps of an emotion-based knowledge graph generation process. Referring to Figure 2 , an example emotion-based knowledge graph generation process 200 is shown. The method can include and the system can involve one or more steps of process 200. Process 200 can be performed by one or more modules of system 100B. Process 200 can be performed by one or more modules of semantic lift portion 102, as shown. Figure 1B Process 200 can begin at step 202.

[0110] In step 202 of process 200, a user can select a multimedia object, such as a movie and / or a movie file, stored on a DVD or any suitable storage medium to generate a knowledge graph. The movie can be of any suitable genre, such as an action movie and / or a thriller.

[0111] In step 204, video pre-processing can include extracting video, audio, and / or subtitles from the movie. Video pre-processing can provide multimedia object data. Multimedia object data can include any suitable data, such as one or more chapter text files detailing the start times of chapters of the multimedia object; multiple audio files for the movie, such as one audio file for one available soundtrack; a relatively large number of smaller audio files that can include segments of actor voices (dialogue) and / or segments of non-speech portions. Step 204 can be performed by automatic media processing module 106 and / or NLP module 108 of semantic lift portion 102, as shown. Figure 1B Step 204 can include natural language processing of the extracted movie data.

[0112] In step 206, movie metadata can be automatically scrubbed from the loaded movie without pre-processing in step 204. Alternatively or additionally, movie metadata can be scrubbed from an external source, such as external source 212. Step 206 can be performed by automatic metadata scrubbing module 116 of semantic lift portion 102, as shown. Figure 1B

[0113] Step 208 shows an implementation of an algorithm and annotation toolkit for annotating the movie according to emotions. Step 208 can be performed by automatic annotation tool module 110 of semantic lift portion 102, as shown. As described above with respect to Figure 1B Figure 1A ​​Step 208 can include one or more user interactions with a specialized graphical user interface. Step 208 can include one or more steps of data analysis. Step 208 can facilitate a user performing any or all of the steps of process 200.

[0114] In step 210, a semantic enrichment module, such as module 114 shown in FIG. 1, can be used to represent a movie in semantics and support analysis of the movie. Implementing an abstract semantic model can include a scene ontology. The scene ontology can include a media annotation ontology. The semantic representation can facilitate navigation between data sources through a resource description framework (RDF) association. Step 210 includes semantic enrichment of data and facilitates semantic representation of movie content. Step 210 can include semantic enrichment of raw or pre-processed movie data. Figure 1B

[0115] Step 212 of process 200 includes interlinking with external data sources of movie metadata. Interlinking with external sources enables the knowledge graph to have semantically richer external information that is not available from the movie content or movie metadata. The external sources of movie metadata can be scraped using movie database application programming interfaces (e.g., IMDB, DBPEDIA, WIKIDATA, and / or other open source data). Step 212 can be mediated by one or more modules of semantic enrichment section 102 shown in FIG. 1, such as automatic metadata scraping module 116 and / or interlinking identification module 118. Figure 1B

[0116] Step 214 depicts generation of an ontology-driven movie knowledge graph. In step 214, a user can follow semantic web and associated data best practices. The user can create novel ontologies and / or reuse one or more existing ontologies. With this architecture, the generated knowledge graph can have rich semantics and machine interoperability and can comprehensively describe the content of a movie, such as scenes and activities within scenes, as well as metadata, such as descriptions of movies, actors. Step 214 can be mediated by one or more modules of semantic enrichment section 102 shown in FIG. 1, such as knowledge graph generation module 120. Figure 1B

[0117] Data analysis can include one or more processes that resolve a multimedia object into action-based and / or emotion-based components. Figure 3 is a graphical depiction of an exemplary emotion tagging process 300. Process 300 can begin at step 302. Data analysis of a multimedia object, such as step 101 of FIG. 1, can include any or all of the steps of process 300.

[0118] ​​​Step 302 includes analyzing a movie having a start time T0 and an end time Tn. As shown in step 302, upon determining (e.g., through data analysis), the movie or other multimedia object can be determined to include a dynamic complex structure. In process 300, the entire movie can be marked as starting at time T0 and ending at time Tn.

[0119] In step 304, it is determined through data analysis that the structure includes one or more component scenes (Sm). Step 304 may include labeling the scenes of the multimedia object. The scenes may have fixed start times and end times. Step 304 may include dividing the movie into component scenes. For example, in step 101 of Figure 1, the data analysis may include dividing the multimedia object into component scenes. The method may include labeling one or more scenes in the component scenes. Labeling may be based on one or more fixed emotions detected in one or more scenes.

[0120] like Figure 3 As shown, it can be determined through data analysis that the first scene (S0) starts at time T0 and the last scene (S m ) can be at time T n End. Intermediate Scene S Tn-50 and S Tn-10 It can be determined through data analysis that at time T n-50 and T n-10 Ending. A scene can follow an overall narrative. A scene transition can be obvious or subtle. A shift in narrative can herald a new scene. A scene transition can be marked, for example, by a change in location, tone, mood, event, narrative stage, and / or major narrative development.

[0121] In step 306, it can be determined through data analysis that each scene includes one or more time-fixed, observable actions and / or activities, such as action A w 、A x 、A y and A z Actions may overlap with each other (not shown).

[0122] In step 308, data analysis can be used to determine that each scene includes one or more temporally fixed, observable, and potentially overlapping emotions. A scene's emotions can differ from the overall genre of the multimedia object. A scene's emotions can set the tone for a particular scene. Table 1 below shows illustrative scene emotions.

[0123] Labeling can be based on one or more emotion indicators, such as auditory indicators of the pitch of the soundtrack in a given scene, and / or audiovisual behavioral indicators of the characters' actions and / or facial expressions. For example, the loudness and / or intensity of the portion of the soundtrack that overlaps the scene can indicate anger as a scene emotion. Soft music can be gentle, melancholic, and / or fearful. High pitches can indicate joy and / or excitement. Low pitches can indicate sadness and / or seriousness. Auditory and / or visual signals of laughter can indicate a joyful atmosphere in the scene, while crying and frowning facial expressions can indicate sadness. Table 1 below shows illustrative emotion indicators.

[0124] Table 1. Illustrative emotion indicators. Music intensity and timbre can represent relative values. Pitch of music is expressed in approximately hertz, and tempo is expressed in beats per minute.

[0125]

[0126] Alternatively and / or additionally, the marking may be performed based on explicit and / or implicit cues of emotions included in the subtitles of the multimedia object.Alternatively and / or additionally, the marking may be performed based on scraped external data of the media database.

[0127] The method may include creating semantic annotations using a semantic annotation architecture. For example, to generate movie metadata, a data resource such as a MovieDB API may be used to scrape the movie metadata and promote it into a knowledge graph. A media annotation ontology (prefixed with "ma" in the figure below) may capture the required information. The knowledge graph may include other predicates such as "hasDirector". The knowledge graph may include a set of categories to represent different professions in the film industry, such as director. These concepts may be formally defined as part of a multimedia scene ontology. The method may include generating one or more features of the scene ontology. The knowledge graph may include a scene ontology that may be Figure 1A The process 100A is shown as created during step 103 and / or may be Figure 1B One or more modules (eg, modules 112, 118, and 120) of the semantic enhancement portion 102 of the illustrated system 100B generate. Figure 4 , shows an illustrative scene annotation ontology 400; refer to Figure 5A , showing an exemplary diagram 501 of the resource description framework architecture provided by the present invention; Figure 5B ,like Figure 5A As shown in diagram 501 , illustrative emotional moment definitions 503 are shown.

[0128] The ontology 400 semantically defines concepts that capture moments in a movie to translate raw multimedia information of a video into a meaningful and semantically rich representation. The ontology 400 can be designed to be machine interoperable and also human understandable.

[0129] As shown in Figure 4 , according to the legend 501 shown in Figure 5A , in the ontology 400, the scene 402 is determined to be an existing Media Annotation Ontology (MA) media segment 404 that corresponds to an existing MA media resource 406 during the time interval 408. The media resource 406 is based on an existing RDFS resource 410. The scene 402 is associated with an existing location 412 through the MA "ma:hasRelatedLocation". For example, the scene 402 is determined to include a moment 414 of the time interval 416 through data analysis. The moment 414 is determined and labeled to include an emotional moment 418 and / or an observable action 420. According to the emotional moment definition 503 shown in Figure 5B , the emotional moment 418 can be determined and labeled as a property with an associated emotion 521 according to the legend 501 shown in Figure 5A . The moment 414 can be scored according to importance, for example, using an Extensible Markup Language (XML) Schema Definition (XSD) of an XML Boolean score 422. The score 422 can be according to the emotion 521 and / or the action 420 and / or other relevant cues for the scene to convey the importance of the main narrative of the movie, for example, the actor 424 in the scene. Figure 5B

[0130] The creation of the scene ontology of the ontology 400 as shown in Figure 4 may be implemented through one or more algorithmic processes. Reference is made to Figure 6 , Figure 6 is a flowchart of an exemplary process and / or algorithm 600. The algorithm 600 can be used to annotate a movie. The algorithm 600 can start from step 601. Step 601 can include starting the algorithm 600. The algorithm 600 can include any and / or all steps of the processes 100A, 200, and 300 described with reference to Figure 1A , Figure 2 and Figure 3 respectively. The processes 100A, 200, and 300 can include any and / or all steps of the algorithm 600. The algorithm 600 can be performed by one or more modules of the portion 102 of the system 100B shown in Figure 1B . ​

[0131] In step 602, a specific movie may be selected and / or loaded by the user. A stored annotation dataset may also be loaded to annotate the movie. Metadata for the movie may also be loaded. Step 602 may include information about Figure 2 Any and / or all features described in step 202 of process 200 are shown.

[0132] The user may add and / or edit annotations for the movie in step 604. The annotations may be saved in an annotation database.

[0133] In step 606, the algorithm and / or the user may populate the media ontology. The ontology may include information about Figure 4 Any and / or all features described for the illustrated body 400 .

[0134] In step 608, the user and / or algorithm may derive a movie knowledge graph. The knowledge graph may be stored as RDF triples. The knowledge graph may be derived by Figure 1B The module 120 of the system 100B shown is executed. Step 608 may include Figure 1A Step 103 of process 100A shown and / or Figure 2 Any and / or all features of steps 212 and / or 214 of process 22 are shown. Step 103 of process 100A may include any and / or all features of step 608.

[0135] The steps of the algorithm process can be implemented by the user through one or more user interfaces. The user interface may include information about Figure 2 Any or all features of the GUI described in step 101 of process 100A and / or step 208 of process 200 are shown. Figure 1B The illustrated module 110 of the system 100B may include one or more features of a user interface.

[0136] refer to Figure 7 , an exemplary graphical user interface (GUI) 700 for an exemplary annotation toolkit for facilitating annotation based on emotions is shown. The user interface may include any or all of the features of the GUI 700. The GUI 700 may facilitate the process of performing the process with respect to 100A, Figure 2 200 and Figure 3 The annotation toolkit may provide annotators with options to record information about activities and emotions observed in a movie. The GUI 700 includes four main features.

[0137] Features 702 include illustrative global actions, moods, and entities for a loaded movie. A user can select predefined actions, moods, and / or entities in features 702. A user can define new actions, moods, and / or entities in features 702. Saved actions, moods, and / or entities can be available in the movie.

[0138] Features 704 are used to display the movie. Features 704 can include a device for viewing the movie; a device for listening to the soundtrack; a device for pausing the movie; a device for skipping scenes; and / or any other suitable movie display device. Features 704 can facilitate a user to detect moods and / or actions in scenes viewed within a specified time frame.

[0139] Features 706 are used to display and / or edit annotation details according to expressed aspects, such as scenes, actions, and / or moods. Various configuration options are provided in features 706 according to a user’s selection of annotation requirements in features 708. Each annotation aspect includes corresponding options. For example, if a user wants to define a new action, the provided options include determining a start point and an end point of a timeline editor. In addition, entities involved in the action, such as a receiving entity and a performing entity, are also assigned to a specific time frame.

[0140] GUI 700 is used to facilitate semantic modeling of movie content. Annotators make selections from performing entities, participating activities, and / or receiving entities of associated actions defined in features 702. Control buttons can be included to facilitate access to subsequent and / or previous options, to delete annotations, and / or other suitable operations. The operations can be implemented according to one or more selected objects and / or regions that have been selected in features 708. Features 708 are used to facilitate switching between scenes, actions, and / or moods.

[0141] A panel of features 708 is used to create and / or define scenes, to determine mood temporal states, and to annotate activities that occur within a specified time. As shown, a horizontal side (from left to right) of the panel represents a timeline. A vertical face is divided into three sections: “Scene,” “Action,” “Mood.”

[0142] A user interface, such as GUI 700, can be included in an annotation tool. The tool can include software and / or hardware for performing media annotation of multimedia objects. Referring to Figure 8 , an illustrative system architecture 800 of an annotation tool is shown. The annotation tool can include the previously described annotation tool kit described with respect to Figure 7 . Architecture 800 can include one or more features of system 100B shown in Figure 1B . Architecture 800 can be used to perform the respective Figure 1Aone or more steps of the process 100A, Figure 2 the process 200, Figure 3 the process 300 and Figure 6 the process 600. The architecture 800 can be used to generate Figure 7 the ontology 700.

[0143] In the system architecture 800, the scrub metadata 802 of a movie is presented through a user interface 804. The user interface 804 can include one or more features of the GUI 700 of Figure 7 The user interface 804 interacts with the annotation module 806 of the movie in the timeline editor 808. The data of the annotation tool is then stored as RDF triples in the knowledge graph 810.

[0144] The methods and systems of the present invention can include semantic analysis. The methods and systems of the present invention can include text mining. The methods and systems of the present invention can include deep learning. The knowledge graph data can be indexed according to searches, returning concepts closely related to the search terms. For example, “punching” and “kicking” can be returned for a more general “fighting” search. Semantic analysis can avoid semantic gaps in the generated activities if the activities are just string labels, so that the machine cannot understand synonyms. Semantic analysis can also facilitate searching for general concepts, for example in the summarization task.

[0145] Semantic analysis can include building a Simple Knowledge Organization System (SKOS) taxonomy so that searches can expand from a broad range of terms / concepts (e.g., fighting) to more detailed terms (e.g., punching). Using SKOS can enrich the movie activities with richer semantic meaning and context. The taxonomy can be created by pre-processing action movie scripts to create a large vocabulary.

[0146] A skip-gram word2vec model can be built with a corpus. The skip-gram model can take each word in a movie script and surrounding words in a pre-defined window to construct word pairs. The word pairs can be used as training data for a single layer neural network. The trained neural network can then be used to predict the probability of related words.

[0147] Referring to Figure 9 , an example 900 of a single layer neural network input output is depicted, referring to Figure 10, depicts a visualization 1000 of a portion of an illustrative SKOS taxonomy. Example 900 includes exemplary inputs and outputs of a trained neural network. Example 900 illustrates the calculated probability of the more specific word "punch" appearing near the broader word "fight" based on the text of an illustrative script including the broader action-oriented word "fight" and the more specific action-oriented word "punch." Based on a threshold of a higher probability of the more specific word appearing near the broader word, a user can incorporate more specific concepts associated with the more specific word into the broader category associated with the broader word, labeling scenes with the more specific concept with the broader label of the broader word.

[0148] Figure 10 The structure 1000 shown in Figure 1 shows the concept "skos:narrower" hierarchy. To establish an activity taxonomy, abstract terms can be created in the knowledge graph for manually curated activities. The abstract terms can be used as broad SKOS concepts, while the manual activities can be defined as skos:narrower concepts.

[0149] The method may include generating one or more comprehensive metadata ontologies that interconnect action and / or emotion metadata with the context of the multimedia object.The knowledge graph may include the metadata ontologies.

[0150] Next, refer to Figure 11 、 Figure 12A and Figure 12B . Figure 11 Depicted is an ontology 1100 comprising illustrative segments of multimedia metadata generated in accordance with the principles of the present invention. Figure 12A The main body 1200 shows Figure 12A The scene 1207 shown in Figure 12B 1201) can be converted into a triple of type "so:ObservableAction".

[0151] In ontology 1100, a movie is represented by the identifier "movie123456," as shown in box 1102. In box 1103, a movie is defined as a type of "MediaResource," a concept in the media annotation ontology. Movie metadata is enriched with several movie properties, including: the creation of the movie via the "ma:createdIn" predicate 1105; the director of the movie using the "so:hasDirector" predicate 1107; and the actors using the "ma:features" predicate 1109.

[0152] Similarly, other instances have multiple properties associated with them, including a human-readable label, as shown by dashed box 1104 associated with the "rdfs:label" predicate. Each instance is categorized, with dashed boxes representing types in an external ontology, such as box 1103 representing "ma:MediaResource." Other instances are shown to refer to categories defined in the scenario ontology, such as box 1110 representing "so:FictionalEntity." Furthermore, each instance can be enriched with one or more associations to external resources, such as those shown by dashed box 1108.

[0153] exist Figure 12A In the ontology 1200 shown in FIG, it is shown that movie 1206 includes scene 1207 associated with movie 1204 via the “ma:isFragmentOf” predicate 1202. Scene 1207 includes one or more moments, e.g. Figure 12A The "so:ObservableAction" moment 1204 in the ObservableAction or, for example, "so:MoodMoment" (not shown). The scene and moment identifiers are annotated with timestamps. The timestamps conform to the W3C recommendations for media resource URI specifications. Timestamps are distinguished by the URI fragment symbol "#" and the parameter "t=ts,te", where ts represents the start time of the media segment and te represents the end time. Figure 12A As shown, scene 1207 is marked with a timestamp “scene#t=60,316”, indicating that scene 1207 starts at the 60th second and ends at the 316th second.

[0154] like Figure 12A As shown, the observable action corresponding to the time 1204 included in the scene 1207, such as action 1209, starts at 195 seconds and ends at 209 seconds. For example, the observable action may include a person 1213 walking on a plane wing 1215, as shown in FIG. Figure 12B As shown. Instances of observable actions are associated with scenario instances through the "so:hasMoment" predicate (e.g., predicate 1208). Each "ObservableAction" defines an activity, such as activity 1210, or a "so:performingEntity", such as so:performingEntity 1203. Or a "so:receivingEntity", such as so:receivingEntity 1205, or both (e.g., Figure 12A ). The "ObservableAction" and "Scene" instances are enriched by a set of time triplets. For simplicity of description, Figure 12B These time triplets are omitted in .

[0155] The apparatus can involve and the method can include querying a movie knowledge graph. A semantic enrichment module can be used to represent video content and support analysis of multimedia objects in a semantic manner. The semantic enrichment module can use emotions and / or related activities as core components to enrich movies and create a knowledge graph.

[0156] As described above, an abstract semantic model can be designed according to a scene ontology, which can include one or more aspects of a media annotation ontology. The resulting data generated by the annotation tool can follow the same ontology / architecture. As a result, a semantic knowledge graph can be generated, enabling complex reasoning and querying capabilities, and supporting novel and personalized customer-oriented content representation.

[0157] As disclosed, the annotation process according to emotions can help create a knowledge graph. In some embodiments of the invention, the knowledge graph in turn can help create a shorter video summary of a movie, for example, a summary that includes no more than 25% of the original length of the movie. However, the summary can preserve important narrative parts of the movie.

[0158] In some embodiments of the invention, annotation of a movie can include manual and / or automatic processing and semantic enrichment of the movie. The automatic part of the enrichment can include natural language processing (NLP) techniques, for example, audio narration of the movie.

[0159] Other systems, methods, features, and advantages of the present invention will be or become apparent to one with skill in the art upon examination of the following drawings and detailed description. It is intended that all such additional systems, methods, features, and advantages be included within this description, be within the scope of the present invention, and be protected by the accompanying claims.

[0160] The descriptions of the various embodiments of the present invention have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0161] It is expected that numerous related processes of multimedia object knowledge graph generation and summarization will be developed during the life of this application, and that the scope of the terms multimedia object knowledge graph generation and summarization is intended to include all such new technologies a priori.

[0162] The term "about" as used herein means ± 10%.

[0163] The terms "comprising," "having," "including," and "containing" are to be construed open-ended terms (i.e., meaning "including, but not limited to,") unless otherwise noted.

[0164] The phrase "consisting essentially of" means that the composition or method can include additional ingredients and / or steps, but only if the additional ingredients and / or steps do not materially alter the basic and novel characteristics of the claimed composition or method.

[0165] As used herein, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a compound" or "at least one compound" can include more than one compound, including mixtures thereof.

[0166] As used herein, the term "exemplary" means "serving as an example, instance, or illustration." Any implementation of the disclosure described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.

[0167] As used herein, the term "optionally" means "in some embodiments provided, and in other embodiments not provided." Any particular embodiment of the application can comprise a plurality of "optional" features unless such features contradict each other.

[0168] In this application, various embodiments of the application can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and is to be interpreted -in the context of the specification as a whole. Therefore, this description of a range should be considered as specifically and explicitly disclosed. Thus, for example, a description of a range such as e.g. from 1 to 6 should be considered to have specifically disclosed sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, e.g. 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0169] When a range of numbers is indicated herein, any number (fraction or integer) within the indicated range is included. The phrases "range between X and Y" and "ranging from X to Y" are interchangeable and are used to describe all numbers within the indicated range and all fractions within the range. For example, a range of 1 to 6 is intended to include all numbers between 1 and 6, as well as fractions of the number, e.g., 1.1, 1.2, 1.3, 1.4, 1.5, etc.

[0170] It should be appreciated that for simplicity and clarity of illustration, elements shown in the drawings have not necessarily been drawn to scale. For example, the dimensions of some of the elements are exaggerated relative to the drawings. Further, where considered appropriate, reference numerals have been repeated among the drawings to indicate corresponding or analogous elements. It should be appreciated that the use of any of the

[0171] Herein, all publications, patents and patent documents referred to in this application are incorporated by reference herein in their entirety, as are individual elements of the publications, patents and patent documents, as specifically and individually set forth herein. In the event of inconsistencies between the disclosure of the present application and the disclosure of the above- incorporated by reference publications, patents and patent documents, the disclosure of the present application shall dominate. To the extent that any meaning or definition of a term in any of the documents incorporated by reference herein conflicts with the meaning or definition of that term provided herein, the meaning or definition provided herein shall prevail. Headings of sections provided in this patent application and the summary above are for convenience only and shall have no interpretations therefrom whatsoever.

Claims

1. A system for summarizing a multimedia object having a primary narrative, characterized by, Comprising: a processor executing readable instructions to: perform data analysis to identify in each of a plurality of scenes of the multimedia object one or more scene-related emotions indicated in the scene; generate a knowledge graph associating each of the plurality of scenes with a respective one or more scene-related emotions; compute a plurality of scores using the knowledge graph, each score indicating a relative importance of one of the plurality of scenes to conveying the primary narrative; select a subset of the plurality of scenes according to the plurality of scores; generate a summary of the multimedia object according to the subset; wherein the computing a plurality of scores using the knowledge graph comprises: computing a score for a time instance of a scene ontology in the knowledge graph according to an emotion and / or an action and / or other cues to the importance of the time instance to conveying the primary narrative using an extensible markup language, XML, schema definition, XSD, of an XML Boolean score.

2. The system of claim 1, wherein, The data analysis comprises pre-processing the multimedia object, the pre-processing comprising extracting from the multimedia object at least one of: a video file; a subtitle file; a chapter text file detailing start times of chapters of the multimedia object; an audio file segment of actor speech and non-speech portions.

3. The system of claim 1, wherein, The data analysis comprises: erasing associated metadata describing the multimedia object; analyzing the associated metadata to indicate the one or more scene-related emotions.

4. The system of claim 1, wherein, The data analysis comprises implementing semantic lifting according to a scene ontology to capture raw multimedia information of the multimedia object.

5. The system of claim 1, wherein, The data analysis comprises interlinking with external sources describing features of the multimedia object.

6. The system of claim 5, wherein, The features comprise at least one of: a scene of the multimedia object; an activity in a scene of the multimedia object; an actor performing in the multimedia object; a character depicted in the multimedia object.

7. The system of claim 1, wherein, The data analysis comprises analyzing a descriptive audio soundtrack of the multimedia object to indicate the one or more scene-related emotions.

8. The system of claim 1, wherein, The data analysis comprises extracting the emotions from visual emotion indicators comprising at least one of: facial expression images; body posture images; video sequences of emotion-indicative behavior.

9. The system of claim 1, wherein, The data analysis comprises extracting the emotions from auditory emotion indicators comprising at least one of: emotional indicators of a musical soundtrack; emotionally suggestive vocalization indicators.

10. The system of claim 1, wherein, The data analysis comprises extracting the emotions from textual emotion indicators comprising at least one of: explicit emotional descriptors; suggestionally emotional indicators.

11. A method for summarizing a multimedia object having a primary narrative, characterized by, Comprising: performing data analysis to identify in each of a plurality of scenes of the multimedia object one or more scene-related emotions indicated in the scene; generating a knowledge graph associating each of the plurality of scenes with a respective one or more scene-related emotions; computing a plurality of scores using the knowledge graph, each score indicating a relative importance of one of the plurality of scenes to conveying the primary narrative; selecting a subset of the plurality of scenes according to the plurality of scores; generating an abstract of the multimedia object from the subset; wherein the calculating the plurality of scores using the knowledge graph comprises: using an extensible markup language, XML, schema definition, XSD, of an XML Boolean score, calculating a score for a time instance of a scene ontology in the knowledge graph based on an emotion and / or an action and / or a scene corresponding to the time instance and other cues for the importance of conveying the main narrative.

12. The method of claim 11, wherein, The data analysis comprises pre-processing the multimedia object, the pre-processing comprising extracting from the multimedia object at least one of: a video file; a subtitle file; a chapter text file detailing start times of chapters of the multimedia object; an audio file segment of actor speech and non-speech portions.

13. The method of claim 11, wherein, The data analysis comprises: erasing associated metadata describing the multimedia object; analyzing the associated metadata to indicate the one or more scene-related emotions.

14. The method of claim 11, wherein, The data analysis comprises implementing semantic lifting from a scene ontology to capture raw multimedia information of the multimedia object.

15. The method of claim 11, wherein, The data analysis comprises interlinking with external sources describing features of the multimedia object.

16. The method of claim 15, wherein, The features comprise at least one of: a scene of the multimedia object; an activity in a scene of the multimedia object; an actor performing in the multimedia object; a character depicted in the multimedia object.

17. The method of claim 11, wherein, The data analysis comprises analyzing a descriptive audio soundtrack of the multimedia object to indicate the one or more scene-related emotions.

18. The method of claim 11, wherein, The data analysis comprises extracting the emotion from visual emotion indicators, the visual emotion indicators comprising at least one of: facial expression images; body posture images; video sequences of emotion-indicative behavior.

19. The method of claim 11, wherein, The data analysis comprises extracting the emotion from auditory emotion indicators, the auditory emotion indicators comprising at least one of: emotional indicators of a musical soundtrack; emotionally suggestive vocalization indicators.

20. The method of claim 11, wherein, The data analysis comprises extracting the emotion from textual emotion indicators, the textual emotion indicators comprising at least one of: explicit emotional descriptors; suggestionally emotional indicators.

Citation Information

Patent Citations

  • System And Method For Determining Sentiment Expressed In Documents

    CN102812475A

  • A method for automatically creating multi-level event and scene map characteristics, a device and an application thereof

    CN109255385A