Method for the automated creation of a content summary, system for carrying out the method, and vehicle with the system
The method and system address the challenge of incomplete media consumption by automatically generating summaries from user-defined positions, leveraging AI and neural networks for efficient media content summarization.
Patent Information
- Application Number
- DE102024124758
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2026-03-05
AI Technical Summary
Existing media consumption methods fail to efficiently summarize and make accessible unconsumed parts of media content consumed in segments over a longer period, leading to potential memory gaps and incomplete context perception.
A method and system utilizing artificial intelligence and neural networks to automatically create content summaries from definable positions within media content, enabling output and storage of these summaries via dedicated units.
Enables seamless access to unconsumed media content parts, filling memory gaps by providing automated summaries that enhance user experience and context retention.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] Media content is sometimes consumed not in its entirety, but only in excerpts. For example, in a video, one might skip to an interesting part and then watch essentially that section. Similarly, with longer texts or audio content, such as audiobooks, it can happen that this media content is not consumed completely or in one go, but rather read or listened to in sections.
[0002] As with traditional books, electronic books, or e-books, can also be read intermittently over a longer period, rather than in one sitting. This can also be the case with video or audio content. For example, an audiobook might be played in a car, so that on each journey, even those long apart, only a portion of the audiobook is listened to. Even with shorter texts, videos, or audio content, changing consumption habits can lead to only brief excerpts being viewed, meaning the context or the rest of the content is not actively perceived.
[0003] The invention aims to improve the consumption of media content by making unconsumed parts of the content accessible to the user. Similarly, for media content that is consumed in segments over a longer period, potential memory gaps can be easily filled by making previous or future parts of the content available.
[0004] To achieve the invention, a method, a system, and a vehicle are provided according to the independent claims. Further developments of the invention are the subject of the dependent claims.
[0005] In a first aspect, a method for the automated creation of a content summary of media content is provided, wherein at least one position for the media content is definable, and wherein the at least one position defines a point up to and / or from which the automated content summary is to be created, wherein the content of the media content is automatically summarized up to and / or from the at least one position and is preferably output and / or stored in a storage unit in a retrievable manner.
[0006] At least one definable position can be defined by the playback progress of the media content.
[0007] At least one position can be defined by a user interaction at an input unit.
[0008] Multiple positions can be defined, with the multiple positions defining sections of the media content, and the content of the defined sections being automatically summarized.
[0009] The automated content summarization can preferably be carried out using an application of artificial intelligence and / or an artificial neural network via a computing unit.
[0010] The media content and the content summary can be output via one output unit. The media content and the content summary can be output via different output units.
[0011] The media content can be at least video content, audio content and / or text.
[0012] The media content can first be transcribed into text. A preliminary summary can then be automatically generated from this text. Relevant sections of the media content can be identified and / or selected for this preliminary summary and output as an automated summary of the media content.
[0013] In a further aspect, a system comprising at least one output unit, one input unit, one storage unit, and one computing unit is provided, wherein the system is configured to automatically create a content summary of media content depending on the position, wherein the media content is at least partially output on the output unit, and wherein at least one position for the media content can be defined by means of the input unit, which defines a point in the media content up to and / or from which the automated content summary is to take place, wherein the computing unit is configured to automatically create the content of the media content up to and / or from the at least one position and preferably output it on the at least one output unit and / or another output unit and / or store it in the storage unit in a retrievable manner.
[0014] In yet another aspect, a vehicle is provided with a system as described herein.
[0015] The invention is now also described with regard to the figures. They show: Fig. 1 schematically a representation of a media playback application according to the invention. Fig. 2 a further media playback device which is configured to carry out the method according to the invention.
[0016] The invention makes it possible, for example, to consume media content and to create a content summary from a certain position, for example from the current playback position.
[0017] For example, video content can be summarized from the current playback position to the end and / or to the beginning. This is also possible for text and audio content. For audio and video content, the section to be summarized can first be transcribed into a textual representation and then summarized using a text summarization application, such as an AI model. When summarizing audio and video content, it is important that the automatic summarization incorporates position markers, such as timestamps, into the summary so that the relevant parts of the media content—for example, the sections of a video or audio track—can be identified.One or more compilations of the audio and / or video content can then be automatically created and presented to the user as a video or audio summary.
[0018] In the simplest case, a summary is provided in text form. This may first require generating an automated transcription of the media content. Specifically, it is possible to convert the transcribed text, or the summary of the transcribed text, into an audio output using audio synthesis. For example, it is possible to automatically create a summary of previously listened-to chapters of an audiobook while driving, meaning a summary up to a user-defined playback point or position, such as the beginning of the audiobook. The process can then transcribe the audio content into text.The transcription is then summarized by the process, and a summary is output, for example, after the generation of speech output using speech synthesis, so that the vehicle user can first listen to a summary of what has happened so far in the vehicle and then continue with the playback of the audiobook.
[0019] The same method can also be applied, for example, to the playback of e-books. When playing an e-book, for instance on an e-reader, the user can define a point up to which the content of the e-book should be summarized. Similarly, if there isn't enough time to finish the book, the user can specify that a summary should be created starting from the playback point or position defined by the user.
[0020] It is important that the user can define one or more positions for the media content, and thus define sections that should be automatically grouped. For example, multiple sections of audio, video, and / or text content can be defined using an input device to be grouped together. In a media player application, these multiple sections can be defined, for example, using or on a media playback bar that displays the playback progress. The user can also define parts of the media content, such as sections that are less interesting to them, that should be grouped together. For example, pages of an eBook that the user would skip over in a traditional book can be summarized.
[0021] As in Fig. As shown in Figure 1, in a media playback application (MWA) executable on a system according to the invention, a first position P1 can be defined, for example, on a media playback bar (MWA). A content summary of the area of media content MI1 prior to the first position P1, which is located to the left of the first position P1 in the figure shown, can then be automatically generated using a button B. Alternatively or additionally, a second position P2 can also be defined by the user on the media playback bar (MWA). A content summary of the area of media content MI1 prior to the second position P2, specifically the area located to the left of the second position P2 in the figure, can then be automatically generated using button B.
[0022] Of course, multiple buttons can be provided, or the configuration of button B can be chosen differently. For example, it is possible that when button B1 is pressed, the media content is created starting from the first position P1 or the second position P2, as specified in the... Fig. The first position (P1) and the second position (P2) are located to the right of the first position (P1) and the second position (P2), respectively. Specifically, the first position (P1) and the second position (P2) represent the current playback position of the media content (MI1). It is also possible to define or select both the first position (P1) and the second position (P2) simultaneously on the media playback bar (MWL). Using a button (B2), the area between the first position (P1) and the second position (P2) can then be automatically summarized, creating a content summary of that area.
[0023] Fig. Figure 2 schematically shows a media playback device MWG as a system for playing back a second media content MI2, in particular an electronic text. Fig. Figure 2 is an example of a current third position P3, which represents the current position up to which a user has consumed the media content MI2, or which the user has freely selected.
[0024] The user is presented with the options available via the input unit of the media playback device (MWG), which allow for the selective summarization of the media content MI2. For example, a first option O1 can be offered by a media playback application of the MWG, which allows the media content MI2 to be summarized from the third position P3 to the beginning of the media content MI2. Alternatively or additionally, a second option O2 can be offered, which allows the media content MI2 to be summarized from the current third position P3 to the end of the media content MI2.
Claims
[1] Method for the automated creation of a content summary of a media content (MI1), wherein at least one position (P1) for the media content (MI1) is definable, and wherein the at least one position (P1) defines a point up to and / or from which the automated content summary is to be created, wherein the content of the media content (MI1) is automatically summarized up to and / or from the at least one position (P1) and is preferably output and / or stored in a storage unit in a retrievable manner. [2] Method according to claim 1, wherein the at least one definable position (P1) is defined by a playback progress of the media content (MI1). [3] Method according to claim 1 or 2, wherein the at least one position (P1) is defined by a user interaction at an input unit. [4] Method according to any of the preceding claims, wherein multiple positions are definable, the multiple positions defining the sections of the media content (MI1), wherein the contents of the defined sections are automatically summarized. [5] Method according to any of the preceding claims, wherein the automated content summarization is preferably carried out by means of an application of artificial intelligence and / or by means of an artificial neural network using a computing unit. [6] Method according to any of the preceding claims, wherein the media content (MI1) and the content summary are provided for output on an output unit. [7] Method according to any of the preceding claims, wherein the media content (MI1) and the content summary are provided for output on different output units. [8] Method according to any of the preceding claims, wherein the media content (MI1) and the content summary is at least video content, audio content and / or text. [9] Method according to one of the preceding claims, wherein first a transcription of the media content (MI1) into a text content is carried out. [10] Method according to claim 9, wherein a preliminary content summary is automatically generated from the text content, and wherein relevant parts from the media content (MI1) are selected for the preliminary content summary and are output as a content summary of the media content (MI1). [11] System comprising at least one output unit, one input unit, one storage unit and one computing unit, wherein the system is configured to automatically create a content summary of a media content (MI1) depending on the position, wherein the media content (MI1) is at least partially output on the output unit and, wherein at least one position (P1) for the media content (MI1) can be defined by means of the input unit, which defines a location of the media content (MI1) up to and / or from which the automated content summary is to be performed, wherein the computing unit is configured to automatically create the content of the media content (MI1) up to and / or from the at least one position (P1) and preferably output it on the at least one output unit and / or another output unit and / or store it in the storage unit in a retrievable manner. [12] Vehicle with a system according to claim 11.
Citation Information
Patent Citations
Systems and methods for content curation in video based communications
US20180359530A1
Preview streaming of video data
US9635307B1
Browser for use in navigating a body of information, with particular application to browsing information represented by audiovisual data
WO1998027497A1