Information processing device, information processing method, and computer program
The technology addresses the inefficiencies of single-modal video summarization by integrating visual and audio data to create a detailed summary video, ensuring no important information is missed, with flexible viewing options for users.
Patent Information
- Application Number
- PCT/JP2025/000086
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-16
- Filing Date
- 2025-01-06
- Publication Date
- 2025-08-21
AI Technical Summary
Conventional video summarization technologies often overlook important information by focusing solely on either visual or audio data, leading to inefficient summarization and potential loss of critical content, especially in videos where both audio and visual information are integral.
An information processing device and method that extracts both visual and audio information from videos, using optical character recognition, object detection, and speech recognition to generate a summary video, adjusting the summarization rate based on the amount of visual and audio content, and incorporating speaker adaptation for seamless transitions.
Generates a comprehensive summary video that includes all important elements, allowing users to efficiently grasp the content without missing key information, with the ability to seamlessly switch between summary and original videos for personalized learning.
Smart Images

Figure JP2025000086_21082025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and computer program
[0001] The technology disclosed in this specification (hereinafter referred to as "the present disclosure") relates to an information processing device, an information processing method, and a computer program for processing moving images.
[0002] The widespread use of video-sharing platforms has led to the publication of a vast number of videos online, increasing opportunities to learn information through them. While the benefits of video-based learning methods have been recognized, it can be difficult to select appropriate videos from a large number of videos and adjust the learning pace to suit individual comprehension and interests. For users with limited time or interests, being able to quickly review only the key points of a long video can help them decide whether the entire video is worth watching and quickly grasp the characteristics of many videos. Furthermore, users can expect to improve their learning efficiency by grasping only the gist of the parts of a video they already fully understand and rewatching or viewing at a slower pace the important parts they have yet to fully understand.
[0003] Video summarization technologies have been widely developed to help users quickly understand the content of videos. Most conventional video summarization technologies focus on either the visual or audio information of a video to extract important parts of the video. However, important information does not necessarily exist in the same place in both the video and the audio. Therefore, if users try to extract important parts by focusing only on either the visual or audio information, they run the risk of missing important information that only appears in the other information. In videos such as lecture videos and news, both audio and visual information are important. For example, in a lecture video, the teacher's remarks, the visual information on the blackboard and slides, and the audio information spoken by the teacher are all important, but the locations where each piece of information appears are not necessarily synchronized.
[0004] For example, a video condensing device has been proposed that includes a speech-to-text processor that performs speech recognition on audio included in a video to obtain text, a sentence summarizing unit that summarizes the text obtained from the text into multiple sentences, and a video condenser that generates a video condense by obtaining time segments corresponding to each of the multiple sentences, extracting video segments corresponding to each time segment from the video, and combining the extracted video segments (see Patent Document 1). This video condenser generates a video condense by cutting the original video, but the cut video often contains redundant or unnecessary parts, resulting in poor summarization efficiency. Furthermore, this video condenser cannot generate a video condense by taking into account visual information included in the video (e.g., visual information from a blackboard or slides in a lecture video).
[0005] Japanese Patent Application Laid-Open No. 2023-5038
[0006] Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever "Robust Speech Recognition via Large-scale Weak Supervision"(arXiv:2212.04356v1 [eess.AS] 6 Dec 2022)GPT-4 Technical Report (arXiv:2303.08774v4 [cs.CL] 19 Dec 2023)
[0007] An object of the present disclosure is to provide an information processing device, an information processing method, and a computer program that perform processing to generate a digest video.
[0008] The present disclosure has been made in consideration of the above-mentioned problems, and a first aspect thereof is an information processing device comprising: a division unit that divides a video into segments; a visual information extraction unit that converts visual information included in the video for each segment into first character information; an audio information extraction unit that converts audio information included in the video for each segment into second character information; a summary sentence generation unit that generates a summary sentence for each segment of the video based on the first character information and the second character information; and a summary video generation unit that generates a summary video of the original video based on the summary sentence generated by the summary sentence generation unit.
[0009] The visual information extraction unit performs optical character recognition on character information included in video frames and detects objects included in the video frames using an object detection model to convert the detected objects into first character information, the audio information extraction unit converts audio information included in the video into second character information using a speech recognition model, and the summary generation unit generates a summary from the first character information and the second character information using a large-scale language model.
[0010] The information processing device according to a first aspect further includes a summary length adjustment unit that adjusts the number of characters output by the summary generation unit based on an amount of information in each segment. The summary length adjustment unit adjusts the number of characters output by the summary generation unit based on at least one of the number of characters of the second text information output by the audio information extraction unit, the number of objects detected from the video frames by the visual information extraction unit, and the number of words recognized as characters from the video frames by the visual information extraction unit.
[0011] The summary video generation unit generates the summary video using videos of segments selected based on the importance of each segment and audio of summaries of the segments. For example, the summary video generation unit selects one or more segments to be used in the summary video based on the length of the synthesized audio.
[0012] A second aspect of the present disclosure is an information processing method having: a division step of dividing a video into segments; a visual information extraction step of converting visual information contained in the video for each segment into first character information; an audio information extraction step of converting audio information contained in the video for each segment into second character information; a summary sentence generation step of generating a summary sentence for each segment of the video based on the first character information and the second character information; and a summary video generation step of generating a summary video of the original video based on the summary sentence generated in the summary sentence generation step.
[0013] Furthermore, a third aspect of the present disclosure is a computer program written in a computer-readable format to cause a computer to function as: a division unit that divides a video into segments; a visual information extraction unit that converts visual information contained in the video for each segment into first character information; an audio information extraction unit that converts audio information contained in the video for each segment into second character information; a summary sentence generation unit that generates a summary sentence for each segment of the video based on the first character information and the second character information; and a summary video generation unit that generates a summary video of the original video based on the summary sentence generated by the summary sentence generation unit.
[0014] A computer program according to a third aspect of the present disclosure defines a computer program written in a computer-readable format to perform predetermined processing on a computer. The computer program can be provided to a computer capable of executing various program codes in a computer-readable format via a storage medium or communication medium, such as an optical disk, a magnetic disk, or a semiconductor memory, or a communication medium such as a network. By installing the computer program according to the third aspect of the present disclosure on a computer via any of these media, a cooperative effect is exerted on the computer, and the same effects as those of the information processing device according to the first aspect of the present disclosure can be obtained.
[0015] Furthermore, a fourth aspect of the present disclosure is an information processing device including: a reception unit that receives an instruction to switch between a summary video generated based on a summary of each segment generated from visual information and audio information extracted for each segment into which the video is divided and the original video; and a presentation unit that presents either the summary video or the original video on a screen based on the switching instruction received by the reception unit.
[0016] The presentation unit further presents a thumbnail and a summary of each segment, and in response to the reception unit receiving an instruction to switch to any of the segments, switches the screen to the video of the segment instructed to switch.
[0017] Furthermore, a fifth aspect of the present disclosure is an information processing method having: a receiving step of receiving an instruction to switch between a summary video generated based on a summary of each segment generated from visual information and audio information extracted for each segment into which the video is divided and the original video; and a presentation step of presenting either the summary video or the original video on a screen based on the switching instruction received in the receiving step.
[0018] Furthermore, a sixth aspect of the present disclosure is a computer program written in a computer-readable format to cause a computer to function as: a reception unit that receives an instruction to switch between a summary video generated based on a summary of each segment generated from visual information and audio information extracted for each segment into which the video is divided and the original video; and a presentation unit that presents either the summary video or the original video on a screen based on the switching instruction received by the reception unit.
[0019] According to the present disclosure, it is possible to provide an information processing device, an information processing method, and a computer program that perform processing to generate a summary video based on both visual information and audio information contained in a video.
[0020] It should be noted that the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited to these. Furthermore, the present disclosure may also bring about additional effects in addition to the effects described above.
[0021] Further objects, features, and advantages of the present disclosure will become apparent from the following detailed description based on the embodiments and accompanying drawings.
[0022] FIG. 1 is a diagram showing the functional configuration of a summary video generation system 1000. FIG. 2 is a flowchart showing a processing procedure for the summary video generation system 1000 to generate a summary video. FIG. 3 is a diagram showing an example of a prompt to be input to a large-scale language model. FIG. 4 is a diagram showing one aspect of a viewing device. FIG. 5 is a diagram showing another aspect of a viewing device. FIG. 6 is a diagram showing yet another aspect of a viewing device. FIG. 7 is a diagram showing yet another aspect of a viewing device. FIG. 8 is a diagram showing an example of the configuration of a GUI screen provided by the viewing device. FIG. 9 is a diagram showing an example of the hardware configuration of an information processing device 2000.
[0023] Hereinafter, embodiments of the present disclosure will be described in the following order with reference to the drawings.
[0024] A. Overview B. Summary video generation system B-1. System configuration B-2. System operation B-3. Video segmentation B-4. Extraction of visual and audio information B-4-1. Extraction of visual information B-4-2. Extraction of audio information B-5. Summary text generation B-6. Calculation of summary length B-7. Summary video generation B-8. Speaker adaptation C. Video viewing using summary videos C-1. Aspects of viewing device C-2. User interface and interaction C-3. Video viewing function D. Contribution points of the summary video generation system E. Configuration of information processing device
[0025] A. Overview When attempting to extract important segments from a video by focusing on either the visual or audio information of the video, there is a risk of overlooking important information in a video in which both audio and visual information are important. Therefore, this disclosure proposes a technology for generating a video summary by taking into account both the visual and audio information of the video.
[0026] In this disclosure, a summary video is generated by combining a speech synthesis of a summary of a video that takes into account both visual and audio information with the corresponding part of the original video. In this process, a summary video that reflects the voice quality of the speaker of the original video is generated, thereby enabling seamless switching between the summary video and the original video.
[0027] In addition, the present disclosure adjusts the summarization rate by reflecting the amount of visual and audio information, allowing for the generation of a summary video that prevents information overload and deficiency while maintaining content important to the user. For example, a more detailed summary video is generated for content-rich sections of video, such as when a speaker speaks for a long time or when slides containing a lot of text are displayed in the video. This reduces the risk of the user missing important visual or audio information.
[0028] In short, according to the present disclosure, a video summary can be generated using a transcript of the audio of a video and images and text from the video, enabling individual users to view the summary without missing any important information. One of the major features of the video summary technology according to the present disclosure is that it does not create a summary by cutting out important parts from the original video, but instead generates a new summary video that includes important elements.
[0029] The present disclosure can further provide interactions to provide a user-centered learning experience. For example, a function for switching between a summary video and the original video is provided for each chapter of a video. Therefore, a user can flexibly adjust the pace of their learning by selecting an appropriate video for each chapter based on the content of the video, their own interests, and their level of understanding. For example, a summary video allows a user to quickly understand the overview of a chapter, and if more detailed information is needed, they can watch the original video for that chapter. Of course, a user can also view only the summary video to understand the overview of the video without switching between videos. Furthermore, the present disclosure can display a list of the title, summary, and thumbnail of each chapter, allowing a user to directly select and watch the chapter of interest.
[0030] B. Summary Video Generation System Next, a summary video generation system to which the present disclosure is applied will be described in detail.
[0031] B-1. System Configuration FIG. 1 schematically illustrates the functional configuration of a summary video generation system 1000 to which the present disclosure is applied. The illustrated summary video generation system 1000 includes a segment division unit 1001, a visual information extraction unit 1010, an audio information extraction unit 1020, a summary sentence generation unit 1030, and a summary video generation unit 1040 as basic components, and may further include components such as a summary length adjustment unit 1031 and a speaker characteristic extraction unit 1043. The summary video generation system 1000 may be configured using one or more information processing devices (e.g., personal computers (PCs)). Each component of the summary video generation system 1000 is implemented, for example, as a computer program running on a computer, but may also be implemented as a dedicated hardware device.
[0032] The video input to the summary video generation system 1000 is in a data format that combines video data and audio data, such as MPEG (Moving Picture Experts Group). The segmentation unit 1001 divides the input video into segments. A segment corresponds to a so-called "chapter." Even if chapters have been created in advance in the original video, the segmentation unit 1001 may divide the video into new segments, or the chapters of the original video may be used as segments as they are.
[0033] The visual information extraction unit 1010 extracts visual information contained in the video data of each segment and converts it into character information (hereinafter, the visual information extraction unit 1010 is referred to as "first character information"). The visual information extraction unit 1010 includes a character recognition unit 1011 and an object detection unit 1012. The character recognition unit 1011 optically recognizes characters (e.g., characters written on a blackboard) reflected in the video of the segment. The object detection unit 1012 recognizes objects (e.g., cars, people, etc.) reflected in the video of the segment and outputs the object names. Therefore, the optical character recognition results by the character recognition unit 1010 and the object detection results by the object detection unit 1012 become the first character information. Of course, in addition to character recognition and object detection, the visual information extraction unit 1010 may also include recognition or detection means for extracting visual information contained in the video to obtain the first character information. For example, the visual information extraction unit 1010 may further perform facial recognition, facial attribute analysis, gesture recognition, scene recognition, action recognition, etc. of people included in the video, and output these recognition and text history results as first text information.
[0034] The audio information extraction unit 1020 extracts the audio information contained in each segment and converts it into text information (hereinafter, the audio information extraction unit 1020 is referred to as "second text information"). Specifically, the audio information extraction unit 1020 includes a speech recognition unit 1021 that converts input speech into text using a speech recognition model, generates a transcript of the audio data of each segment, and outputs it as second text information.
[0035] The summary generation unit 1030 generates a summary for each segment from the first text information and the second text information using a large-scale language model such as GPT (Generative Pre-trained Transformer)-4 (see Non-Patent Document 2). By replacing the visual information and audio information contained in the video with text information, processing using a large-scale language model becomes possible.
[0036] The summary generation unit 1030 may generate a summary of a length (number of characters) that appropriately reflects the amount of information in the original video of each segment. For example, the summary length adjustment unit 1031 may adjust the length of the summary by adjusting the number of words L that the character recognition unit 1011 recognizes from the video of the segment. c , the number of objects detected by the object detection unit 1012 from the video of the segment L o , and the length L of the sentence transcribed by the speech recognition unit 1021 from the speech data of the segment. t The summary generator 1030 may be instructed on the number of characters N of the summary for each segment based on the above.
[0037] The summarized moving image generating unit 1040 generates a summarized moving image of the original moving image based on the summary of each segment generated by the summary generating unit 1030. The system includes an audio generating unit 1041 and a moving image creating unit 1042.
[0038] The speech generation unit 1041 uses speech synthesis technology to generate a reading voice of the summary of each segment generated by the summary generation unit 1030. When synthesizing the speech, speaker adaptation technology may be used. That is, the speech generation unit 1041 may generate a reading voice similar to the speech of the original video based on the speaker's speech features extracted from the speech of the segment by the speaker feature extraction unit 1043.
[0039] The video creation unit 1042 selects segments to be used in the summarized video based on the importance of each segment. For example, the video creation unit 1042 selects an optimal segment from the original video as a summarized video based on the length of the audio reading of the summary of each segment generated by the audio generation unit 1041, and determines the length of the audio reading of the summary of that segment as the summary of the summarized video. The video creation unit 1042 then creates a video of the summarized video that matches the determined audio reading of the summary. For example, the video creation unit 1042 may use the video of the segment from which the summary was selected as the video of the summarized video. Alternatively, the video creation unit 1042 may extract segments from multiple time positions, such as the beginning, middle, or end of the original video, to match the length of the summarized video (the length of the audio reading of the summary), and connect them together to generate the video of the summarized video. Furthermore, the video creation unit 1042 may generate video data of the selected segment from the summary of the segment using a generation model that generates video from text.
[0040] In this way, the summary video generation system 1000 outputs a summary video that combines the video created by the video creation unit 1042 in the summary video generation unit 1040 with the audio of the summary generated by the audio generation unit 1041. The summary video generation system 1000 can also generate and output "video segments" that combine the video of the segment, the summary, and the audio of the summary being read, for each segment into which the original video is divided.
[0041] B-2. System Operation Next, a description will be given of the system operation for generating a summary video from a video on the summary video generation system 1000 shown in Fig. 1. Fig. 2 shows, in the form of a flowchart, the processing procedure for the summary video generation system 1000 to generate a summary video.
[0042] In order to generate a digest video by summarizing the video into coherent segments (chapters), the segmentation unit 1001 first detects scene transitions and silent periods within the video and divides it into segments (step S201).
[0043] Then, a processing loop is started to generate a summary for each segment.
[0044] In this processing loop, first, the visual information extraction unit 1010 extracts visual information contained in the video data of the segment and converts it into first character information (step S202).
[0045] In the visual information extraction process of step S202, the character recognition unit 1011 optically recognizes characters reflected in the segment image, and the object detection unit 1012 recognizes objects such as cars and people reflected in the segment image, and first character information including the character recognition result and the object detection result is output.
[0046] Next, the speech information extraction unit 1020 extracts speech information included in the segment and converts it into second text information (step S203). The speech information extraction unit 1020 outputs the second text information consisting of a sentence transcribed from the speech data of the segment using a speech recognition model.
[0047] Next, the summary length adjustment unit 1031 calculates the number of characters N of the summary of each segment so that the length (number of characters) of the summary of each segment appropriately reflects the amount of information in the original video of the segment, and instructs the summary generation unit 1030 to do so (step S204).
[0048] Next, the summary generation unit 1030 generates a summary of each segment from the first character information and the second character information using the large-scale language model (step S205).
[0049] Next, the speech generation unit 1041 uses speech synthesis technology to generate a reading voice of the summary of each segment generated by the summary generation unit 1030 (step S206). The speech synthesis may utilize speaker adaptation technology. That is, the speech generation unit 1041 may generate a reading voice similar to the speech of the original video, based on the speaker's speech features extracted from the speech of the segment by the speaker feature extraction unit 1043.
[0050] When the processing loop for generating a summary for each segment is completed, a summary and its voice reading have been generated for each segment divided in step S201.
[0051] Then, the video creation unit 1042 selects segments to be used in the summary video based on the importance of each segment (step S207). For example, the video creation unit 1042 selects an optimal segment from the original video as the summary video based on the length of the audio read out of the summary of each segment generated by the audio generation unit 1041, and determines the length of the audio read out of the summary of that segment as the summary of the summary video.
[0052] Next, the video creation unit 1042 creates a video summary of the video that matches the determined summary sentence (step S208), and ends this process.
[0053] In step S208, the video creation unit 1042 may use the video of the segment from which the summary was selected as the video of the summarized video. Alternatively, the video creation unit 1042 may extract segments from multiple time positions, such as the beginning, middle, or end of the original video, according to the length of the summarized video (the length of the audio reading of the summary), and connect them to generate the video of the summarized video. Furthermore, the video creation unit 1042 may generate video data of the selected segment from the summary of that segment using a generative model that generates video from text.
[0054] 2, the summary video generation system 1000 can output a summary video that combines video created by the video creation unit 1042 from the original video with audio reading of the summary generated by the audio generation unit 1041. Although not shown in detail in FIG. 2, the summary video generation system 1000 can generate and output "video segments" that combine video of the segment, a summary, and audio reading of the segment for each segment into which the original video is divided.
[0055] B-3. Video Segment Division In the digest video generation system 1000 according to the present disclosure, in order to generate a digest video by summarizing the video into coherent segments (chapters), the segment division unit 1001 first detects scene transitions and silent periods within the video and divides it into segments. Section B-3 provides a detailed description of the processing performed by the segment division unit 1001.
[0056] For scene transition detection, the segmentation unit 1001 analyzes variations in color distribution (histogram) between consecutive frames and determines segment boundaries based on the magnitude of the variations. Clear variations between consecutive frames are known to indicate scene changes, and histogram variations are likely to indicate changes in the mood or background of the video.
[0057] Furthermore, taking into consideration that visual divisions and audio divisions do not necessarily coincide, the segmentation unit 1001 uses an algorithm that detects periods in which the sound waveform has an amplitude below a specific threshold as silence. The segmentation unit 1001 detects periods of low amplitude or the absence of specific frequencies to identify silent or quiet segments.
[0058] The segmentation unit 1001 divides the video at good points and segments it, taking into consideration both the visual scene transitions and silent periods in the audio. By using this function, the segmentation unit 1001 can automatically divide the video into chapters (segment division) without attaching tags to each coherent segment.
[0059] The segment division unit 1001 may perform processing to divide a video into segments using an existing library that is saved, shared, and made public, for example, through a source code management service.
[0060] B-4. Extraction of Visual and Audio Information In order to generate a digest video that takes into account both the visual and audio information of a video, the digest video generation system 1000 according to the present disclosure converts this information into text information.
[0061] B-4-1. Extraction of Visual Information The visual information extraction unit 1010 uses optical character recognition (OCR) and object detection technologies to acquire visual information (i.e., text information in the video) from the video. For this purpose, the visual information extraction unit 1010 includes a character recognition unit 1011 and an object detection unit 1012.
[0062] The character recognition unit 1011 uses OCR to detect characters that appear in the video of the segment and converts them into text. For example, the character recognition unit 1011 is effective for acquiring text information from a video of a lecture, such as the contents of slides or handwritten notes on a whiteboard. By including this information in the first text information, the themes and contents of the text in the video can be accurately reflected in the summary of the segment.
[0063] Meanwhile, the object detection unit 1012 detects and identifies specific objects or entities from each frame of the video of the segment and outputs the object names. Objects in the video are often closely related to the focus or theme of the story of the segment. Therefore, by including the object detection result by the object detection unit 1012 in the first text information, it becomes possible to identify key scenes and important entities in the video and generate an effective summary of the segment based on that information.
[0064] In the summarized video generation system 1000 according to the present disclosure, for example, the character recognition unit 1011 applies the Tesseract-OCR Engine for OCR. Furthermore, the object detection unit 1012 uses Faster R-CNN, which has ResNet-50-FPN as its backbone, as its object detection model. However, the present disclosure is not limited to specific OCR or object detection model techniques. Furthermore, CLIP or GPT-4V (GPT-4 with vision) or the like may be used to extract information other than characters and objects from an image, and this information may be included in the first information and used to generate a summary of the segment. The character recognition unit 1011 and the object detection unit 1012 may perform character recognition processing and object detection processing, respectively, using existing libraries that are stored, shared, or made public, for example, through a source code management service.
[0065] B-4-2. Extraction of Audio Information Accurate acquisition of audio information is also required for the spoken content of videos, particularly academic content such as lectures and seminars. In the summary video generation system 1000 according to the present disclosure, the audio information extraction unit 1020 includes a speech recognition unit 1021 that converts input audio into text using a speech recognition model. The audio of the video is input to the speech recognition unit 1021, and a transcript of the audio data for each segment is generated and output as second text information. Because the accuracy of the speech recognition here affects the quality of the summary, it is desirable to convert the audio signal into text data with high accuracy.
[0066] Recent advances in speech recognition are attributable to models based on RNNs and transformers. The speech recognition unit 1021 uses Whisper (see Non-Patent Document 1), one of these models, as a speech recognition (Speech2Text) model that converts speech into text. Whisper has been trained with approximately 680,000 hours of multilingual speech collected from the web, and its recognition accuracy has been confirmed to be comparable to that of humans. However, the present disclosure is not limited to a specific speech recognition method, and any method can be used. The speech recognition unit 1021 may perform speech recognition processing using an existing library that is stored, shared, or made public, for example, through a source code management service.
[0067] Although the performance of speech recognition techniques has improved dramatically in recent years, the error rate is not zero. Therefore, the summarized video generation system 1000 according to the present disclosure is designed to cover up errors in speech recognition by using the first character information extracted from the video by the visual information extraction unit 1010 using OCR or an object detection model as described above, so as to prevent important words in the video from being missed from the summary.
[0068] B-5. Summary Generation In typical video summarization, data labeled with important points in a video is used to detect important locations within the video using supervised learning techniques. However, videos available to the public span a wide range of domains, making it difficult to use training data that extracts important parts from all domains. In recent years, in the field of natural language processing, numerous large-scale language models (LLMs) based on transformer architectures such as BERT, T5, and GPT have been proposed. These models are trained using documents and dialogues from various domains, and therefore possess knowledge of a wide range of domains, enabling a variety of natural language processing, including summary generation. Therefore, the summary generation unit 1030 utilizes these LLMs to enable video summarization without the need for training data. The summary generation unit 1030 generates a summary for each segment using both a transcript of the utterance generated by a speech recognition model (i.e., second text information) and visual information acquired through OCR or object detection (i.e., first text information), thereby reflecting the visual information contained in the video of the segment in the summary. The summary generator 1030 generates a summary of each video segment by inputting a prompt to the LLM that generates a summary that takes into account both visual and audio information, as shown in Figure 3. The summary generator 1030 may generate a summary of the transcript using an existing library that is stored, shared, or made public, for example, in a source code management service.
[0069] Along with the instructional statements, the transcript of each video segment and visual information (i.e., textual information in the video) obtained through OCR and object detection are input as text into the LLM. This method allows for the generation of summaries that highlight nuances not contained in the spoken content, speaker movements, visual objects, and text within the slides and visual presentation. Any model, such as ChatGPT, LlaMA, or PaLM, can be used as the LLM.
[0070] B-6. Calculation of Summary Length When generating a summary using the summary generator 1030, it is necessary to determine the length of the document summary in a way that appropriately reflects the amount of information in the original video for each segment. In other words, if the video is long or has a lot of content, the summary must provide sufficient information, while parts that are shortened in the original video must also be shortened to eliminate redundancy.
[0071] Therefore, the summary length adjustment unit 1031 adjusts the number of characters in the summary output from the LLM depending on the amount of text transcribed from the audio data for each segment. Furthermore, because the summary generation unit 1030 generates a summary for each segment based not only on the audio of the video but also on visual information, the summary length adjustment unit 1031 adjusts the number of characters in the summary output from the LLM, taking into account the amount of visual information, i.e., the first text information. For example, if the number of characters on the blackboard or slides shown in the video is large, even if the speaker's remarks are short, it is necessary to accurately reflect the information content on the blackboard or slides in the summary. From this perspective, the summary length adjustment unit 1031 calculates the number of characters N in the summary using the following equation (1):
[0072]
[0073] In the above formula (1), N is the final number of characters in the summary output from the LLM, L t is the length of the transcribed sentence by the speech recognition model, L o is the number of objects identified in a video frame by object recognition, L c is the number of words in the visual elements recognized from the video by OCR. s And lol i is a weighting coefficient for adjusting the importance of each of the visual information (first character information) and the audio information (second character information). s And lol i is selected based on experimental data and user preferences. By changing these two weighting factors, the weighting of visual information and audio information can be adjusted. The weighting factor w s And lol iBy changing the weighting coefficient w, it is possible to provide a personalized summary video. For example, by managing user profiles and user video viewing histories, it is possible to change the weighting coefficient w based on the user information. s And lol i Alternatively, the number of characters to be output, N, may be changed, or the number of characters to be output, N, may be determined. The value "50" in the parentheses on the right side of the above formula (1) is a value for preventing the summary length (number of characters) from becoming shorter than 50, but can be changed to any value depending on the application.
[0074] B-7. Generating Summarized Videos Conventional audio summarization methods for audiobooks and other audio sources involve evaluating the importance of each individual audio transcript, then directly extracting and combining important audio segments. However, this method can pose problems with the continuity of the audio and video when combining the extracted segments. Furthermore, this method cannot create a summary video that reflects the visual presentation of the video.
[0075] Therefore, the summary video generation unit 1040 employs a method using voice synthesis. Specifically, the voice generation unit 1041 receives the summary of each segment generated by the summary generation unit 1030 using visual information and audio information, and generates a voice reading it using voice synthesis technology. After generating the synthesized voice for each segment, the video creation unit 1042 determines the importance of each segment based on the summary and the voice reading it, and selects the segments to use in the summary video.
[0076] For example, the video creation unit 1042 selects an optimal video segment from the original video based on the length of the audio. The video creation unit 1042 may also extract a video segment from the beginning, middle, or end of the original video to match the length of the summarized video. Alternatively, the video creation unit 1042 may extract and splice any video portion to match the generated audio. In this manner, by combining the selected video segment with the synthesized audio, a continuous and concise summarized video can be realized. The video creation unit 1042 may also generate video data for the selected segment from a summary of the segment using a generative model that generates video from text. Examples of generative models that generate video data from text include Video Diffusion Models, ModelScope T2V, Make-A-Video, and Sora.
[0077] The criteria for determining the importance of each segment can be arbitrary in the video creation unit 1042. For example, a user profile, a user's video viewing history, etc. may be managed, and a method for determining the importance of a segment may be determined taking into consideration this user information.
[0078] B-8. Speaker Adaptation When using a summarized video generated by the summarized video generation system 1000 according to the present disclosure, it is expected that users will frequently switch between the original video and the summarized video, so it is necessary to provide a seamless experience. Therefore, by using speaker adaptation technology when synthesizing the reading voice of the summary text, the audio generation unit 1041 generates a summarized video that minimizes changes in the voice from the voice of the original video. That is, the audio generation unit 1041 generates reading voice similar to the speech of the original video based on the speaker's voice features extracted from the audio of the segment by the speaker feature extraction unit 1043. This technique can reduce discrepancies in voice features when switching between the original video and the summarized video, thereby achieving seamless switching.
[0079] C. Viewing Videos Using Summary Videos Section B above described the system configuration and processing operations of the summary video generation system 1000 according to the present disclosure for generating a summary video from an original video. The summary video is composed of a combination of a speech read out by synthesizing a summary generated by taking into account both visual and audio information, and a corresponding portion of the original video. Furthermore, the summary video generation system 1000 according to the present disclosure can obtain a summary generated for each segment into which the original video is divided, and the speech read out of that summary.
[0080] On the viewing device side where the video is viewed, seamless switching between the summarized video and the original full version of the video can be performed by using the summarized video provided by the summarized video generation system 1000 according to the present disclosure. The summarized video generation system 1000 according to the present disclosure can provide an accurate summarized video that does not miss any important information. Therefore, by watching the summarized video created by the summarized video generation system 1000, users who are limited in time or who are interested in a variety of topics can quickly check only the important points of the original long video, and can quickly determine whether it is worth watching the original full version of the video or understand the aspects of many videos.
[0081] The viewing device can also provide a function for quickly switching between the original full-length video and shorter videos with audio summaries for each segment (chapter). Users can quickly grasp the overview of the entire segment through the segment summaries and their audio readings, and can watch the original video for that segment if more detailed information is needed. For example, when watching a lecture video, users can flexibly adjust their learning pace by selecting the appropriate video for each segment based on the content of the video, their own interests, and their level of comprehension. Furthermore, users can improve their learning efficiency by grasping only the gist of the parts of the video they already fully understand and repeating or viewing more slowly the important parts they have not yet fully understood.
[0082] C-1. Configuration of Viewing Device One configuration of a viewing device is shown in Fig. 4. In the example shown in Fig. 4, the viewing device 400 is configured as the same information processing device as the summary video generation system 1000 (i.e., as a device that is physically integrated with the summary video generation system 1000).
[0083] When viewing a video provided by a video providing device 402, the viewing device 401 presents a summary video created by the functions of an internal summary video generating system 1000, so that the user can quickly check only the important points of the original long video and quickly determine whether it is worth watching the original full version of the video.
[0084] In addition, the viewing device 401 provides a summary for each segment created using the functions of the internal summary video generation system 1000 and an audio readout of the summary, allowing the user to quickly grasp the outline of each segment through the summary for each segment and the audio readout, and if more detailed information is required, the user can watch the original video of that segment.
[0085] The video providing device 402 may be, for example, a television broadcasting station that broadcasts video content or a streaming server that distributes video content via streaming. Alternatively, the video may be provided to the viewing device 401 in the form of a recording medium such as a Blu-ray or DVD, rather than a device.
[0086] Fig. 5 shows another aspect of a viewing device. In the example shown in Fig. 5, the viewing device 501 is configured as an information terminal (e.g., a PC, a tablet, a smartphone, etc.) separate from the summary video generation system 1000. In this case, the viewing device 501 and the summary video generation system 1000 may be connected via, for example, a wired or wireless network. The viewing device 501 views videos provided by a video providing device 502.
[0087] 6 shows yet another embodiment of a viewing device. The summary video generation system 1000 may be arranged on a cloud, and a viewing device 601 may be connected to the summary video generation system 1000 via a wide area network such as the Internet. The viewing device 601 then views videos provided by a video providing device 602.
[0088] 7 shows yet another aspect of a viewing device. The summary moving image generation system 1000 is, for example, a server that provides services to a plurality of viewing devices 701-1, 701-2, .... Each of the viewing devices 701-1, 701-2, ... views a moving image provided by a moving image providing device 702.
[0089] 5 to 7, the summary video generation system 1000 creates a summary video and generates a summary for each segment and a voice reading of the summary for that segment for the same video that is being viewed by the viewing devices 501..., and provides the summary to the viewing devices 401.... Therefore, when viewing a video, a user can quickly check only the important points of the original long video from the summary video created by the internal summary video generation system 1000, and quickly determine whether it is worth watching the full original version of the video. In addition, by viewing the summary for each segment and the voice reading of the summary for that segment, the user can quickly grasp the overview of each segment, and if more detailed information is required, the user can watch the original video for that segment.
[0090] 4 to 7, the summary video generation system 1000 may perform pre-processing to create a summary video, a summary for each segment, and a read-out voice for each video provided by the video providing devices 402.... If such pre-processing has been performed, the viewing devices 401... can start viewing the video and at the same time immediately check the original video using the summary video and the summary for each segment.
[0091] 4 to 7 introduced in Section C-1 above, by using the summary video generation system 1000 according to the present disclosure, it is possible to design a user interface and interaction so that the user can switch between the original full version video and the summary video for each video according to the user's own interest and level of understanding. In Section C-2, we will explain the user interface and interaction on the viewing device that can be realized by using the summary video generation system 1000 according to the present disclosure.
[0092] 8 shows an example of the configuration of a GUI (Graphical User Interface) screen provided by a viewing device. The viewing device referred to here may be any of the aspects shown in FIGS. 4 to 7, but it is assumed that the viewing device uses the summarized video generation system 1000 according to the present disclosure. Such a GUI screen may be, for example, an application window or a browser window for viewing a specific website.
[0093] A "video window" denoted by reference numeral 801 is located in the center of the left half of the screen. The video window 801 displays either a summarized video or the original full version of the video. The user can view the video in either the summarized or original format in this video window 801. A "video title" denoted by reference numeral 802 is located at the top of the video window 801 and displays the title of the segment currently being played. Note that the method for creating the segment title is arbitrary. For example, the title may be automatically generated based on the segment summary generated by the summary generation unit 1030.
[0094] Below the video window 801, there are arranged a "Summary" button indicated by the reference numeral 803 and an "Original" button indicated by the reference numeral 804. When the "Summary" button 803 is clicked, the video displayed in the video window 801 switches to a summary video, and when the "Original" button 804 is clicked, the video displayed in the video window 801 switches to the original full version video. In other words, by operating the buttons, it is possible to quickly switch the video displayed in the video window 801. The original full version video is either the entire original video or the original video of a segment selected by the user from the original video, and can be selected by the user.
[0095] When a summary video is displayed in the video window 801, a voice reading out the summary text synthesized by the voice generating unit 1041 is output. Also, when an original video is displayed in the video window 801, the voice of the original video is output (when a segment video is displayed, the voice of the segment is output).
[0096] A "segment window" designated by the reference numeral 805 is located in the right half of the screen. Within the segment window 805, thumbnails of each segment divided from the original video are displayed in a list. Additionally, the segment title and summary are displayed to the right of each segment. If it is not possible to display all the segment thumbnails within the segment window 805, an up and down scroll bar or the like may be provided so that the thumbnails of all the segments can be viewed by scrolling the window.
[0097] The user can get an overview of the entire segment from the thumbnail and summary of each segment in the segment window 805. The user can also quickly access the segment they need. When the user clicks on the thumbnail of a segment for which they need more detailed information, the video window 801 switches to playing the video of that segment.
[0098] A "search window" denoted by reference numeral 806 is located above the segment window 805. The search window 806 includes a text input box for searching for specific keywords within the segment window 805, and a search button. When a user inputs a desired keyword into the text input box and presses the search button, a keyword search is performed on the abstracts of each segment within the segment window 805, and the parts of each abstract that match the keyword are highlighted. Therefore, by further utilizing such search results, a user can quickly access segments for which they require more detailed information or that interest them.
[0099] C-3. Video Viewing Function Through a GUI screen such as that shown in FIG. 8, a user can adjust how much time to spend watching a video, as well as where to spend time and where not to spend time. Therefore, the summary video generation system 1000 according to the present disclosure can provide a personalized learning experience according to each user's interests and level of understanding. Below are some examples of video viewing functions provided using the GUI screen shown in FIG. 8.
[0100] (1) Checking the entire summary By using the Summary button 803, the user can view a summary of the video. By viewing the video in this state, the user can quickly grasp the outline of the video. This function is similar to the act of getting a rough idea of the contents of a book.
[0101] (2) Detailed Review of Specific Sections By using the Segment window 805, users can learn about the content of a segment in advance through related thumbnails and documents. By selecting a segment of interest, users can view only that section in detail. For deeper understanding, users can further select the Original button 804 to view the original complete version of the segment. This function is similar to picking out and reading a specific section or chapter in a book.
[0102] (3) Repeated viewing of specific parts: Users can replay specific segments multiple times to deepen their understanding, similar to the situation when reading a book and repeatedly rereading a passage that they found important.
[0103] (4) Searching for Information Within a Video by Keyword: Users can instantly search for and access segments related to a specific term or keyword through the search window 806. This feature provides an experience similar to finding information by browsing the index of a book.
[0104] (5) Selection of Simplified and Detailed Views Users can freely switch between the summary video and the full original video using the Summary button 803 and the Original button 804. This operation is similar to the combination of speed reading and detailed reading when reading. By combining speed reading and detailed reading, readers can quickly grasp the overall overview while gaining a deeper understanding of the necessary parts. This function applies this reading strategy to video viewing, allowing users to focus on segments of interest without losing the overall context.
[0105] D. Contributions of the Summary Video Generation System The contributions of the summary video generation system 1000 according to the present disclosure are summarized below.
[0106] (1) The video summary generation system 1000 according to the present disclosure can generate a video summary that reflects both important visual and audio information by comprehensively considering the visual and audio information of the video. (2) The ability to freely switch between the original video and the video summary allows users to choose whether to obtain concise information or detailed information in each chapter of the video, providing a user-centered learning experience. (3) By applying the video summary generation system 1000 according to the present disclosure to lecture videos, users can efficiently acquire information from long video content.
[0107] E. Configuration of Information Processing Device The summary video generation system 1000 according to the present disclosure can be configured using one or more information processing devices. Furthermore, a viewing device capable of viewing a summary video can be configured as an information processing device that is physically integrated with the summary video generation system 1000, or can be configured as an information processing device separate (physically separated) from the summary video generation system 1000.
[0108] 9 shows an example of the hardware configuration of an information processing device 2000 that can operate as the summary video generation system 1000 or a viewing device. The information processing device 2000 includes a CPU (Central Processing Unit) 2001, a ROM (Read Only Memory) 2002, a RAM (Random Access Memory) 2003, a host bus 2004, a bridge 2005, an expansion bus 2006, an interface unit 2007, an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013. The information processing device 2000 is configured, for example, by an information terminal such as a personal computer, a tablet, or a smartphone.
[0109] The CPU 2001 controls the overall operation of the information processing device 2000 in accordance with various programs. When performing processing with a high computational load on the information processing device 2000, it is desirable that the CPU 2001 be a multi-core CPU (e.g., Apple M1 Max, etc.), or that the information processing device 2000 further be equipped with a multi-core processor (e.g., NVIDIA's "Quadro A6000") such as a GPU (Graphics Processing Unit) or GPGPU (General-purpose computing on graphics processing units) in addition to the CPU 2001. However, hereinafter, for convenience, these will be collectively referred to simply as the CPU 2001.
[0110] The ROM 2002 stores in a nonvolatile manner programs (such as a basic input / output system) and calculation parameters used by the CPU 2001. The RAM 2003 is used to load programs to be executed by the CPU 2001 and to temporarily store parameters such as working data that change as appropriate during program execution. Programs loaded into the RAM 2003 and executed by the CPU 2001 include, for example, various application programs and an operating system (OS).
[0111] The CPU 2001, ROM 2002, and RAM 2003 are interconnected by a host bus 2004, which includes a CPU bus and the like. The CPU 2001 executes various application programs in an execution environment provided by the OS through the cooperative operation of the ROM 2002 and RAM 2003, thereby realizing various functions and services. If the information processing device 2000 is a personal computer, the OS may be, for example, Microsoft Windows (registered trademark), Unix (registered trademark), or a successor OS. Examples of application programs executed on the information processing device 2000 include the following. Note that the application program or some of the modules in the application program may use existing libraries that are stored, shared, or made public, for example, through a source code management service. (1) A video summary generation program that generates a video summary from a video. (2) A video viewing program that utilizes a video summary to watch the original video or watch it in segments.
[0112] The host bus 2004 is connected to an expansion bus 2006 via a bridge 2005. The expansion bus 2006 is, for example, a PCI (Peripheral Component Interconnect) bus or PCI Express, and the bridge 2005 is based on the PCI standard. However, the information processing device 2000 does not need to be configured so that the circuit components are separated by the host bus 2004, bridge 2005, and expansion bus 2006, and may be implemented so that almost all circuit components are interconnected by a single bus (not shown).
[0113] The interface unit 2007 connects peripheral devices such as an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013 in accordance with the standards of the expansion bus 2006. However, not all of the peripheral devices shown in Fig. 9 are necessarily required, and the information processing device 2000 may further include peripheral devices not shown. Furthermore, the peripheral devices may be built into the main body of the information processing device 2000, or some of the peripheral devices may be externally connected to the main body of the information processing device 2000.
[0114] The input unit 2008 is composed of an input control circuit that generates an input signal based on input from a user and outputs it to the CPU 2001. When the information processing device 2000 is a personal computer, the input unit 2008 may include a keyboard, a mouse, a touch panel, and may further include a camera and a microphone used for remote conferences and face-to-face customer service. The output unit 2009 includes display devices such as a liquid crystal display (LCD) device, an organic electroluminescence (EL) display device, and an LED (light emitting diode), as well as an audio output device such as a speaker. The input unit 2008 is used to input a video to be processed, and the output unit 2009 is used to display a GUI screen (see, for example, FIG. 8 ).
[0115] The storage unit 2010 stores files such as programs (applications, OS, etc.) executed by the CPU 2001 and various data. The storage unit 2010 is configured with a large-capacity storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive), but may also include an external storage device.
[0116] The removable storage medium 2012 is a storage medium configured as a cartridge, such as a microSD card. The drive 2011 performs read and write operations on the loaded removable storage medium 113. The drive 2011 outputs data read from the removable storage medium 2012 to the RAM 2003 or the storage unit 2010, and writes data on the RAM 2003 or the storage unit 2010 to the removable storage medium 2012.
[0117] The communication unit 2013 is a device that performs wireless communication via Wi-Fi (registered trademark), Bluetooth (registered trademark), or cellular communication networks such as 4G and 5G. The communication unit 2013 may also include terminals such as a Universal Serial Bus (USB) or a High-Definition Multimedia Interface (HDMI) (registered trademark), and may further include a function for performing HDMI (registered trademark) communication with USB devices such as scanners and printers, displays, and the like. Programs executed on the information processing device 2000 are installed from an external device, for example, via the communication unit 2013. An acoustic signal that is the subject of the summary generation process according to the present disclosure is captured, for example, via the communication unit 2013.
[0118] The present disclosure has been described in detail above with reference to specific embodiments. However, the present disclosure should not be construed as being limited to the above-described embodiments, and it is obvious that those skilled in the art can modify or substitute the embodiments without departing from the spirit of the present disclosure. Furthermore, the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited thereto, and additional effects not described in this specification may exist.
[0119] The following two examples can be given as application examples of the present disclosure.
[0120] (1) Online education platforms, corporate training and professional development programs, and short-term learning for self-improvement and hobby learning. The present disclosure can also be applied to various categories of videos that contain both audio and visual information, such as news, animation, movies, and cooking videos. (2) Reducing the time required for video editing, automatically creating digest videos, or quickly grasping the parts necessary to create a digest video from a long video. For example, the present disclosure can be applied to video editing by entertainers, original video contributors, and artists, video editing for general users' Vlogs, and the work of filmmakers quickly grasping the parts necessary from past and current works and extracting only the necessary parts.
[0121] The present disclosure can be applied to generating a summarized video of a video in which both audio and visual information are important, such as a lecture video or a news video, but can also be applied to other categories of videos in which both audio and visual information are important. Of course, the present disclosure can also be applied to categories of videos in which both audio and visual information are not necessarily important.
[0122] In short, the present disclosure has been described in the form of examples, and the contents of the specification should not be interpreted as limiting. To determine the gist of the present disclosure, the claims should be taken into consideration.
[0123] The series of processes described in this specification can be executed by hardware, software, or a configuration that combines hardware and software. When executing processes by software, a program recording a processing sequence related to realizing the present disclosure is installed in memory in a computer incorporated in dedicated hardware and executed. It is also possible to install the program in a general-purpose computer capable of executing various processes and execute the processes related to realizing the present disclosure.
[0124] The program can be stored in advance on a recording medium installed in the computer, such as a HDD, SSD, or ROM. Alternatively, the program can be temporarily or permanently stored on a removable recording medium such as a flexible disk, CD-ROM (Compact Disc Read Only Memory), MO (Magneto Optical) disk, DVD (Digital Versatile Disc), BD (Blu-Ray Disc (registered trademark)), magnetic disk, or USB (Universal Serial Bus) memory. Using such a removable recording medium, a program related to the realization of the present disclosure can be provided as so-called package software.
[0125] The program may also be transferred wirelessly or via a wire from a download site to a computer via a network such as a wide area network (WAN) typified by cellular, a local area network (LAN), the Internet, etc. The computer can receive the program transferred in this manner and install it in a large-capacity storage device such as an HDD or SSD within the computer.
[0126] The present disclosure may also be configured as follows.
[0127] (1) An information processing device comprising: a division unit that divides a video into segments; a visual information extraction unit that converts visual information contained in the video for each segment into first character information; an audio information extraction unit that converts audio information contained in the video for each segment into second character information; a summary sentence generation unit that generates a summary sentence for each segment of the video based on the first character information and the second character information; and a summary video generation unit that generates a summary video of the original video based on the summary sentence generated by the summary sentence generation unit.
[0128] (2) The information processing device according to (1), wherein the visual information extraction unit converts character information included in the video frames into first character information through optical character recognition.
[0129] (3) The information processing device according to any one of (1) or (2), wherein the visual information extraction unit detects an object included in a video frame using an object detection model and converts the object into first character information.
[0130] (4) The information processing device according to any one of (1) to (3), wherein the audio information extraction unit converts audio information included in the video into second character information using a voice recognition model.
[0131] (5) The information processing device according to any one of (1) to (4), wherein the summary generation unit generates the summary from the first character information and the second character information using a large-scale language model.
[0132] (5-1) The information processing device described in (5) above, wherein the summary generation unit inputs the first character information and the second character information into a large-scale language model together with a prompt to generate a summary taking into account both the first character information and the second character information, to obtain the summary.
[0133] (6) The information processing device according to any one of (1) to (5) above, further comprising: a summary length adjustment unit that adjusts the number of characters output by the summary generation unit based on the amount of information in each segment.
[0134] (7) The information processing device described in (6) above, wherein the summary length adjustment unit adjusts the number of characters output by the summary generation unit based on at least one of the number of characters of the second character information output from the audio information extraction unit, the number of objects detected from the video frame by the visual information extraction unit, and the number of words recognized as characters from the video frame by the visual information extraction unit.
[0135] (8) The information processing device according to any one of (1) to (7), wherein the summary video generation unit generates the summary video using videos of segments selected based on the importance of each segment and audio of summaries of the segments.
[0136] (9) The information processing device according to (8), wherein the summary video generation unit selects one or more segments to be used in the summary video based on the length of the synthesized speech.
[0137] (10) The information processing device according to (8), wherein the summary video generation unit selects one or more segments to be used in the summary video based on a time position of each segment in the original video.
[0138] (11) An information processing device according to any one of (8) to (10) above, further comprising a speaker characteristic extraction unit that extracts speaker characteristics of the audio of the segment, and the summary video generation unit generates the summary video using a voice reading a summary sentence that has been voice-synthesized to adapt to the speaker characteristics extracted from the audio of the selected segment.
[0139] (12) An information processing method comprising: a division step of dividing a video into segments; a visual information extraction step of converting visual information contained in the video for each segment into first character information; an audio information extraction step of converting audio information contained in the video for each segment into second character information; a summary generation step of generating a summary for each segment of the video based on the first character information and the second character information; and a summary video generation step of generating a summary video of the original video based on the summary generated in the summary generation step.
[0140] (13) A computer program written in a computer-readable format to cause a computer to function as: a division unit that divides a video into segments; a visual information extraction unit that converts visual information contained in the video for each segment into first character information; an audio information extraction unit that converts audio information contained in the video for each segment into second character information; a summary generation unit that generates a summary for each segment of the video based on the first character information and the second character information; and a summary video generation unit that generates a summary video of the original video based on the summary generated by the summary generation unit.
[0141] (14) An information processing device comprising: a reception unit that receives an instruction to switch between a summary video generated based on a summary sentence of each segment generated from visual information and audio information extracted for each segment into which a video is divided and the original video; and a presentation unit that presents either the summary video or the original video on a screen based on the switching instruction received by the reception unit.
[0142] (15) The information processing device according to (14), wherein the presentation unit outputs a voice reading a summary of a segment corresponding to the summary video while presenting the summary video on the screen.
[0143] (16) The information processing device described in any one of (14) or (15) above, wherein the presentation unit further presents a thumbnail and summary of each segment, and in response to the reception unit receiving an instruction to switch to any segment, switches the screen to a video of the segment instructed to switch.
[0144] (17) The information processing device according to (16), wherein the accepting unit accepts input of a search keyword, and the presenting unit presents, on the screen, a search result for the input search keyword in the summary sentence of each segment.
[0145] (18) An information processing method comprising: a receiving step of receiving an instruction to switch between a summary video generated based on a summary of each segment generated from visual information and audio information extracted for each segment into which the video is divided and the original video; and a presentation step of presenting either the summary video or the original video on a screen based on the switching instruction received in the receiving step.
[0146] (19) A computer program written in a computer-readable format to cause a computer to function as: a reception unit that receives an instruction to switch between a summary video generated based on a summary of each segment generated from visual information and audio information extracted for each segment into which the video is divided and the original video; and a presentation unit that presents either the summary video or the original video on a screen based on the switching instruction received by the reception unit.
[0147] 1000... Summary video generation system, 1001... Segment division unit, 1010... Visual information extraction unit, 1011... Character recognition unit, 1012... Object detection unit, 1020... Audio information extraction unit, 1021... Audio recognition unit, 1030... Summary sentence generation unit, 1031... Summary length adjustment unit, 1040... Summary video generation unit, 1041... Audio generation unit, 1042... Video creation unit, 1043... Speaker characteristic extraction unit, 2000... Information processing device, 2001... CPU, 2002... ROM, 2003... RAM, 2004... Host bus, 2005... Bridge, 2006... Expansion bus, 2007... Interface unit, 2008... Input unit, 2009... Output unit, 2010... Storage unit, 2011... Drive, 2012... Removable recording medium, 2013... Communication unit
Claims
1. An information processing device comprising: a division unit that divides a video into segments; a visual information extraction unit that converts visual information contained in the video for each segment into first character information; an audio information extraction unit that converts audio information contained in the video for each segment into second character information; a summary generation unit that generates a summary for each segment of the video based on the first character information and the second character information; and a summary video generation unit that generates a summary video of the original video based on the summary generated by the summary generation unit.
2. The information processing device according to claim 1, wherein the visual information extraction unit converts character information contained in the video frames into first character information through optical character recognition.
3. The information processing device according to claim 1, wherein the visual information extraction unit detects an object included in a video frame using an object detection model and converts the detected object into first character information.
4. The information processing device according to claim 1, wherein the audio information extraction unit converts audio information included in the video into second character information using a voice recognition model.
5. The information processing device according to claim 1, wherein the summary generation unit generates the summary from the first character information and the second character information using a large-scale language model.
6. The information processing device according to claim 1, further comprising a summary length adjustment unit that adjusts the number of characters output by the summary generation unit based on the amount of information in each segment.
7. The information processing device according to claim 6, wherein the summary length adjustment unit adjusts the number of characters output by the summary generation unit based on at least one of the number of characters of the second character information output from the audio information extraction unit, the number of objects detected from the video frame by the visual information extraction unit, and the number of words recognized as characters from the video frame by the visual information extraction unit.
8. The information processing device according to claim 1, wherein the summary video generation unit generates a summary video using videos of segments selected based on the importance of each segment and audio of summaries of the segments.
9. The information processing device according to claim 8, wherein the summary video generation unit selects one or more segments to be used in the summary video based on the length of the synthesized speech read aloud.
10. The information processing device according to claim 8, wherein the summary video generation unit selects one or more segments to be used in the summary video based on the time position of each segment in the original video.
11. An information processing device as described in claim 8, further comprising a speaker characteristic extraction unit that extracts speaker characteristics from the audio of the segment, and wherein the summary video generation unit generates the summary video using a voice reading a summary sentence that has been voice-synthesized to adapt to the speaker characteristics extracted from the audio of the selected segment.
12. An information processing method comprising: a division step of dividing a video into segments; a visual information extraction step of converting visual information contained in the video for each segment into first character information; an audio information extraction step of converting audio information contained in the video for each segment into second character information; a summary generation step of generating a summary for each segment of the video based on the first character information and the second character information; and a summary video generation step of generating a summary video of the original video based on the summary generated in the summary generation step.
13. A computer program written in a computer-readable format to cause a computer to function as: a division unit that divides a video into segments; a visual information extraction unit that converts visual information contained in the video for each segment into first character information; an audio information extraction unit that converts audio information contained in the video for each segment into second character information; a summary generation unit that generates a summary for each video segment based on the first character information and the second character information; and a summary video generation unit that generates a summary video of the original video based on the summary generated by the summary generation unit.
14. An information processing device comprising: a reception unit that receives an instruction to switch between a summary video generated based on a summary of each segment generated from visual information and audio information extracted for each segment into which a video is divided and the original video; and a presentation unit that presents either the summary video or the original video on a screen based on the switching instruction received by the reception unit.
15. The information processing device according to claim 14, wherein the presentation unit outputs a voice reading a summary of a segment corresponding to the summary video while the summary video is being presented on the screen.
16. The information processing device according to claim 14, wherein the presentation unit further presents a thumbnail and summary of each segment, and in response to the reception unit receiving an instruction to switch to any of the segments, switches the screen to a video of the segment instructed to be switched.
17. The information processing device according to claim 16, wherein the reception unit receives input of a search keyword, and the presentation unit presents on the screen the results of searching the input search keyword in the summary sentences of each segment.
18. An information processing method comprising: a receiving step of receiving an instruction to switch between a summary video generated based on a summary of each segment generated from visual information and audio information extracted for each segment into which the video is divided and the original video; and a presentation step of presenting either the summary video or the original video on a screen based on the switching instruction received in the receiving step.
19. A computer program written in a computer-readable format to cause a computer to function as: a reception unit that receives an instruction to switch between a summary video generated based on a summary of each segment generated from visual information and audio information extracted for each segment into which the video is divided and the original video; and a presentation unit that presents either the summary video or the original video on the screen based on the switching instruction received by the reception unit.
Citation Information
Patent Citations
Video abstract generation method and device
CN114078221A
System and method for providing multimedia summaries of video programming
JP2004516753A
Information processing system, information processing method, and computer program
JP2005267278A
Method for processing audio video data, device, electronic apparatus, storage medium, and computer program
JP2022000972A
User interfaces and tools for facilitating interactions with video content
WO2022246450A1