Abstract generation device, abstract model learning device, abstract generation method, abstract model learning method, and program
The system effectively generates summary texts from videos by extracting text from images and audio, using a learned summary model to address the challenge of summarizing videos with both audio and images.
Patent Information
- Application Number
- JP2024504338
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-04
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-04
AI Technical Summary
There is no effective technique for generating summary text from videos that include both audio and images, such as presentation videos, which makes it difficult to quickly grasp the content of long videos.
A system comprising an image processing unit, an audio processing unit, and a summary generation unit that extracts text from images and audio, and uses a learned summary model to generate a summary text from the extracted information.
Enables the generation of accurate summary texts from videos, allowing for quick comprehension of video content without the need to watch the entire video.
Smart Images

Figure 0007683810000001 
Figure 0007683810000002 
Figure 0007683810000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for generating a summary text of a video from the video. [Background technology]
[0002] In recent years, online meetings have become more common, and many videos of presentations from meetings and other events have been made available on the Internet.
[0003] Generally, presentation videos are long, so viewers must watch the video for a long time to understand the content. Therefore, there is a demand for a way to understand the content of presentation videos in a short amount of time. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension Summary of the Invention [Problem to be solved by the invention]
[0005] In order to quickly understand the contents of a presentation video, it may be possible to generate a text that summarizes the presentation video (summary text).
[0006] However, in the prior art, there was no technology that could appropriately generate a summary text from a video that includes audio and images (such as slide images), such as a presentation video.
[0007] The present invention has been made in view of the above points, and aims to provide a technique for appropriately generating a summary text from a video that includes audio and images. [Means for solving the problem]
[0008] According to the disclosed technology, there is provided an image processing unit that receives an image related to a moving image and extracts at least text from the image; an audio processing unit that receives audio from the video and extracts at least text from the audio; a summary generation unit that generates a summary text of the video from information extracted from the image and information extracted from the audio using a trained summary model. And, The summary model is a summary model further trained on a summary model that was previously trained using text in a field related to the video and a correct summary text of the text. Summary Generator is provided. [Effects of the Invention]
[0009] The disclosed technology provides a technology for appropriately generating a summary text from a video that includes audio and images. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram showing the basic process flow for creating a summary text from a presentation video. [Figure 2] 1 is a configuration diagram of a summary generation device 100. FIG. [Figure 3] 1 is a flowchart illustrating the operation of the summary generation device 100. [Figure 4] FIG. 2 is a configuration diagram of a summary model learning device 200. [Figure 5] FIG. 1 illustrates a configuration for pre-training a summary model. [Figure 6] 10 is a flowchart illustrating the operation of the summary model learning device 200. [Figure 7]FIG. 10 is a diagram illustrating an example of input to a summary model and output from the summary model in pre-learning. [Figure 8] FIG. 10 is a diagram for explaining processing for cutting out an image from a moving image. [Figure 9] FIG. 10 is a diagram for explaining extraction of text from an image. [Figure 10] FIG. 1 is a diagram for explaining extraction of text from speech. [Figure 11] FIG. 10 is a diagram illustrating an example of input to a summary model and output from the summary model during learning. [Figure 12] FIG. 10 is a diagram showing the configuration of a data extension unit 400. [Figure 13] 10 is a flowchart for explaining the operation of the data extension unit 400. [Figure 14] FIG. 10 is a diagram illustrating an example of data division. [Figure 15] FIG. 10 is a diagram for explaining learning using a divided learning data set. [Figure 16] FIG. 2 illustrates an example of a hardware configuration of the apparatus. [Figure 17] FIG. 10 is a diagram showing the effect of learning paper data in advance. [Figure 18] FIG. 10 is a diagram illustrating the effect of pre-learning slide summaries. [Figure 19] FIG. 10 is a diagram showing the effect of training the original training data set together with the training data set obtained by division. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, an embodiment of the present invention (the present embodiment) will be described with reference to the drawings. The embodiment described below is merely an example, and the embodiment to which the present invention is applied is not limited to the following embodiment.
[0012] The summary generation device 100 and summary model learning device 200 described below both offer specific improvements over conventional techniques, such as generating summaries from papers, and represent an advancement in the field of video summary generation.
[0013] The data augmentation unit 400 (training data generation device 400) described below provides specific improvements over conventional techniques such as manually generating summaries, and represents an advancement in the field of technology related to techniques for training a summary model for generating summary text of a video.
[0014] In the following, a presentation video is used as the video for which a summary is to be generated, but this is just an example. The technology according to the present invention is not limited to presentation videos and can be applied to videos in general.
[0015] (Outline of the embodiment) In recent years, the number of online conferences has increased, and many videos of presentations from conferences and other events have been made public. Since presentation videos are generally long, there is a demand for a way to understand their content in a short amount of time. To understand the content of a presentation video in a short amount of time, it is desirable to be able to generate a summary of the presentation video.
[0016] Therefore, in this embodiment, a technique for generating a summary text corresponding to a presentation video will be described.
[0017] <Example of a presentation video> As an example, as disclosed in "https: / / slideslive.com / 38928967 / predicting-depression-in-screening-interviews-from-latent-categorization-of-interview-prompts" (accessed February 27, 2022) and "https: / / videolectures.net / " (accessed February 27, 2022), a typical presentation video consists of images of slides containing the presentation content, an image of the presenter, and the presenter's audio. Note that in many cases, the presenter's image is not displayed.
[0018] <Basic process flow for creating summary text from a presentation video> The basic process flow for creating a summary text from a presentation video will be explained with reference to Figure 1. In the following explanation, for the sake of convenience, the presentation video will be referred to as the "video" and the summary text will be referred to as the "summary."
[0019] First, from the video to be summarized, (A) presentation slides, (B) images extracted from the video, and (C) audio are prepared as input data to the summary generating unit 130.
[0020] It is assumed that the presentation slides (A) are separate files from the video. Although a summary can be generated with at least one of the three input data (A), (B), and (C), it is preferable to have three of them, or two of them (A), (B), and (C), or two of them (B) and (C), in order to generate a more accurate summary.
[0021] Next, the input data converted into text by image recognition / voice recognition is input to summary generation unit 130, which outputs the summary text. Summary generation unit 130 is a functional unit included in summary generation device 100, which will be described later.
[0022] <About summary generation technology> In this embodiment, the summary generation unit 130 uses a neural network model (called a summary model) to generate a summary from text.
[0023] Any summarization model can be used as long as it inputs text and outputs summary text. In this embodiment, as an example, a model based on BART disclosed in Non-Patent Document 1 is used.
[0024] BART is a model consisting of an encoder and a decoder. By using a trained model, when text is input to the encoder, a summary text is output from the decoder.
[0025] <About the assignment> Although there have been technologies that input text and output a summary, there are no technologies that output a summary from multimodal input data. In other words, there has been no technology in the past that can appropriately generate a summary text from a video that includes audio and images (slide images, etc.), such as a presentation video.
[0026] If the above-mentioned problems are divided into more specific problems from the viewpoint of the embodiment, they can be divided into the following problems 1 to 3.
[0027] Issue 1: The cost of creating training data containing correct summary text to be used when training a summarization model to generate summaries for videos is high.
[0028] Problem 2: There is no summary generation technology that uses a summary model to extract audio and images from video and output summary text using these as input.
[0029] Issue 3: Even if it were possible to collect correct summary text from an external server, etc., to use when training a summary model for generating summaries for videos, the amount of data would be small, resulting in a small amount of training data and making it impossible to generate an accurate summary model.
[0030] Below, we will explain the configuration and operation of a summary generation device 100 that generates a summary from a presentation video, and a summary model training device 200 that generates (trains) a summary model used in the summary generation device 100. The technology described below solves the above problems 1 to 3.
[0031] (Configuration and Operation of Summary Generation Device 100) Fig. 2 shows a configuration diagram of a summary generation device 100 according to this embodiment. As shown in Fig. 2, the summary generation device 100 includes an image processing unit 110, an audio processing unit 120, a summary generation unit 130, and a summary model DB (database) 140. The summary model DB 140 stores trained summary models. Note that the DB in this specification may also be referred to as a storage unit or a storage unit.
[0032] The flow of operations of the summary generation device 100 shown in FIG. 2 will be described with reference to the flowchart of FIG.
[0033] Audio information and image information are extracted from the video for which a summary is to be created, and in S101, the image information is input to the image processing unit 110, and the audio information is input to the audio processing unit 120. In the example of Fig. 2, it is assumed that the functional unit that extracts audio information and image information (particularly image information) from the video is located outside the summary generation device 100, but the functional unit may also be provided inside the summary generation device 100.
[0034] In S102, the image processing unit 110 extracts text from the image using image recognition technology. In addition to the text, the image processing unit 110 may also extract accompanying auxiliary information (such as the color of the text in the slide).
[0035] In S103, the voice processing unit 120 extracts text from the voice using a voice recognition technique. Note that the order of the processes of S102 and S103 may be reversed, or S102 and S103 may be executed simultaneously.
[0036] The text extracted in S102 and the text extracted in S103 are input to the summary generation unit 130. In S104, the summary generation unit 130 generates a summary from the text extracted in S102 and the text extracted in S103 using a summary model read from the summary model DB 140. As will be explained in the section on learning the summary model, in addition to the text, information that adds one, more than one, or all of character placement features, image features, and audio features may be used as input to the summary model. Note that the actual substance of the "summary model" is data consisting of functions and weight parameters that constitute a neural network. In S104, the summary generation unit 130 outputs the generated summary.
[0037] As described above, by using both audio and image information obtained from a video, a high-quality summary can be generated.
[0038] The processing in the functional units that extract audio information and image information from video, image processing unit 110, and audio processing unit 120, is the same as the processing in learning data input unit 220, image processing unit 230, and audio processing unit 240 of summary model learning device 200, which will be described later, and therefore these processing details will be explained in the explanation of summary model learning device 220.
[0039] The summary generation device 100 of this embodiment solves the above-mentioned problem 2 and realizes a summary generation technology using a summary model that extracts audio and images from video and outputs summary text using these as input. The summary model is trained by a summary model training device 200, which will be described below.
[0040] (Configuration and operation of the summary model learning device) Fig. 4 shows an example of the configuration of summary model learning device 200 in this embodiment. As shown in Fig. 4, summary model learning device 200 includes a data acquisition unit 210, a learning data input unit 220, an image processing unit 230, an audio processing unit 240, a summary model learning unit 250, a data extension unit 400, a model setting unit 270, a summary model DB 280 that stores pre-trained summary models, and a summary model DB 290 that stores summary models currently being trained.
[0041] In this embodiment, when training a summary model, a summary model is created by learning a large amount of abstracts from papers that are thought to have a high similarity in content to the presentation, and then this summary model is fine-tuned using a small amount of summary data from the presentation. This makes it possible to achieve high accuracy even with a small amount of correct summary data for the presentation video.
[0042] Note that performing pre-learning as described above is one method for solving Problem 3. Problem 3 can also be solved by using further training data generated by the data extension unit 400 (described later) without performing pre-learning. Performing pre-learning and using further training data generated by the data extension unit 400 (described later) may be combined.
[0043] 4 shows a configuration for performing the above-mentioned pre-learning, but it is also possible to perform learning using the learning data generated by the data augmentation unit 400 without performing pre-learning. Also, it is possible to perform learning using the learning data generated by the data augmentation unit 400 on a summary model that has undergone pre-learning.
[0044] The configuration for pre-learning is shown in Fig. 5. As shown in Fig. 5, the configuration for pre-learning includes a summary model pre-learning unit 310 and a summary model DB 320 that stores summary models undergoing pre-learning.
[0045] A summary model pre-training device (a device separate from summary model training device 200) may be configured that includes summary model pre-training unit 310 and summary model DB 320, or summary model pre-training unit 310 and summary model DB 320 may be included within summary model training device 200.
[0046] The flow of operations of the summary model training device 200 and the summary model pre-training unit 310 will be described with reference to the flowchart of Fig. 6. Detailed processing will be described later.
[0047] S201 and S202 are processes in the configuration for pre-learning shown in Fig. 5. In S201, pre-learning data is input to the summary model pre-learning unit 310. The pre-learning data is, for example, the text of a paper related to the presentation and a summary of that paper (correct answer data).
[0048] In S202, summary model pre-training unit 310 uses input data to train (pre-train) a summary model. The pre-trained summary model is stored in summary model DB 280 in summary model training device 200.
[0049] Steps S203 to S207 are processes performed by the summary model learning device 200 shown in FIG. 4. In the input process of step S203, access information (e.g., the URL where the paper and presentation video are publicly available) is input to the data acquisition unit 210. The data acquisition unit 210 uses the access information to acquire learning data, for example, from a server on the network, and inputs the data to the learning data input unit 220. The learning data is, for example, a presentation video related to the paper and a correct summary text corresponding to the video. In step S203, the learning data input unit 220 further performs a process of separating the presentation video into image information and audio information, and inputs the image information to the image processing unit 230, the audio information to the audio processing unit 240, and the correct summary to the summary model learning unit 250.
[0050] The image information input by the learning data input unit 220 to the image processing unit 230 may be slide images or the like in a separate file from the presentation video, or slide images or the like extracted from the presentation video. In either case, the images may be referred to as "images related to the video." In either case, text can be extracted from the "images related to the video" by image recognition processing.
[0051] In the following description, it is assumed that the image information input to the image processing unit 230 is a slide image extracted from a presentation video.
[0052] In S204, the image processing unit 230 extracts text from the image using image recognition technology. In addition to the text, the image processing unit 230 may extract accompanying auxiliary information (such as the color of the text in the slide), character layout features, image features, and the like.
[0053] In S205, the speech processing unit 120 extracts text from the speech using speech recognition technology. The speech processing unit 120 may extract speech features in addition to the text. Note that the order of the processes of S204 and S205 may be reversed, or S204 and S205 may be executed simultaneously.
[0054] The text extracted in S204 and the text extracted in S205 are input to the summary model training unit 250. The correct summary is also input to the summary model training unit 250.
[0055] Here, model setting unit 270 reads out a pre-trained summary model from summary model DB 280, and stores the pre-trained summary model in summary model DB 290. Using the parameters in this pre-trained summary model as initial values, the following learning (fine tuning) is performed.
[0056] In S206, the summary model learning unit 250 uses the summary model read from the summary model DB 290 to generate a summary from the text extracted in S204 and the text extracted in S205, and also learns the summary model (updates its parameters) so as to minimize the error between the generated summary and the correct summary.
[0057] When the learning is completed, the summary model learning unit 250 stores the learned summary model in the summary model DB 140 of the summary generation device 100.
[0058] In the above example, pre-learning is performed to fine-tune the pre-trained learning model, but as mentioned above, pre-learning is not essential. Processing may start from S203 in Fig. 6 without performing pre-learning. When pre-learning is not performed, the initial values of the parameters of the summary model may be random values or values other than random values.
[0059] The processing content of each step from S201 to S207 will be explained in more detail below.
[0060] (S201, S202: Pre-learning) A detailed example of pre-learning performed by the summary model pre-learning unit 310 shown in Fig. 5 will now be described. In pre-learning, the summary model is trained using text from a field related to the field of the presentation video to be summarized (referred to as related field text) and its correct summary. Examples of related field text include paper text (text from the main body of a paper) and slide text.
[0061] An example of input to and output from the summary model when a research paper text is used as the related field text is shown in Fig. 7. As described above, the summary model according to this embodiment is a model consisting of an encoder and a decoder.
[0062] As shown in Figure 7, the text of a paper is input to the encoder, and a summary text is output from the decoder. The summary model is trained to minimize the error between the output summary text and the correct summary text. When slide text is used as input, the processing is the same as when paper text is used.
[0063] When text is input to the encoder, the text token sequence is first converted into a fixed-dimension vector, and then converted into summary text through the encoder-decoder.
[0064] An example of input paper text is shown below.
[0065] 「We assume familiarity with basic notions of graph theory (see, for instance, 1]) and with elementary notions of polyhedral combinatorics (see, for instance, 6]).", "Our graphs will be undirected and simple (no loops and no multiple edges).", "As usual, K n denotes the complete graph with n vertices; K n;m denotes the complete bipartite graph with n + m vertices and n m edges.", "Let G be a graph; G is connected if for every pair of distinct vertices there exists a path in G joining them; G is twoconnected if for every vertex v of G, the graph G ?", "v is connected; G is planar if it can be embedded in the plane.", "A subgraph H of a G is spanning if the vertex sets of H and G are the same.", "Subdivision of an edge uv of G consists of removing edge uv, and adding a new vertex w and the two edges uw and vw; w is called subdivision vertex.", "If G and H are two graphs, we say that G contains a subdivision of H, if H arises by subdivision of the edges of some subgraph of G.As usual, (u) denotes the set of all edges that are incident in the vertex u.", "In automatic graph drawing the following problem arises: nd in a complete graph with weights on its edges a two-connected planar spanning subgraph with weight as Partially supported by DFG-Grant JU204 / 7-1 Forschungsschwerpunkt \" E ziente Algorithmen f ur diskrete Probleme und ihre Anw…". An example of the output (or summary text, which is the correct answer data) for the above input is shown below.
[0066] "The problem of finding a two-connected planar spanning subgraph of maximum weight in a complete edge-weighted graph is important in automatic graph drawing.", "We investigate the problem from a polyhedral point of view." On presentation video websites, slide files can sometimes be obtained as separate files from the video. Furthermore, slide files often contain both the slide data itself (slide text) and a summary of the slide (summary text). In such cases, a summary model can be pre-trained by using the slide text as input to the encoder-decoder and the summary text as the correct answer.
[0067] An example of input slide text is shown below.
[0068] 「[["ssn"], ["MASTERS", "IN", "AUTOMOTIVE"], ["ENGINEERING"], ["Karthiek", "Nagaraj"], ["PRESENTED", "AT", "IRIS", ",", "DEPARTMENT", "OF", "MECHANICAL", "ENGINEERING"], ["SSN"], ["WHY", "AUTOMOBILE", "ENGINEERING", "?"], ["Its", "scope", "is", "irrefutable", "and", "job", "prospects", "are", "very", "strong", "in", "any", "part", "of", "the", "world", ".", "Also", "the", "prospect", "of", "returning", "to", "India", "to", "work", "is", "bright", "as", "the", "indian", "automotive", "industry", "is", "making", "tremendous", "progress", "."], [">", "It", "is", "a", "stream", "which", "blends", "passion", "for", "vehicles", "and", "technical", "knowledge", ",", "thus", "making", "it", "all", "the", "more", "interesting", "."], ["It", "is", "an", "interdisciplinary", "field", "which", "encompasses", "mechanical", "engineering", ",", "electrical", "and", "electronics", "engineering", "and", "software", "engineering", ".", "This", "again", "adds", "to", "the", "interest", "factor", "."], ["A", "multitude", "of", "research", "options", "are", "on", "offer", ",", "especially", "in", "hybrid", "powertrains", "and", "fuel", "cells", "."], ["PRESENTED", "AT", "IRIS", ",", "DEPARTMENT", "OF", "MECHANICAL", "ENGINEERING"], ["2"], ["SSN"], ["KEY", "AREAS", "OF", "AUTOMOTIVE", "ENGINEERING"], ["Vehicle", "Propulsion", "~", "Internal", "combustion", "engines"], ["Powertrain", "dynamics", "and", "control"], ["Vehicle", "dynamics", "~", "Handling", "response"], ["~", "Advanced", "transmission"], ["systems"], ["~", "Hybrid", "propulsion", "systems"], ["~", "Terrain", "modelling"], ["~", "Fuel", "cells"], ["~", "Drivetrain", "control", "systems"], ["~", "NVH", "modelling"], ["Automotive", "body", "structures", "~", "Material", "selection"], ["Automotive", "safety", "~", "Active", "and", "passive", "safety"], ["systems"], ["~", "Crash", "worthiness"], ["~", "Human", "factor", "engineering"], ["and",". An example of the output (or slide summary, which is the correct answer data) for the above input is shown below.
[0069] "A guide to Masters in Automotive Engineering at International Destinations" (S203: Input processing of summary model learning device 200) Next, a detailed example of the processing by the data acquisition unit 210 and the processing by the learning data input unit 220 in the summary model learning device 200 shown in FIG. 4 will be described.
[0070] The data acquisition unit 210 accesses, for example, a website for presentation videos on the Internet and acquires the presentation videos and the summaries of the correct answers corresponding to the videos from the website. An example of a website from which such videos and summaries can be acquired is "https: / / aclanthology.org / " (accessed February 27, 2022).
[0071] As described above, by obtaining presentation videos and their summaries from a server on the network, it is possible to create training data without manually creating summaries, thereby solving the aforementioned problem 1.
[0072] The learning data input unit 220 performs processing to separate the presentation video acquired by the data acquisition unit 210 into image information and audio information, and inputs the image information to the image processing unit 230 and the audio information to the audio processing unit 240 .
[0073] Although the image information is not limited to a specific image, it is assumed here that the image information is a slide image in a presentation video.
[0074] An example of processing performed by the learning data input unit 220 to extract an image from a presentation video will be described with reference to FIG.
[0075] S203(1-1): The learning data input unit 220 extracts images from the presentation video every k seconds, where k is a predetermined real number greater than 0. The top row of Fig. 8 shows six images extracted every k seconds.
[0076] S203(1-2): The learning data input unit 220 compares the images extracted in S203(1-1) in order for each time, and if the similarity between the t-th image and the t-1-th image is equal to or greater than a threshold, determines that these images are the same image. Note that any method may be used to determine the similarity between images. Figure 8 shows an example of the similarity between each pair of six images.
[0077] S203(1-3): The learning data input unit 220 repeats S203(1-1) and S203(1-2) to extract a set of different images. In Fig. 8, image 1, image 4, and image 6 are shown as the set of different images when the threshold value is 25. The obtained set of images is input to the image processing unit 230.
[0078] (S204: Image processing) Next, a detailed example of image processing performed by the image processing unit 230 will be described. The image processing unit 230 performs OCR (Optical Character Recognition) processing on the set of different images input from the learning data input unit 220, and acquires information such as text, character color, character size, and character position from each image in the set of different images, as shown in Fig. 9. Note that the acquired information may be only text.
[0079] (S205: Audio processing) Next, a detailed example of the speech processing performed by the speech processing unit 240 will be described. As shown in Fig. 10, the speech processing unit 240 performs speech recognition processing on the speech input from the learning data input unit 220, and obtains text of the speech recognition result.
[0080] (S206: Learning process) Next, a detailed example of the learning process performed by the summary model training unit 250 will be described. The summary model training unit 250 combines the text obtained by the image processing unit 230 and the text obtained by the audio processing unit 240, and inputs the combined text into the summary model. The summary model training unit 250 trains the summary model so as to minimize the error between the summary text output from the summary model and the correct summary text. For input to the summary model, information obtained by adding character placement features, image features, character size and color information, etc. obtained by the image processing unit 230 to the combined text may be used. Information obtained by adding audio features obtained by the audio processing unit 240 to the combined text may also be used.
[0081] The initial state of the summary model is the summary model pre-trained in S202. However, as described above, it is possible not to perform pre-training, so the initial state of the summary model does not have to be the summary model pre-trained in S202. If pre-training is not performed, learning may be performed using further training data generated by the data extension unit 400, which will be described later.
[0082] An example of input to and output from the summary model is shown in Fig. 11. As described above, the summary model according to this embodiment is a model consisting of an encoder and a decoder.
[0083] As shown in Figure 11, the encoder receives the text combined by [SEP], character size, and color information, and the decoder outputs a summary text. The summary model is trained to minimize the error between the output summary text and the correct summary text.
[0084] When inputting text to the encoder, the text token sequence is first converted into a fixed-dimension vector, and then converted into a summary text through the encoder-decoder. Also, character size and color information may not be available in the input.
[0085] The text obtained by the voice processing unit 240 may be called ASR (Automatic Speech Recognition) text, and the text obtained by the image processing unit 240 may be called OCR text.
[0086] An example of ASR text is shown below.
[0087] “So to put in context to put my presentation in the context, I will, I would like to begin with the word decision support or decision-making. And first ask the question who, or what is making decisions and obviously we get two branches here. One is that we have a human decision maker who makes a decision and all of us are decision makers and then we are also talking about the decision systems. So robots.” An example of OCR text is shown below. The example below is an example of text obtained from a slide image disclosed at "http: / / videolectures.net / site / normal_dl / tag=1005123 / icml2015_schmidt_time_framework_01.pdf" (accessed February 26, 2022).
[0088] “Structured sparsity sparsity is widely used in signal processing, machine learning, and statistics (compressive sensing, sparse linear regression, etc.) Examples of sparsity….” Below is an example of the summary text (or its correct answer) that is output when ASR text and OCR text are combined and input into the summary model.
[0089] 「Decision Support is a discipline concerned with human decision making: it aims to provide methods and tools that support, rather than replace, people in making difficult decisions. One of the widely used decision-support approaches relies on decision models, which are developed in the decision process and used to evaluate and analyse decision alternatives. In this lecture, we shall present the method DEX (Decision EXpert), which was heavily influenced by ideas from Artificial Intelligence. DEX is a hierarchical, qualitative, rule-based, multi-criteria modelling method, suitable particularly for solving classification decision problems. DEX combines traditional approaches with those from expert systems and machine learning. DEX is supported by the software called DEXi and has been used in hundreds of real-world decision-making studies. The presentation will be illustrated by recent applications in the areas of electric energy production, food safety and health care.」 (Configuration and Operation of Data Expansion Unit 400) Below, we will explain one technology that solves Problem 3, which is a technology for automatically generating additional training datasets.
[0090] Fig. 12 shows the configuration of the data extension unit 400 in the summary model learning device 200 shown in Fig. 4. As shown in Fig. 12, the data extension unit 400 has a learning data generation unit 410, a key sentence extraction unit 420, and a task information assignment unit 430. Note that the data extension unit 400 may be a functional unit within the summary model learning device 200, or may be a separate device external to the summary model learning device 200. When the data extension unit 400 is within the summary model learning device 200, the summary model learning device 200 may be referred to as the learning data generation device 400. When the data extension unit 400 is a separate device external to the summary model learning device 200, the separate device may be referred to as the learning data generation device 400.
[0091] The operation flow of the data expansion unit 400 (training data generation device 400) shown in Fig. 12 will be described with reference to the flowchart in Fig. 13. In S301, ASR text obtained by speech processing, OCR text obtained by image processing, and the corresponding correct summary text are input to the training data generation unit 410.
[0092] In S302, the data division unit 410 performs a learning data generation process (which may also be called a data division process) on the input data. In S302, a key sentence extraction process is also performed by the key sentence extraction unit 420. Note that the key sentence extraction unit 420 may be included in the learning data generation unit 410.
[0093] In S303, the task information assigning unit 430 assigns task information to the generated training data set, and in S304 outputs the training data set with the task information assigned. The output data is input to the summary model training unit 250 and used for training the summary model. The processing of each of the above steps will be explained in more detail below.
[0094] (S301: Input, S302: Data division) For one presentation video, one set of data, "OCR text, ASR text, and correct summary text," is input to the training data generation unit 410. A data set for training is called a training data set.
[0095] Based on the above input data, the training data generation unit 410 generates the following five training data sets as shown in FIG. 14. Note that (1) is the original training data set. Since each training data set represents a task, the training data set may also be called a task. Note that the following five are examples, and it is sufficient if at least one further training data set is generated in addition to the original training data set. In addition to the following, (6) OCR text, OCR key sentences, and (7) ASR text, ASR key sentences may also be generated.
[0096] (1) OCR text, ASR text, and correct summary text (2) OCR text, correct summary text (3) ASR text and correct answer summary text (4) OCR text, ASR key text (5) ASR text, OCR important text Both the ASR key sentences and the OCR key sentences are examples of pseudo-correct answer information. Both the ASR key sentences and the OCR key sentences are created by the key sentence extraction unit 420. An example of how to create these key sentences will be described below.
[0097] Regarding ASR key sentences, the key sentence extraction unit 420 extracts ASR key sentences by matching the summary text with the ASR text. For example, the key sentence extraction unit 420 extracts, from the ASR text, parts that have a high similarity to the summary text as ASR key sentences.
[0098] Regarding the OCR key sentences, the key sentence extraction unit 420 extracts the OCR key sentences by matching the summary text with the OCR text. For example, the key sentence extraction unit 420 extracts, from the OCR text, a portion that has a high similarity to the summary text as the OCR key sentence.
[0099] Any method can be applied as a matching method for extracting ASR / OCR key sentences, but it is also possible to use the method used to create extractive summarization data, such as the method described in Fine-tune BERT for Extractive Summarization (https: / / arxiv.org / pdf / 1903.10318v2.pdf, retrieved February 27, 2022).
[0100] (S303: Assign task information) The task information assigning unit 430 assigns identification information (which may be called a label) for identifying a task to each training data set generated by the training data generating unit 410. The identification information is a special token. In the examples (1) to (5) above, for example, identification information such as [task0] is assigned as follows:
[0101] (1) [task0] OCR text, ASR text, correct summary text (2) [Task 1] OCR text, correct summary text (3) [Task 2] ASR text, correct answer summary text (4) [Task 3] OCR text, ASR key sentences (5) [task 4] ASR text, OCR important text (S304: Output, (and learning)) Each task (each learning data set) to which identification information has been assigned in S303 is output to the summary model learning unit 250.
[0102] Summary model training unit 250 trains the summary model using each training data set with identification information. The training method for each training data set is the same as the training method in S206 described above. However, here, as shown in FIG. 15, text with the above identification information is used as input to the decoder. FIG. 15 shows an example of training for task (2) of the above five tasks. Such training is performed for each of (1) to (5).
[0103] This allows the amount of training data to be increased, and a highly accurate summary model to be generated.
[0104] (Example of hardware configuration) The summary generation device 100, the summary model learning device 200, and the training data generation device 400 can all be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on the cloud. Hereinafter, the summary generation device 100, the summary model learning device 200, and the training data generation device 400 will be collectively referred to as the "devices."
[0105] That is, the device can be realized by executing a program corresponding to the processing performed by the device using hardware resources such as a CPU and memory built into a computer. The program can be recorded on a computer-readable recording medium (such as a portable memory) and stored or distributed. The program can also be provided via a network such as the Internet or email.
[0106] Fig. 16 is a diagram showing an example of the hardware configuration of the computer. The computer in Fig. 16 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, and the like, all of which are interconnected by a bus BS.
[0107] A program for realizing processing on the computer is provided by a recording medium 1001 such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001, but may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files, data, etc.
[0108] The memory device 1003 reads and stores the program from the auxiliary storage device 1002 when an instruction to start the program is received. The CPU 1004 realizes the functions related to the light touch maintenance device 100 in accordance with the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network, etc. The display device 1006 displays a GUI (Graphical User Interface) or the like according to the program. The input device 1007 is composed of a keyboard, mouse, buttons, a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the results of calculations.
[0109] (Effects of the embodiment) As described above, the technology according to the present embodiment makes it possible to appropriately generate a summary text from a video including audio and images, such as a presentation video. It also makes it possible to automatically generate additional training data for training a summary model that generates a summary text from a video.
[0110] In particular, in this embodiment, the accuracy of the summary model can be improved by performing pre-learning or data expansion (generation of additional learning data by data division).
[0111] Below, we explain the effects based on the experimental results when pre-training was performed and when data division was performed. Below, we use ROUGE-1, ROUGE-2, and ROUGE-L as evaluation indices, which are abbreviated as R1, R2, and RL, respectively.
[0112] Figure 17 shows the effect of pre-training with paper data. For comparison, "ASR+OCR" shows the evaluation results when no paper data was pre-trained. "+Paper Abstracts (300,000)" and "+Paper Abstracts (500,000)" show the evaluation results when 300,000 and 500,000 paper summaries were pre-trained, respectively. As shown in Figure 17, it can be seen that accuracy is improved by pre-training with paper data.
[0113] Figure 18 shows the effect of pre-training slide summaries. For comparison, "ASR+OCR(4096)" shows the evaluation results when slide summaries are not pre-trained. "+slideshare" shows the evaluation results when slide summaries are pre-trained. As shown in Figure 18, it can be seen that accuracy is improved by pre-training slide summaries.
[0114] FIG. 19 shows the effect of training the original training dataset together with the additional training dataset obtained by splitting. For comparison, "ASR+OCR(4096)" shows the evaluation result when only the original training dataset was trained. "ASR+OCR(4096)+extend" shows the evaluation result when the original training dataset was trained together with the additional training dataset obtained by splitting. As shown in FIG. 19, it can be seen that accuracy is improved by training the original training dataset together with the training dataset obtained by splitting.
[0115] (Addendum) The following additional clauses are disclosed in relation to the above-described embodiment. (Additional note 1) Memory and at least one processor coupled to said memory; Including, The processor: inputting an image relating to the video and extracting at least text from the image; Inputting audio from the video and extracting at least text from the audio; Using the trained summarization model, a text summary of the video is generated from the information extracted from the image and the information extracted from the audio. Summary generator. (Additional note 2) Memory and at least one processor coupled to said memory; Including, The processor: inputting an image relating to the video and extracting at least text from the image; Inputting audio from the video and extracting at least text from the audio; A summary model is trained using information extracted from the image, information extracted from the audio, and a correct summary text of the video. Summary model learning device. (Additional note 3) The processor acquires the video and the summary text of the correct answer from a server on a network. 3. The summary model learning device according to claim 2. (Additional note 4) The processor pre-trains the summary model using text in a field related to the video and a correct summary of the text. 3. The summary model learning device according to claim 2. (Additional note 5) The processor generates at least one further training data set from the training data set consisting of the information extracted from the images, the information extracted from the audio, and a ground truth summary text of the video. 3. The summary model learning device according to claim 2. (Additional note 6) 1. A computer-implemented method for generating a summary, comprising: an image processing step for extracting at least text from images relating to the video; an audio processing step for extracting at least text from the audio in the video; a summary generation step of generating a summary text of the video from information extracted from the image and information extracted from the audio using a trained summary model; A method for generating a summary comprising: (Additional note 7) A non-transitory storage medium storing a program executable by a computer to execute the summary generation process in the summary generation device described in appended claim 1. (Additional note 8) A non-transitory storage medium storing a program executable by a computer to execute the summary model learning process in the summary model learning device according to any one of appended claims 2 to 5.
[0116] Although the present embodiment has been described above, the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims. [Explanation of symbols]
[0117] 100 Summary Generator 110 Image processing section 120 Audio Processing Unit 130 Summary generator 140 Summary Model DB 200 Summary Model Learning Device 210 Data Acquisition Unit 220 Learning data input section 230 Image Processing Unit 240 Audio Processing Unit 250 Summary Model Learning Unit 270 Model Setting Section 280 Summary Model DB 290 Summary Model DB 310 Summary Model Pre-training Section 320 Summary Model DB 400 Data Extension 410 Learning data generation unit 420 Important sentence extraction part 430 Task information assignment unit 1000 Drive Device 1001 Recording media 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input Device 1008 Output Device
Claims
1. An image processing unit that inputs an image related to a video and extracts at least text from the image, An audio processing unit that inputs the audio in the video and extracts at least text from the audio, A summary generation device comprising a summary generation unit that generates a summary text of the video from the information extracted from the image and the information extracted from the audio using a pre-trained summary model, The summary model is a summary model that is further trained with respect to a summary model that has been pre-trained using text in the field related to the video and the correct summary text of the text Summary generation device.
2. An image processing unit that inputs an image related to a video and extracts at least text from the image, An audio processing unit that inputs the audio in the video and extracts at least text from the audio, A summary model learning device comprising a summary model learning unit that learns a summary model using the information extracted from the image, the information extracted from the audio, and the correct summary text of the video, The summary model to be learned by the summary model learning unit is a summary model that has been pre-trained using text in the field related to the video and the correct summary text of the text Summary model learning device.
3. An image processing unit that inputs an image related to a video and extracts at least text from the image, An audio processing unit that inputs the audio in the video and extracts at least text from the audio, A summary model learning unit that learns a summary model using a learning dataset consisting of the information extracted from the image, the information extracted from the audio, and the correct summary text of the video, and at least one additional learning dataset generated from the learning dataset Summary model learning device comprising.
4. A data acquisition unit that acquires the video and the correct summary text from a server on a network The summary model learning device according to claim 3, further comprising.
5. A summary generation method executed by a computer, An image processing step of extracting at least text from an image related to a video, An audio processing step of extracting at least text from the audio in the video, A summary generation method comprising: a summary generation step of generating summary text of the video from information extracted from the image and information extracted from the audio using a learned summary model. The summary model is a summary model that is further learned with respect to a summary model that has been pre-trained using text in the field related to the video and the correct summary text of the text. Summary generation method.
6. A summary model learning method executed by a computer, comprising: An image processing step of extracting at least text from an image related to a video; An audio processing step of extracting at least text from audio in the video; A summary model learning method comprising: a summary model learning step of learning a summary model using information extracted from the image, information extracted from the audio, and the correct summary text of the video. The summary model to be learned in the summary model learning step is a summary model that has been pre-trained using text in the field related to the video and the correct summary text of the text. Summary model learning method.
7. A summary model learning method executed by a computer, comprising: An image processing step of extracting at least text from an image related to a video; An audio processing step of extracting at least text from audio in the video; A summary model learning step of learning a summary model using a learning dataset composed of information extracted from the image, information extracted from the audio, and the correct summary text of the video, and at least one additional learning dataset generated from the learning dataset. A summary model learning method comprising the above.
8. A program for causing a computer to function as each part in the summary generation device according to claim 1, or as each part in the summary model learning device according to any one of claims 2 to 4.
Citation Information
Patent Citations
Information acquisition method and device, computer equipment and storage medium
CN112069309A
Presentation analysis device and presentation viewing system
JP2008152605A
Apparatus and method for summary presentation
JP2008252322A
Analysis device, analysis method, and program
JP2018156473A
Learning device, generation device, learning method, generation method, learning program, and generation program
JP2019008742A