program
By generating and saving short videos during distribution, viewers can flexibly re-watch missed sections of video content, overcoming the need for pre-creation, ensuring uninterrupted re-viewing and real-time continuation.
Patent Information
- Application Number
- JP2022009012
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-24
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2042-01-24
AI Technical Summary
Existing systems fail to efficiently re-watch, during the technical problem of providing viewers with the convenience of flexibly re-watching specific sections even while distribution is in progress, without the distributor having to take the effort of creating the distributor having to take the distributor having to take the distributor creating the distribution content as a short video in advance.
The solution is to generate the already distributed portion as a short video through parallel processing during distribution, and a short video generation program P1 that provides a means for viewers to temporarily re-view a short video generation program P2 that generates and saves the distributor as a short video viewing program P2 that generates and saves the distributor as a short video viewing program P2 that generates and distributes the video as a short video that generates and distributes the video as a slideshow video.
This solution allows viewers to flexibly re-watch specific sections of a video distribution in progress, without the distributor having to create the content as a short video in advance, providing uninterrupted re-viewing and real-time viewing within the scheduled distribution time.
Smart Images

Figure 0007794648000001 
Figure 0007794648000002 
Figure 0007794648000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to the generation and viewing of various types of video content, such as slideshow videos, for distributing lessons, lectures, commentaries, etc. over a communication network. [Background technology]
[0002] Conventional slideshow videos are either pre-created slideshows that are played back with audio commentary and distributed in real time, or the situation is recorded and audio is generated to generate video data that is then played back and distributed. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent Publication No. 2012-109820 "Lecture video content processing device and program" Summary of the Invention [Problem to be solved by the invention]
[0004] To provide a means for enabling viewers to temporarily re-watch, during the time that the distribution is in progress, a progress portion that they missed or want to reconfirm when engaging in an activity of transmitting some knowledge or information such as a class, lecture, or commentary through video distribution, and, once the re-viewing of the progress portion is completed, to continue viewing the unviewed portion from the start of the re-viewing. [Means for solving the problem]
[0005] In the video distribution of classes, lectures, and commentaries that is the subject of this invention, the distributor, such as a teacher, lecturer, or commentator, shares and displays a slideshow material screen from their own distribution terminal, and proceeds with the distribution by moving the slideshow pages in accordance with their own speaking pace.However, if a viewer misses part of the progress, the content being distributed is successively divided into multiple short videos, which are generated and saved so that the missed part can be quickly re-viewed while the distribution is in progress.
[0006] This means is realized by a short video generation program P1 that divides the video and audio of a video being distributed according to the semantic division of the distribution content and generates them as multiple short videos, and a short video viewing program P2 that accelerates and plays back the short video of the progress part desired by the viewer.
[0007] In the short video generation program P1, in order to create a short video divided into sections that are easy to understand in light of the progression of the content of the content, the distribution process is divided by identifying the section division positions on the timeline by combining the division positions based on content analysis of the material data with the dividing positions on the timeline that follow the progression of the live commentary.
[0008] In the short video viewing program P2, when a viewer transmits a re-viewing request, the short video of the section requested to be re-viewed is played back at an accelerated speed, and the video and audio being distributed after the re-viewing request is transmitted are buffered, and after the accelerated playback of the re-viewing section is completed, the buffered video and audio are also delivered at an accelerated speed to catch up with the ongoing distribution, thereby providing uninterrupted re-viewing of the already-distributed section and real-time viewing of the currently-distributed section within the scheduled distribution time. [Effects of the Invention]
[0009] According to the present invention, it is possible to provide viewers with the convenience of flexibly re-watching specific sections even while distribution is in progress, without the distributor having to take the effort of creating the distribution content as a short video in advance. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a conceptual diagram of the hardware configuration of the entire system of the present invention. [Figure 2] FIG. 2 is a conceptual diagram of the hardware configuration of a distributor terminal H1 included in FIG. [Figure 3] FIG. 2 is a conceptual diagram of the hardware configuration of a server H3 included in FIG. [Figure 4] FIG. 2 is a conceptual diagram of the hardware configuration of a viewer terminal H4 included in FIG. [Figure 5] FIG. 1 is a conceptual diagram showing three use cases assumed in the present invention. [Figure 6] FIG. 6 is a conceptual diagram of use case 1 included in FIG. 5. [Figure 7] FIG. 6 is a conceptual diagram of Use Case 2 included in FIG. 5. [Figure 8] FIG. 6 is a conceptual diagram of Use Case 3 included in FIG. 5. [Figure 9] 1 is a conceptual diagram of the software configuration of the entire system of the present invention. [Figure 10] FIG. 10 is a conceptual diagram of software processing of the short video generation program P1 included in FIG. 9. [Figure 11] FIG. 10 is a conceptual diagram of software processing of the short video viewing program P2 included in FIG. 9. [Figure 12] FIG. 10 is a conceptual diagram of operation steps for use case 1 and use case 2 in the short video viewing program P2 included in FIG. 9. [Figure 13] FIG. 10 is a conceptual diagram of operation steps for Use Case 3 in the short video viewing program P2 included in FIG. 9. [Figure 14] FIG. 11 is a conceptual diagram of the hardware configuration involved in the execution of distribution block S1 included in short video generation program P1 shown in FIG. [Figure 15] FIG. 11 is a conceptual diagram of the software processing of the material analysis block S2 included in the short video generation program P1 shown in FIG. [Figure 16]FIG. 16 is a conceptual diagram showing the processing contents from data extraction S2.1 to data classification S2.2 included in the document analysis block S2 shown in FIG. [Figure 17] FIG. 16 is a conceptual diagram showing an on-screen area that is the target of processing from data extraction S2.1 to data classification S2.2, which is included in the document analysis block S2 shown in FIG. [Figure 18] FIG. 16 is a conceptual diagram of each page type classified by page type determination S2.3 included in the document analysis block S2 shown in FIG. [Figure 19] FIG. 16 is a process flow diagram of the first step of page type determination S2.3 included in the document analysis block S2 shown in FIG. [Figure 20] FIG. 16 is a process flow diagram of the second step of page type determination S2.3 included in the document analysis block S2 shown in FIG. [Figure 21] FIG. 16 is a process flow diagram of the third step of page type determination S2.3 included in the document analysis block S2 shown in FIG. [Figure 22] FIG. 16 is a conceptual diagram of a document structure tree D3 generated by page type determination S2.3 included in the document analysis block S2 shown in FIG. [Figure 23] FIG. 23 is a conceptual diagram showing the attribute classification of each inter-page connection point derived from the document structure tree D3 shown in FIG. [Figure 24] FIG. 16 is a conceptual diagram of the keyword extraction S2.4 process included in the document analysis block S2 shown in FIG. [Figure 25] FIG. 16 is a conceptual diagram of an on-screen area that is the target of keyword extraction S2.4, included in the document analysis block S2 shown in FIG. 15, and a document keyword stack D5 that is generated. [Figure 26] FIG. 16 is a conceptual diagram of an example in which the relevance of both pages before and after a connection point is high in the processing method of connection point attribute determination S2.5 included in the document analysis block S2 shown in FIG. [Figure 27]FIG. 16 is a conceptual diagram of an example in which the relevance of both pages before and after a connection point is low in the processing method of connection point attribute determination S2.5 included in the document analysis block S2 shown in FIG. [Figure 28] 16 shows an example of semantic continuity determination criteria that are referred to in the process of connection point attribute determination S2.5 included in the document analysis block S2 shown in FIG. [Figure 29] This is a conceptual diagram showing a processing method for determining the semantic continuity of connection points classified as ``undetermined'' on the material structure tree D3 shown in Figure 23, in the processing of connection point attribute determination S2.5 included in the material analysis block S2 shown in Figure 15. [Figure 30] FIG. 11 is a conceptual diagram of the software processing of the audio analysis block S4 included in the short video generation program P1 shown in FIG. [Figure 31] FIG. 31 is a conceptual diagram showing an overview of character string conversion S4.1 included in the speech analysis block S4 shown in FIG. 30. [Figure 32] FIG. 31 is a conceptual diagram showing an outline of voice character string keyword extraction S4.2 included in the voice analysis block S4 shown in FIG. 30. [Figure 33] FIG. 31 is a conceptual diagram showing an overview of page description section delimiter position estimation S4.3 included in the audio analysis block S4 shown in FIG. 30. [Figure 34] FIG. 31 is a conceptual diagram showing an overview of the temporal continuity determination S4.4 included in the voice analysis block S4 shown in FIG. 30. [Figure 35] 31 shows an example of a time continuity determination criterion referenced in time continuity determination S4.4 included in the speech analysis block S4 shown in FIG. [Figure 36] FIG. 31 is a process flow diagram of the initial step (process 4.0) of determining temporal continuity S4.4, which is included in the speech analysis block S4 shown in FIG. [Figure 37] FIG. 31 is a process flow diagram of the first step (process 4.1) of the temporal continuity determination S4.4 included in the speech analysis block S4 shown in FIG. [Figure 38]FIG. 31 is a process flow diagram of the second step (process 4.2) of the temporal continuity determination S4.4 included in the speech analysis block S4 shown in FIG. [Figure 39] FIG. 31 is a process flow diagram of the third step (process 4.3) of the temporal continuity determination S4.4 included in the voice analysis block S4 shown in FIG. [Figure 40] FIG. 31 is a process flow diagram of the fourth step (process 4.4) of the temporal continuity determination S4.4 included in the voice analysis block S4 shown in FIG. [Figure 41] FIG. 31 is a process flow diagram of the fifth step (process 4.5) of the temporal continuity determination S4.4 included in the voice analysis block S4 shown in FIG. [Figure 42] 11 is a determination table referenced in the division position determination block S5 included in the short video generation program P1 shown in FIG. [Figure 43] 11 is a conceptual diagram illustrating the contents and application status of a division position stack D11 generated by a division position determination block S5 included in the short video generation program P1 shown in FIG. [Figure 44] FIG. 11 is a conceptual diagram of the software configuration of division block S6 included in short video generation program P1 shown in FIG. [Figure 45] FIG. 11 is a conceptual diagram of the software configuration of the post-editing block S7 included in the short video generation program P1 shown in FIG. [Figure 46] FIG. 46 is a conceptual diagram of a processing method in material combination S7.4 included in the post-editing block S7 shown in FIG. [Figure 47] FIG. 13 is a conceptual diagram showing the transition of the operation screen for use case 1 in the short video viewing program P2 shown in FIG. [Figure 48] FIG. 13 is a conceptual diagram showing the operation screen transition of use case 2 in the short video viewing program P2 shown in FIG. [Figure 49] FIG. 13 is a conceptual diagram of the function of Use Case 1 "I want to re-watch part of it" in the short video viewing program P2 shown in FIG. [Figure 50] FIG. 13 is a conceptual diagram of the function of Use Case 2 "I want to rewatch from the beginning" in the short video viewing program P2 shown in FIG. [Figure 51] This is a conceptual diagram showing the screen transition when "Partial Rewatch" is selected from the operations of Use Case 3 "Want to rewatch after distribution" in the short video viewing program P2 shown in Figure 13. [Figure 52] This is a conceptual diagram showing the screen transition when "Rewatch the entire video" is selected from the operations of Use Case 3 "I want to rewatch after distribution" in the short video viewing program P2 shown in Figure 13. DETAILED DESCRIPTION OF THE INVENTION
[0011] The present invention is composed of a short video generation program P1 that, when distributing one piece of video content, sequentially generates the already distributed portion as a short video through parallel processing during distribution, and a short video viewing program P2 that provides a function that allows the viewer to temporarily view the short video of the already distributed portion during distribution.
[0012] The short video generation program P1 is conceptually composed of a character string analysis means, an image analysis means, a semantic analysis means, a progress monitoring means, a segment generation means, a supplementary editing means, an audio analysis means, and a sequential storage means.
[0013] As will be explained in the following embodiments, the character string analysis means is a means for extracting keywords from the character data contained in each page of a document slideshow, and in the case of videos other than slideshow format, it is a means for extracting keywords from character data obtained from general video and audio by the image analysis means and audio analysis means.
[0014] As will be explained in the following embodiments, the image analysis means is a means for extracting keywords from the image and video data contained on each page of a document slideshow, and for videos other than slideshow format, it is a means for extracting keywords from general video content by image recognition.
[0015] As described in the following embodiments, the semantic analysis means is a means for calculating the semantic continuity between two consecutive pages in a slideshow, and for videos other than slideshow format, is a means for identifying semantic divisions based on video and audio.
[0016] As described in the embodiments below, the progress monitoring means is a means for identifying the position on the distribution timeline of the explanatory audio section for each page of the slideshow, and in the case of videos other than slideshow format, is a means for identifying the position on the distribution timeline of the semantic division identified based on the video and audio being distributed.
[0017] As described in the embodiment below, the segment generation means is a means for determining the explanatory sections that should be contained consecutively within each short video based on a general assumption of a viewing time length that is easy for viewers to handle, and for dividing the video data.
[0018] The supplementary editing means, as will be described in the following embodiments, is a means for completing the divided video data of a distribution by adding beginning video data to the beginning and ending video data to the end, thereby creating a single short video.
[0019] As described in the following embodiments, the audio analysis means is a means for generating text data from the audio being streamed and comparing it with the semantic summary data of each page of the slideshow. In the case of videos other than slideshow format, the audio analysis means is a means for comparing the semantic categories identified from the text data generated from the audio being streamed with the semantic categories identified from the analysis results of the video being streamed.
[0020] The sequential storage means, as will be described in the following embodiments, is a means for sequentially storing video and audio data being distributed as material for a short-side moving image.
[0021] The short video viewing program P2 is conceptually composed of a viewer request receiving means, a requested section identification means, an accelerated playback time calculation means, a requested section accelerated playback means, a tracking section accelerated playback means, a viewer interaction means, and an exception handling means.
[0022] The viewer request receiving means, as will be described in the following embodiments, is a means for receiving a request from a viewer to re-view a part of a portion that has already been distributed.
[0023] The requested section specifying means, as will be described in the following embodiments, is a means for allowing the viewer to select a section they wish to rewatch from a list of short videos.
[0024] The accelerated playback time calculation means, as will be described in the following embodiment, is a means for calculating the total time required for accelerated playback of the section requested for re-viewing, and the time required for accelerated playback thereafter to track and synchronize with the running being distributed.
[0025] As described in the following embodiment, the requested section accelerated playback means is a means for individually delivering accelerated playback of a specific delivered section that has been requested to be played back to the viewer terminal that issued the request, at an acceleration rate set arbitrarily by the distributor.
[0026] As described in the following embodiment, the tracking section accelerated playback means is a means for tracking a distribution by continuing accelerated playback of a section currently being distributed that remains unviewed while responding to the accelerated playback of the requested section after the accelerated playback of the requested section is completed, until the section is synchronized with the ongoing normal speed distribution.
[0027] The viewer interaction means, as will be explained in the following embodiment, is a means for displaying a menu on the screen. It is a means for identifying a section that a viewer requests to re-watch.
[0028] The exception processing means is a means for making a decision to play back the replay request section and the tracking section at actual speed when it is calculated that the end point of accelerated playback of the replay request section and the tracking section will be after the end point of distribution, as described in the following embodiment.
[0029] A typical form for implementing this invention is to process a single distribution of content into a set of re-viewable content consisting of multiple short videos by playing a pre-prepared slideshow of materials during classes, lectures, commentaries, etc. that are distributed over a network, while recording the video and audio of the distribution situation, along with audio explanations, and then sequentially dividing the content into multiple short videos.
[0030] In this case, before distribution begins, keywords are extracted from the character string data contained on each page of the document slideshow, as well as from the character string data obtained by image recognition processing of the image and video data contained on each page, by matching them with a pre-prepared terminology dictionary. This goes through a process of generating semantic summary data for each page, and then the semantic summary data of each of two consecutive pages in the document slideshow is compared to calculate the semantic relevance of the two pages as a constant, comparable numerical value.
[0031] When distribution begins, the text data generated from the audio being distributed is compared with the semantic summary data for each page of the slide show to determine the position on the distribution time axis of the explanatory section for each page of the slide show, i.e., the timestamps for the start and end positions of that section.
[0032] Then, based on general assumptions about the viewing time length that viewers find psychologically or conveniently manageable in video viewing experiences such as classes, lectures, and commentaries, the system determines the page sections that should be included consecutively within each short video in light of the length of each page explanation section on the distribution timeline, relative to the viewing time standard for one short video that the distributor defines arbitrarily.Then, based on the combination determination information for these page sections, the system adds video and audio from the beginning and end of the video in a certain format to the video and audio of the divided distribution, completing it into a single short video.
[0033] Furthermore, during the time when the distribution is in progress, viewers can select a video they wish to watch according to their own purpose from among 1 to N short videos that have already been generated sequentially over the course of time, and can quickly watch the video by accelerating playback at an acceleration rate set by the distributor, and then continue to watch the portion of the video that is currently being distributed.
[0034] As shown in FIG. 1, the hardware of the present invention has a configuration in which a distributor terminal H1 is connected to a server H3 and a viewer terminal H4 via a network H2.
[0035] As shown in FIG. 2, within the distributor terminal H1, a processor H1.1, a voice input device H1.2, a display device H1.3, a storage device H1.4, and a communication device H1.5 are connected to one another via a bus H1.6.
[0036] As shown in FIG. 3, within the server H3, a processor H3.1, a storage device H3.2, and a communication device H3.3 are connected to one another via a bus H3.4.
[0037] As shown in FIG. 4, within the viewer terminal H4, a processor H4.1, an audio output device H4.2, a display device H4.3, a storage device H4.4, and a communication device H4.5 are connected to one another via a bus H4.6.
[0038] Based on the above hardware configuration, the present invention implements three use cases, listed in FIG. 5 and described below.
[0039] Use case 1: As shown in Figure 6, a viewer wants to quickly rewatch a specific part of a stream, perhaps to refresh their memory, and then continue watching the ongoing stream.
[0040] Use Case 2: As shown in Figure 7, a viewer starts watching after the start of a broadcast and wants to quickly watch the parts they missed from the beginning to the start of viewing, and then continue watching the ongoing broadcast.
[0041] Use Case 3: As shown in Figure 8, after the distribution has ended, you want to rewatch the entire video or a specific part of it.
[0042] Server H3 of this system (computer) above of As shown in Figure 9, the software is divided into two parts: a short video generation program P1 and a short video viewing program P2.
[0043] As shown in Figure 10, the short video generation program P1 is composed of processing blocks: distribution block S1, material analysis block S2, buffering block S3, audio analysis block S4, division position determination block S5, division block S6, post-editing block S7, and recording block S8.
[0044] As shown in FIG. 11, the short video viewing program P2 is composed of processing blocks including a viewer request receiving block S9, a requested section identification block S10, an accelerated playback time calculation block S11, a requested section accelerated playback block S12, and a tracking section playback block S13.
[0045] The short video viewing program P2 executes individual processing steps according to the three use cases described above, as shown in Figures 12 and 13.
[0046] The short video generation program P1 provides, via the material analysis block S2, a character string analysis means for extracting keywords from character string data contained in material data used for videos in an online distribution service for video content consisting of video and audio including text, and from character string data generated by image recognition from image and video data, and a semantic analysis means for detecting semantic divisions that exist in the overall flow of video material data by verifying the locations and frequency of occurrence of keywords in the material data; a distribution block S1, a buffering block S3, and an audio analysis block S4 for detecting the positions on the video distribution time axis where each keyword that constitutes the detected semantic division appears in the video distribution process, thereby providing a distribution progress monitoring means for detecting distribution time division information consisting of the division start position, division end position, and division distribution duration based on these that each semantic division has on the video distribution time axis; a division position determination block S5 and a division block S6 for multiplying the semantic division information on the video material data by the distribution time division information at the time of video distribution, and calculating the optimum start position information and end position information for each division based on the efficiency of transmission of semantic content and ease of handling in terms of time; A video segment generating means is provided for dividing the entire video into a plurality of video segments, and a post-editing block S7 adds segment title video data generated by processing a part of the distributed video data, reminder video data summarizing the contents of the immediately preceding video segment, and segment ending video data to each of the generated plurality of video segments, thereby providing a supplementary editing means for processing each video segment into an independent but serialized short video work.
[0047] The distributor prepares slideshow-type material data D1 in advance, describing the content to be distributed (lectures, speeches, explanations, etc.), and uploads it to server H3 via network H2. Material data D1 is then stored in storage device H3.2 within server H3.
[0048] When the distributor activates distribution block S1 of short side video generation program P1 at the distribution start time, as shown in Figure 14, the material data D1 stored in storage device H3.2 in server H3 is distributed to one or more viewer terminals H4 via network H2 using either unidirectional streaming communication or bidirectional conference communication.
[0049] At the same time, the processes of the data analysis block S2, buffering block S3, and audio analysis block S4 are all executed in parallel as described below.
[0050] In the document analysis block S2, as shown in Figure 15, the document data D1 is processed in the following steps in each of the processes: data extraction S2.1, data classification S2.2, page type determination S2.3, keyword extraction S2.4, and connection point attribute determination S2.5, to generate a semantic continuity stack D6.
[0051] First, from the start of distribution and in parallel with the distribution process, the document data D1 stored in storage device H3.2 in server H3 is scanned sequentially, starting from the first page, and as shown in Figure 16, all data objects on each page are classified into character string data E1.1 or graphic / image / video data E1.2. The character string data E1.1 is left as is, while the graphic / image / video data E1.2 is converted into character string data representing the image content by performing image recognition processing. All of these are then grouped together by page and stored in data stack D2.
[0052] At this time, when detecting and classifying data objects, the screen of each page is As shown in Figure 17, the document is divided into three areas: Header E2.1, Body E2.2, and Footer E2.3. In the processing of the document analysis block, analysis is performed by detecting and classifying data objects in the Body E2.2 area.
[0053] By comparing the order in which each page appears with the contents of the data objects detected, classified, and listed within the body E2.2 (Body) area of each page, and going through the page type determination S2.3, Step 1, Step 2, and Step 3 described in the following section, each page is classified into the following four page types, as shown in Figure 18, and page type data is assigned to each. Page Type A: Cover Page Type B: Heading (Index) Page Type C: Part Cover Page Type D: Content
[0054] In the first step of the page type classification process (S2.3), as shown in FIG. 19, it is first verified whether the page in question is at the beginning of the data. If it is at the beginning, it is determined to be page type A: cover, and the page type data is assigned.
[0055] Page type determination S2.3, second step: All pages that are not the first page proceed to the page type determination process shown in Figure 20, where it is verified whether there is a list-formatted string on the page.If there is, it is determined to be page type B: heading (Index), and the page type data is assigned.
[0056] Page type determination S2.3, third step: All pages that do not fall into page type A or page type B proceed to page type determination S2.3, third step shown in Figure 21, where it is first verified whether a page of page type B: Index exists before the page in question, and if so, it is verified whether the same string as the item contained in the list-type string on the previous page type B: Index page is placed at the beginning of the first character block in body E2.2, and if so, it is determined to be page type C: Part Cover, and the page type data is assigned.
[0057] On the other hand, if there is no page of page type B: index before the page, the page is determined to be of page type D: content, and the page type data is assigned.
[0058] Next, as shown in the example of FIG. 22, a document structure tree D3 is generated based on these page types.
[0059] All slideshow materials display content from the first page to the last page by page breaks, and in a video format where audio accompanies the screen transitions, in order to divide the entire video into a collection of short videos, the video must be separated somewhere at the page break positions between adjacent pages in the slideshow.In this invention, these page break positions are called "junction points," and each junction point is assigned two types of "junction point attributes," "continuous" or "split," which serve as basic information for determining where it is appropriate to divide the video.
[0060] At this stage, based on the determined page type of each page, the connection point attributes of some connection points within the slide show are automatically determined.
[0061] For example, as shown in Figure 23, the content page PG02 located immediately after the cover page PG01 is usually presented and read aloud in conjunction with the cover page, so both of these pages should be included in the same short video, and the connection point between the cover page PG01 and the content page PG02 is determined to have the attribute of "continuous."
[0062] On the other hand, with regard to the connection point between the content page PG02 and the index page PG03, since it is appropriate for the index page to be located at the beginning of a short video, it is necessary to split the video immediately after the content page PG02 so that the index page PG03 is at the beginning of the next short video, and therefore it is determined that the connection point between the content page PG02 and the index page PG03 should have the attribute of "split".
[0063] However, in this example, after the index pages such as PG03, PG08, and PG13 there are part cover pages or content pages, and since it is not possible to determine at this stage whether the connection points between these pages should have the attribute of "continuous" or "split," the provisional attribute "undetermined" is temporarily assigned to these connection points.
[0064] Following the above process, the document analysis block, as shown in Figure 24, compares the data of each page stored in the data stack D2 with the terminology dictionary D4 equipped for each field of distribution content, extracts keywords present on each page, and stores them in the document keyword stack D5.
[0065] As shown in the example of Figure 25, the document keyword stack D5 stores the keywords that appear on each page in association with the number of times they appear on that page.At this time, for each keyword, the total number of times they appear in multiple character data objects on the page is stored.
[0066] After generating a document keyword stack D5 for each page for all pages, the document keyword stacks D5 of two adjacent pages are compared and collated to determine the "semantic continuity" between the two pages.
[0067] As shown in Figures 26 and 27, for example, when PAGE_A and PAGE_B, which are located consecutively, are compared and collated in the document keyword stacks D5 of both pages, common keywords present on both pages are detected as follows: [KEY_01]: Appears twice on PAGE_A: Appears once on PAGE_B [KEY_02]: Appears 4 times on PAGE_A: Appears 1 time on PAGE_B [KEY_04]: Appears 3 times on PAGE_A: Appears 1 time on PAGE_B
[0068] In this regard, the number of times the detected common keywords appear on both pages is calculated as follows, and the "continuity" of both pages is quantified. (Number of times [KEY_01] appears on PAGE_A x number of times [KEY_01] appears on PAGE_B) + (Number of times [KEY_02] appears on PAGE_A × Number of times [KEY_02] appears on PAGE_B) + (Number of times [KEY_04] appears on PAGE_A × Number of times [KEY_04] appears on PAGE_B) =Continuity between PAGE_A and PAGE_B
[0069] In this example, for three pages, PAGE_A, PAGE_B, and PAGE_C, which are positioned consecutively within the same slideshow, the continuity between PAGE_A and PAGE_B is calculated in Figure 26, and the continuity between PAGE_B and PAGE_C is calculated in Figure 27. Considering these two figures together, we can see that the continuity between PAGE_A and PAGE_B is calculated to be 9, while the continuity between PAGE_B and PAGE_C is calculated to be 1, which means that the continuity between PAGE_A and PAGE_B is higher than the continuity between PAGE_B and PAGE_C.
[0070] In this way, once the "continuity" has been quantified, the semantic continuity of the preceding and following pages in a consecutive positional relationship is classified into three levels: "low," "medium," and "high," in accordance with the standards set by the provider based on values optimized for the type of material content and writing style, as shown in the example in Figure 28.
[0071] Then, as shown in Figure 29, semantic continuity is determined for all connection points whose attributes were provisionally determined to be "undetermined" in the process shown in Figure 23, and the assignment of semantic continuity to all connection points in the slideshow material is completed.The results are stored in the semantic continuity stack D6, as shown in Figure 15.
[0072] In the distribution block S1, when the distributor performs live distribution by speaking while displaying the slide show, the following processing is simultaneously executed in the voice analysis block S4 as shown in FIG.
[0073] String conversion S4.1: As shown in FIG. 31, the spoken voice is converted into a string, and the string is stored in the voice string stack D7 in a form linked with elapsed time information.
[0074] Spoken string keyword extraction S4.2: As shown in FIG. 32, the spoken string stack D7 is compared with the term dictionary D4 to extract keywords present in the spoken string and store them in a spoken string keyword stack D8.
[0075] Estimation of page explanation section boundary positions S4.3: As shown in Figure 33, by comparing the distribution of keywords extracted from the audio and stored in the audio string keyword stack D8 and the sentence boundary positions detected by language recognition with the document keyword stack D5 obtained in the document analysis block S2, the boundary positions on the time axis of the explanatory audio for each page content are estimated and stored in the page boundary position stack D9.
[0076] Temporal continuity determination S4.4: Furthermore, as shown in Figure 34, the time width of each page description section calculated from the difference between each page break position timestamp stored in the page break position stack D9 is compared with the temporal continuity determination criteria set by the distributor, optimized for the content of the document and the description format, etc., and the temporal continuity between description sections on previous and subsequent pages that are in a consecutive position relationship within the document is determined on a three-level scale of "high," "medium," or "low," and stored in the temporal continuity stack D10.
[0077] The time continuity criteria are defined in the following format as shown in FIG. Criteria example) Regarding the length of one explanation section: Anything over 18 minutes (18m00s~) is not permitted. 15 minutes or more but less than 18 minutes (15m00s-17m59s) is classified as exceptionally long. The standard classification is 10 minutes or more but less than 15 minutes (10m00s-14m59s). Exceptionally short distances are those between 7 minutes and 10 minutes (7m00s-9m59s). A time of less than 7 minutes (0m01s - 6m59s) is not permitted. However, the specific numerical values included in the above-mentioned standard examples are determined to be optimal values by the distributor based on the content of the distribution.
[0078] On the premise that this criterion for determining temporal continuity is applied, the following temporal continuity determination S4.4, initial step (process 4.0) and subsequent steps are executed.
[0079] Temporal continuity determination S4.4, initial step (processing 4.0): As shown in Figure 36, the length of the current page description section is determined to be one of the following, and the process proceeds to the next processing step: If it is less than 7 minutes (0m01s to 6m59s), proceed to the first step (process 4.1). If the time is between 7 minutes and 10 minutes (7m00s to 9m59s), proceed to the second step (process 4.2). If the time is between 10 minutes and 15 minutes (10m00s to 14m59s), proceed to the third step (process 4.3). If the time is between 15 minutes and 18 minutes (15m00s to 17m59s), proceed to the fourth step (process 4.4). If it is 18 minutes or longer (18m00s~), proceed to the fifth step (process 4.5).
[0080] In the first step (process 4.1), as shown in Figure 37, the following is repeated: Step 4.1.0: Determine whether the total time falls into one of the following categories when adding the time of the next page description section, and perform the following process accordingly. Process 4.1.1: If the duration is less than 7 minutes (0m01s to 6m59s), assign a "high" attribute to the temporal continuity with the next page and return to process 4.1.0. Process 4.1.2: If the duration is between 7 minutes and 10 minutes (7m00s - 9m59s), assign a "high" attribute to the temporal continuity with the next page and return to process 4.1.0. Process 4.1.3) If the duration is between 10 and 15 minutes (10m00s to 14m59s), assign a "high" attribute to the temporal continuity with the next page and return to process 4.1.0. Process 4.1.4) If the duration is between 15 minutes and 18 minutes (15m00s - 17m59s), assign a "medium" attribute to the temporal continuity with the next page and return to process 4.1.0. Process 4.1.5) If the duration is 18 minutes or longer (18m00s~), assign a "low" attribute to the temporal continuity with the next page and return to process 4.1.0.
[0081] In the second step (step 4.2), as shown in Figure 38, the following is repeated: Step 4.2.0: Determine whether the total time falls into one of the following categories when adding the time of the next page description section, and perform the following process accordingly. Process 4.2.1: If the duration is between 7 minutes and 10 minutes (7m00s - 9m59s), assign a "high" attribute to the temporal continuity with the next page and return to process 4.2.0. Process 4.2.2: If the time is between 10 and 15 minutes (10m00s-14m59s), assign a "high" attribute to the temporal continuity with the next page and return to process 4.2.0. Process 4.2.3: If the duration is between 15 and 18 minutes (15m00s-17m59s), assign a "medium" attribute to the temporal continuity with the next page and return to process 4.2.0. Process 4.2.4: If the duration is 18 minutes or longer (18m00s~), assign a "low" attribute to the temporal continuity with the next page and return to process 4.2.0.
[0082] In the third step (Process 4.3), as shown in Figure 39, the following is repeated: Process 4.3.0: When adding the time of the next page description section, determine which of the following the total time falls under, and perform the following process accordingly. Process 4.3.1: If the time is between 10 and 15 minutes (10m00s-14m59s), assign a "high" attribute to the temporal continuity with the next page and return to process 4.3.0. Process 4.3.2: If the duration is between 15 and 18 minutes (15m00s-17m59s), assign a "medium" attribute to the temporal continuity with the next page and return to process 4.3.0. Process 4.3.3: If the duration is 18 minutes or longer (18m00s~), assign a "low" attribute to the temporal continuity with the next page and return to process 4.3.0.
[0083] In the fourth step (process 4.4), as shown in FIG. 40, the following is repeated: Process 4.4.0: When the time of the next page description section is added, it is determined whether the total time falls into one of the following cases, and the following process is carried out accordingly. Process 4.4.1: If the duration is between 15 and 18 minutes (15m00s-17m59s), assign a "medium" attribute to the temporal continuity with the next page and return to process 4.4.0. Process 4.4.2: If the duration is 18 minutes or longer (18m00s~), assign a "low" attribute to the temporal continuity with the next page and return to process 4.4.0.
[0084] In the fifth step (step 4.5), as shown in Figure 41, the following is done: Process 4.5.0: The audio string stack D7 is linguistically analyzed, and the position of the speech break (end of sentence) just before the 18-minute mark in the current page description section is determined as the "additional page break position." Process 4.5.1: For the page description section that ends at the "additional page break position," assign a "low" attribute to the temporal continuity with the next page. Process 4.5.2: For the page description section beginning with the "Additional Page Break Point", return to the initial step (Process 4.0).
[0085] The division position determination block S5 multiplies the semantic continuity stack D6 generated in the material analysis block S2 with the temporal continuity stack D10 generated in the speech analysis block S4 to generate a division position stack D11 that determines the continuity or division attribute for each connection point.
[0086] The criteria for determining the division position are as follows, as shown in Figure 42: - If either "semantic continuity" or "temporal continuity" is "low", the segment is split. - Split if both "Semantic continuity" and "Temporal continuity" are "Medium". If one of "semantic continuity" and "temporal continuity" is "high" and the other is "medium," then the two are contiguous.
[0087] As shown in Figure 43, the data format of the split position stack D11 is as follows: Connection point number / Connection point timestamp / Continuity attribute 0 or division attribute 1 Connection point number / Connection point timestamp / Continuity attribute 0 or division attribute 1 Connection point number / Connection point timestamp / Continuity attribute 0 or division attribute 1 Note: The above data string continues for the number of connection points.
[0088] An example of division processing based on the division position stack D11 in this data format is shown in FIG. 43, and in this example, the entire video is divided into eight short videos.
[0089] In the division block S6, as shown in FIG. 44, the distributed video C3 stored in the buffering block S3 is divided at the timestamp position to which the division attribute 1 is assigned on the division position stack D11, and the generated N video segments 1 to N (C4.1 to C4.N) are sent to the post-editing block S7.
[0090] In the post-editing block S7, the following processing is carried out as shown in FIG.
[0091] Short video title creation S7.1: The video segments recorded in the buffering block S3 and divided and generated in the division block S6 are cut out from the document title page video at the beginning of the first video segment 1 (C4.1) at a certain number of seconds (e.g., 3 to 5 seconds), and a short video number string (e.g., "Vol. 1") is superimposed on the cutout, and the result is saved as short video titles 1 to N (C5.1 to C5.N).
[0092] Extraction of page change time S7.2: N video segments 1 to N (C4.1 to C4.N) are scanned, and timestamps of the times when page changes occur are extracted by image recognition and recorded in a page change time stack D12.
[0093] Reminder video generation S7.3: From each video data, a reminder video (C6.1 to C6.N) is generated that displays a still image of each page scene for only a certain number of seconds (e.g., 3 to 5 seconds) according to the timestamp of the page switching time stack D12.
[0094] Thereafter, as shown in FIG. 46, the following material combination S7.4 is performed, which is optimized for each of the moving image section 1 (C4.1) and the moving image section N (C4.N).
[0095] Material combination S7.4, for video category 1 (C4.1): Background music C7, which the distributor provides at their discretion, is added to short video title 1 (C5.1), and video 1 (C4.1) is then concatenated at the end. A guide string containing the next short video number 2 (for example, "Next is Vol. 2") is then superimposed on the short video title 1 (C5.1) screen at the end, and the resulting video with background music C7 added is concatenated, and saved in storage device H3.2 within server H3 as the completed version of short video 1 (C8.1).
[0096] Material combination S7.4, for video segments 2 to N (C4.2 to C4.N): Background music C7 prepared by the distributor is added to the short video titles 2 to N (C5.2 to C5.N), and then the reminder video 1 to N-1 (C6.1 to C6.N-1) of the previous video is linked to the end, and background music C7 prepared by the distributor is added to that, and then video segments 2 to N (C4.2 to C4.N) are linked to the end, and then a guide string including the next short video number 3 to N (for example, "Next is Vol. 3") is superimposed on the short video titles 2 to N (C5.2 to C5.N) screen behind them, and background music C7 is added to the linked pieces, and the completed version of short videos 2 to N (C8.2 to C8.N) is saved in storage device H3.2 in server H3.
[0097] As shown in Figure 11, when a request to re-view a specific portion of the video that has already been distributed is sent from any viewer terminal, the short video viewing program P2 provides a viewer request receiving block S9 with a re-viewing request receiving means for receiving the request, a requested section identification block S10 with a requested section identification means for identifying the requested portion to be re-viewed, an accelerated playback time calculation block S11 with an accelerated playback time calculation means for calculating the total time required to accelerate playback of the re-viewing requested section at an acceleration rate arbitrarily set by the distributor, and then continue to accelerate playback of the unviewed section until the video catches up with the distribution speed at the same acceleration, a requested section accelerated playback block S12 with a requested section accelerated playback means for individually accelerating playback of the re-viewing requested section for the viewer terminal, and after accelerated playback of the re-viewing requested section is completed, a tracking section accelerated playback block S13 with a tracking section accelerated playback means for accelerating playback of the unviewed section from the completion of accelerated playback of the re-viewing requested section until the video catches up with the distribution speed.
[0098] For the three use cases shown in Figure 5, viewers use short videos 1 to N (C8.1 to C8.N) stored in storage device H3.2 in server H3, following the processing steps shown in Figure 12 for use cases 1 and 2, and the processing steps shown in Figure 13 for use case 3, as described below.
[0099] Processing steps for Use Case 1: If a viewer wants to rewatch part of a portion that has already been distributed, as shown in Figure 47, the viewer first presses the "Rewatch" button on the screen of this system to display the first operation menu M1, and then presses the "Partial Rewatch" button there to send a request.
[0100] A request to "partially rewatch" can be made when at least one short video has already been generated, and the request can remain available until the end of distribution.
[0101] When a "partial replay" request is issued, the following occurs:
[0102] The screen displays a list of short videos 1 to N (C8.1 to C8.N) stored in storage device H3.2 in server H3, along with their respective time information. When the viewer selects a short video 1 to N for the part they wish to watch and presses the "Start Watching" button to start watching, the following process is executed, as shown in Figure 49.
[0103] First, the first short video N (C8.N) on the timeline among the short videos selected by the viewer is played at an acceleration rate of X, which is set by the streamer. However, the spoken audio is naturalized, and the tone of the spoken audio is heard as if it were spoken at X times the normal speed, while maintaining a tone close to the speaker's own voice (using existing audio filter technology).
[0104] After the accelerated playback of the first short video N (C8.N) selected for replay is completed, the second short video N+1 (C8.N+1) is accelerated, followed by the third short video N+2 (C8.N+2), and so on, with all selected short videos (1 to N) being accelerated continuously along the timeline.
[0105] Once accelerated playback of all selected short videos (1 to N) has been completed, the accelerated playback of the unwatched parts since the first time the "Rewatch" button was pressed and a rewatch request was sent will continue, and once the playback catches up with the running position of the ongoing broadcast, normal broadcasting will resume at actual speed.
[0106] Processing steps for Use Case 2: If a viewer wants to rewatch a portion that has already been distributed from the beginning, as shown in Figure 48, the viewer presses the "Rewatch" button on the screen of this system to display the first operation menu M1, and then presses the "Rewatch from the beginning" button there to send the request.The following processing will be executed, as shown in Figure 50.
[0107] When a viewer requests to "rewatch from the beginning," the short video viewing program P2 reads the distributed video C3 stored in the buffering block S3 of the short video generation program P1, and accelerates playback from the beginning at an acceleration rate of X set by the distributor. During this process, the spoken voice is naturalized, and the tone of the spoken voice is heard as if it were spoken at X times the normal speed, while maintaining a tone close to the speaker's actual voice (using existing voice filter technology).
[0108] After the accelerated playback from the beginning reaches the point when the "replay" button is pressed for the first time to transmit a replay request, the accelerated playback continues for the part that has not been viewed since the request was transmitted, and when it catches up with the running position of the ongoing distribution, it returns to normal distribution at the actual speed.
[0109] Processing steps for Use Case 3: If a viewer wants to rewatch after the broadcast has ended, as shown in Figure 51, the viewer presses the "Rewatch" button on the screen of this system, and a second operation menu M2 appears with the options to "Rewatch Part" or "Rewatch the Entire Program."
[0110] When "Partial Replay" is selected on the second operation menu M2, a list of short videos 1 to N (C8.1 to C8.N) stored in storage device H3.2 in server H3 is displayed on the screen, along with their respective time information. The viewer selects the short video 1 to N for the part they wish to watch and presses the "Replay at normal speed" or "Replay at accelerated speed" button.
[0111] When you press the "Rewatch at Actual Speed" button, the selected short videos (1 to N) will be played at actual speed in order of oldest to newest distribution time.
[0112] When you press the "Accelerate Rewatch" button, the selected short videos 1 to N will be played at an accelerated speed of X, set by the streamer, in order of oldest to newest.
[0113] Also, in use case 3, as shown in FIG. 52, when "Replay the entire program" is selected on the second operation menu M2, buttons for "Replay at normal speed" and "Replay at accelerated speed" are displayed.
[0114] When you press the "Rewatch at Actual Speed" button, all short videos will be played at actual speed in order of their release time.
[0115] When you press the "Accelerate Rewatch" button, all short videos will be played in order of oldest to newest at an accelerated speed of X, set by the streamer. [Explanation of symbols]
[0116] H1 Streamer Device H1.1 processor H1.2 Audio input device H1.3 Display device H1.4 Storage device H1.5 Communication equipment H1.6 Bus H2 Network H3 Server H3.1 processor H3.2 Storage device H3.3 Communication equipment H3.4 Bus H4 Viewer terminal H4.1 processor H4.2 Audio output device H4.3 Display device H4.4 Storage Device H4.5 Communication equipment H4.6 Bus C1 Speech C2 Delivery Stream C3 Streamed Videos C4.1 Video Category 1 C4.2 Video Category 2 C4.N Video Category N C5.1 Short video title 1 C5.2 Short video title 2 C5.N Short video title N C6.1 Remind video 1 C6.2 Remind video 2 C6.N-1 Remind Video N-1 C7 Background Music C8.1 Short Video 1 C8.2 Short Video 2 C8.3 Short Video 3 C8.4 Short Video 4 C8.5 Short Video 5 C8.6 Short Video 6 C8.N Short Video N C9 full video D1 Document Data D2 Data Stack D3 Document Structure Tree D4 Glossary D5 Resource Keyword Stack D6 Semantic Continuity Stack D7 Audio String Stack D8 Voice String Keyword Stack D9 Page Break Stack D10 Temporal Continuity Stack D11 Split position stack D12 Stack at page change time D13 Request timestamp D14 Estimated playback completion timestamp P1 Short video generation program P2 Short Video Viewing Program M1 First operation menu M2 Second operation menu S1 Delivery Block S2 Data Analysis Block S3 Buffering Blocks S4 Audio Analysis Block S5 Split position determination block S6 Split Block S7 Post-editing Block S8 Recording Block S9 Viewer request reception block S10 Request section identification block S11 Accelerated playback time calculation block S12 Request section acceleration playback block S13 Tracking section acceleration playback block
Claims
1. A computer, a character string analysis means for extracting keywords from character string data included in material data used for videos in an online distribution service of video content consisting of video and audio including text, and from character string data generated by image recognition from image and video data; a semantic analysis means for detecting semantic divisions that exist in the overall flow of the video material data by verifying the locations and frequency of occurrence of keywords in the material data; a distribution progress monitoring means for detecting distribution time segment information consisting of a segment start position, a segment end position, and a segment distribution duration based on the segment start position and the segment end position of each semantic segment on the video distribution time axis by detecting a position on the video distribution time axis where each keyword constituting the detected semantic segment appears during the video distribution process; a video segmentation generating means for dividing the entire video into a plurality of video segments based on optimal start and end position information of each segment, which is calculated based on the efficiency of transmitting semantic content and ease of handling in terms of time, by multiplying semantic segmentation information on the video material data with distribution time segmentation information at the time of video distribution; This program functions as a supplementary editing means for processing each of the generated multiple video segments into an independent yet serialized short video work by adding to each of the generated multiple video segments segment title video data generated by processing a part of the distributed video data, reminder video data summarizing the content of the immediately preceding video segment, and segment ending video data.
2. 2. The program according to claim 1, wherein the progress monitoring means has at least a voice analysis means for converting uttered voice data into a character string.
3. 2. The program according to claim 1, wherein said supplemental editing means has at least a sequential storage means for storing video distribution data.
Citation Information
Patent Citations
Network conference system, conference server, recording server and conference terminal
JP2005244522A
Lecture-video-content processing apparatus and program
JP2012109820A
Document correlation device and document correlation method
WO2005069171A1
Information processing device and information processing method
WO2019155695A1