Video generation device, video generation method, and program

The system automatically generates explanatory text from presentation materials by analyzing character attributes and using a trained model, addressing the burden of user-prepared text and enabling precise video playback time control.

JP7794515B1Active Publication Date: 2026-01-06VALUE UPDATE CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2025118472
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2026-01-06
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

Existing methods for generating explanatory videos from presentation materials require users to prepare explanatory text for each slide, which is burdensome.

Method used

A system that automatically generates explanatory text from material data by detecting character information and using a trained model to create commentary, including attributes like character size, position, and color, without requiring user-prepared text.

Benefits of technology

Enables automatic generation of explanatory text, reducing user burden and allowing for precise control over video playback time, thus enhancing the efficiency of video creation from materials lacking audio information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007794515000001_ABST
    Figure 0007794515000001_ABST
Patent Text Reader

Abstract

To automatically generate explanatory text without preparing text for explanatory audio in advance when generating explanatory video based on materials that do not have audio information. [Solution] The video generation device 10 includes an information acquisition unit 112 that acquires material data that is the basis for generating a video, a material analysis unit 113 that detects information for generating explanatory text from the material data, an explanatory text generation unit 114 that generates prompts for generating explanatory text based on the detected information and generates explanatory text by using the generated prompts in a trained model, and a video generation unit 116 that generates a video using the material data and explanatory text.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a moving image generating device, a moving image generating method, and a program. [Background technology]

[0002] In recent years, services and platforms that handle videos, such as YouTube (registered trademark) and Instagram, have become widely used by the general public, and are now being used not only by individuals but also in the business world. In the business world, presentation materials have traditionally been created using presentation creation applications such as Microsoft PowerPoint, or converted to PDF format using Adobe Acrobat applications.

[0003] However, presentation materials often do not contain audio information, and the content is difficult to convey unless the recipient actively reads the materials. For users accustomed to viewing videos, there is an emerging need to watch videos of presentation materials in order to understand them more efficiently and without the hassle.

[0004] To meet such demands, for example, Patent Document 1 proposes a file generation method for generating video with explanatory audio from presentation materials consisting of multiple slides. Specifically, Patent Document 1 discloses that presentation materials with audio are generated by preparing in advance each slide of the presentation material and a note containing explanatory text for that slide, and adding audio data obtained by speech synthesis of the note to the slide. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent No. 7048141 Summary of the Invention [Problem to be solved by the invention]

[0006] The file generation method disclosed in Patent Document 1 requires the user to prepare explanatory text for each slide in advance using a note function, etc. This places a heavy burden on the user.

[0007] The present invention has been made in consideration of the above circumstances, and aims to provide a video generation device, a video generation method, and a program that can automatically generate explanatory text when generating an explanatory video based on materials that do not contain audio information, without having to prepare text for the explanatory audio in advance. [Means for solving the problem]

[0008] One aspect of the present invention is to and one or more still images The system includes an information acquisition means for acquiring material data, a material analysis means for detecting information for generating explanatory text from the material data, an explanatory text generation means for generating a prompt for generating explanatory text based on the detected information and using the generated prompt in a trained model to generate explanatory text, and a video generation means for generating a video using the material data and the explanatory text. The material analysis means detects character information including character strings and accompanying information of the character strings from each still image constituting the material data, and adds pre-registered attribute information to each character string based on the accompanying information, and the comment generation means creates a prompt using the attribute information of the character string to generate the comment, and the accompanying information includes at least one of the size of the character, the display position in the still image, the arrangement, and the color. It is a video generation device.

[0009] One aspect of the present invention is to and one or more still images Get material data Information acquisition and detecting information for generating a description from the material data. Material Analysis generating a prompt for generating an explanatory sentence based on the detected information, and generating the explanatory sentence by using the generated prompt in a trained model. Explanation generation and generating a video using the material data and the commentary. Video Generation The process is carried out by a computer The material analysis step detects character information including a character string and accompanying information of the character string from each still image constituting the material data, and adds pre-registered attribute information to each character string based on the accompanying information. The comment generation step creates a prompt using the attribute information of the character string to generate the comment, and the accompanying information includes at least one of a character size, a display position in the still image, an arrangement, and a color. This is a video generation method.

[0010] One aspect of the present invention is a program for causing a computer to function as the moving image generating device. [Effects of the Invention]

[0011] According to the present invention, when generating an explanatory video based on materials that do not contain audio information, explanatory text can be automatically generated without having to prepare text for the explanatory audio in advance. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a diagram showing an example of the overall configuration of a moving image generation system according to a first embodiment of the present invention. [Figure 2] 1 is a diagram illustrating an example of a hardware configuration of a moving image generating device according to a first embodiment of the present invention. [Figure 3] 1 is a functional configuration diagram showing an example of functions provided in a moving image generating device according to a first embodiment of the present invention. [Figure 4] FIG. 10 is a diagram showing an example of a video generation screen. [Figure 5] 10 is a flowchart illustrating an example of a processing procedure of the moving image generating device in response to an input operation by a user on a moving image generating screen. [Figure 6] 10 is a flowchart showing an example of a procedure of a material acquisition process executed by an information acquisition unit according to the first embodiment of the present invention. [Figure 7] 4 is a flowchart showing an example of a procedure of a material analysis process executed by a material analysis unit according to the first embodiment of the present invention. [Figure 8] FIG. 10 is a diagram showing an example of a still image of one page constituting a presentation data file, which is material data. [Figure 9] FIG. 10 is a diagram illustrating an example of a page information management table. [Figure 10] FIG. 10 is a diagram showing an example of a video playback table. [Figure 11] 10 is a flowchart showing an example of a procedure of a comment generation process executed by a comment generation unit according to the first embodiment of the present invention. [Figure 12]FIG. 10 is a diagram illustrating a method for generating an explanation. [Figure 13] FIG. 10 is a diagram illustrating a method for generating an explanation. [Figure 14] 4 is a flowchart showing an example of a procedure of a voice data generation process executed by a voice data generation unit according to the first embodiment of the present invention. [Figure 15] 4 is a flowchart showing an example of a procedure of a moving image generation process executed by a moving image generation unit according to the first embodiment of the present invention. [Figure 16] FIG. 10 is a diagram showing an example of an editing screen for editing an explanation. [Figure 17] FIG. 10 is a diagram showing an example of an editing screen for editing breath timing. [Figure 18] FIG. 10 is a diagram showing an example of an explanatory text with breath marks inserted therein. [Figure 19] FIG. 10 is a diagram showing an example of an editing screen for editing intonation. [Figure 20] FIG. 10 is a diagram showing an example of a playback time setting screen. [Figure 21] FIG. 10 is a diagram showing another example of the playback time setting screen. [Figure 22] FIG. 10 is a functional configuration diagram showing an example of functions provided in a moving image generating device according to a second embodiment of the present invention. [Figure 23] FIG. 10 is a diagram showing an example of a material data editing screen. [Figure 24] FIG. 10 is a functional configuration diagram showing an example of functions provided in a moving image generating device according to a third embodiment of the present invention. [Figure 25] FIG. 10 is a diagram showing an example of an element data input screen. [Figure 26] FIG. 10 shows an example of a prompt for generating decoration data and an example of the result when the prompt is given to a generation AI. [Figure 27] FIG. 10 is a diagram showing an example of an avatar confirmation screen for confirming an avatar. [Figure 28] FIG. 10 is a diagram showing an example of a background image confirmation screen for confirming a background image. [Figure 29]FIG. 10 is a diagram showing an example of a video playback table in which paths of avatar images and background images are registered. [Figure 30] FIG. 10 is a diagram showing an example of an explanatory video after decoration. DETAILED DESCRIPTION OF THE INVENTION

[0013] [First embodiment] A moving image generating device, a moving image generating method, and a program according to a first embodiment of the present invention will be described below with reference to the drawings.

[0014] [Network configuration] FIG. 1 is a diagram showing an example of the overall configuration of a video creation system 1 according to this embodiment. As shown in FIG. 1, the video creation system 1 includes a video creation device 10 and a client terminal 100. The video creation device 10 and the client terminal 100 are connectable via a network 3, and are configured to be able to send and receive information to and from each other. Examples of the network 3 include a Bluetooth network, the Internet, Wi-Fi, Li-Fi, a mobile communication system (3G, 4G, 5G, 6G, LTE, etc.), a wireless LAN, and a wired LAN. The connection may also be via a dedicated line, VPN, etc. Although one client terminal 100 is shown as an example in FIG. 1, the number of client terminals 100 connected to the video generation device 10 is not limited to this.

[0015] The client terminal 100 is an information processing device used by a user of the service provided by the video production device 10, and is configured with an environment that allows access to the video production device 10, such as the Internet or an in-house system. Examples of the client terminal 100 include a desktop PC, a notebook PC, a tablet terminal, and a smartphone. The client terminal 100 includes an information processing circuit including a CPU, a display for displaying information, an input unit for a user to perform input operations, and a communication unit for communicating with other devices. Note that the configuration of an information processing device is well known, and therefore a detailed description thereof will be omitted here.

[0016] The moving image generation device 10 generates moving images based on instructions received from the client terminal 100. The moving image generation device 10 may be configured as a system on a so-called on-premise server, or may be configured as a virtualized system on a cloud system.

[0017] 2 is a schematic diagram showing an example of the hardware configuration of the moving image generation device 10 according to this embodiment. The moving image generation device 10 is a so-called computer, and as shown in FIG. 2, includes, for example, a CPU (Central Processing Unit: processor) 201, a main memory 202, a secondary storage device 203, and a communication device 204. Furthermore, the video generation device 10 may also include an external interface for connecting external devices, an input device as a man-machine interface, a display for displaying data, and the like. These units are interconnected directly or indirectly via a bus, and work together to execute various processes.

[0018] The computer system may include one or more CPUs 201. When multiple CPUs 201 are provided, they may cooperate with each other to realize processing. The computer system may also include other processors such as a GPU.

[0019] The main memory 202 is configured by a writable memory such as a RAM (Random Access Memory), and is used as a work area for reading out the execution program of the CPU 201, writing data processed by the execution program, etc. A plurality of main memories 202 may be provided.

[0020] The secondary storage device 203 is, for example, a non-transitory computer readable storage medium. Examples of the secondary storage device 203 include semiconductor memory, magnetic recording media, optical recording media, and magneto-optical recording media. Specific examples include flash memory, SSD (Solid State Drive), ROM (Read Only Memory), HDD (Hard Disk Drive), CD, DVD, and Blu-ray (registered trademark) Disc. Furthermore, some of the secondary storage devices 203 may be provided as cloud storage.

[0021] The secondary storage device 203 stores, for example, an OS, applications, and various data and files for realizing the functions of the video generation device 10. A plurality of secondary storage devices 203 may be provided, and the programs and data for realizing the processes (functions) described below may be divided and stored in each secondary storage device 203.

[0022] The communication device 204 functions as an interface for connecting to a network to communicate with other devices and sending and receiving information. The communication device 204 communicates with other devices, for example, via a wired or wireless connection. Examples of wireless communication include Bluetooth (registered trademark), Wi-Fi, and communication using a dedicated communication protocol. An example of wired communication is a wired local area network (LAN).

[0023] FIG. 3 is a functional configuration diagram showing an example of functions provided in the video production device 10 according to this embodiment. A series of processes for realizing various functions of the video generation device 10 described below is stored in the secondary storage device 203 in the form of a program, for example, and the CPU 201 reads this program into the main memory 202 and executes information processing and arithmetic operations to realize various functions.

[0024] Furthermore, the programs for realizing various functions of the video generation device 10, which will be described later, may be provided in a form in which they are pre-installed in the secondary storage device 203, or in a form in which they are installed in a state in which they are stored in a computer-readable external storage (recording medium), or in a form in which they are distributed via wired or wireless communication means, etc. Examples of computer-readable storage media include magnetic disks, optical disks, magneto-optical disks, CD-ROMs, DVD-ROMs, and semiconductor memories.

[0025] As shown in FIG. 3, the video production device 10 includes, for example, a storage unit 110, an information acquisition unit 112, a material analysis unit 113, a commentary production unit 114, an audio data production unit 115, a video production unit 116, and an editing unit 117. The moving image generating device 10 also includes a material data storage unit 120 for storing material data, an audio data storage unit 121 for storing audio data, and a moving image data storage unit 122 for storing moving image data.

[0026] The storage unit 110 stores data and various parameters required for generating a moving image. For example, it stores various input information acquired by an information acquisition unit 112 (described later). The storage unit 110 is realized by, for example, a main memory 202 and a secondary storage device 203. Some of the information is temporarily saved in the main memory 202, and the other information is stored in the secondary storage device 203.

[0027] The information acquisition unit 112 acquires information related to video generation from the client terminal 100. For example, the information acquisition unit 112 causes the client terminal 100 to display various input screens (e.g., video content generation screens) that allow the user to input various data required for video generation. Then, the information acquisition unit 112 acquires the data input via these input screens.

[0028] For example, the information acquisition unit 112 acquires material data that is the basis for generating a moving image from the client terminal 100. Examples of the material data include presentation materials created using presentation material creation software. The file format of the presentation materials is not particularly limited, but examples include data files created using PowerPoint provided by Microsoft, PDF files, and the like. The material data may consist of, for example, one or more pages (still images) that do not include audio. The material data may also include additional information to explain the content of each page. Such additional information is supplementary information and is not essential for generating a video.

[0029] The material analysis unit 113 detects information for generating explanatory text from the material data. The commentary generation unit 114 automatically generates a commentary based on the information detected by the material analysis unit 113. For example, the commentary generation unit 114 generates a prompt for generating a commentary based on the information detected by the material analysis unit 113, and generates a commentary by using the generated prompt in a trained model (for example, a generation AI).

[0030] The audio data generating unit 115 generates audio data from the commentary generated by the commentary generating unit 114 . The moving image generating unit 116 generates moving images using, for example, material data and audio data. The editing unit 117 provides the commentary generated by the commentary generating unit 114 in a state that allows the user to edit it, and edits the commentary based on the editing instructions of the user, for example.

[0031] Next, a video generation screen, which is a user interface used for generating a video according to this embodiment, will be described with reference to the drawings. Fig. 4 is a diagram showing an example of a video generation screen 300.

[0032] This video generation screen 300 may be constructed as a Web application on the video generation device 10, or may be constructed as a native application on the client terminal 100. In each case, the processing of the information acquisition unit 112 will be slightly different, which will be described later.

[0033] On the video creation screen 300, the material input area 301 functions as an input section for importing the source material data from which the video will be created, such as a data file of presentation materials. By dragging and dropping the file into this area, the data file can be uploaded from the client terminal 100.

[0034] The video generation screen 300 may be provided with an input area for setting conditions for the video to be generated, an area for checking a text explanation of the generated video, and the like.

[0035] For example, the video generation screen 300 may be provided with an input area for the user to input information related to the video playback time. For example, the video generation screen 300 is provided with an input area 302 for inputting information related to the playback time of the entire video, and an input area 303 for inputting information related to the playback time per page. Note that it is not necessary for both of these input areas 302 and 303 to be provided, and it is sufficient that at least one of them is provided.

[0036] An input area 302 is an input area for the user to input the total time of the video in advance. An input area 303 is an input area where the user can input the playback time (display time) for each page. For example, if 10 seconds is input in the input area 303 and the material data consists of 10 pages, a video is generated so that the total video playback time is 100 seconds.

[0037] The user uploads material data on video creation screen 300, and if necessary, inputs information about video playback time in either input area 302 or 303, and then presses video creation start button 304. This causes video creation device 10 to start creating a video based on the input information.

[0038] Specifically, input information entered on the video generation screen 300 is acquired by the information acquisition unit 112, and this information is analyzed by the material analysis unit 113. Then, based on the analysis results, an explanatory text is generated by the explanatory text generation unit 114, and the explanatory text is converted into audio data by the audio data generation unit 115.

[0039] On the video creation screen 300 in FIG. 3, a video playback area 305 provides a function for playing back and checking the created video. Furthermore, the commentary generated by the commentary generating unit 114 is displayed for each page in the commentary confirmation area 306. This allows the user to check and edit the content. For example, the user selects a portion of the text displayed using an input device such as a mouse (not shown), changes the text using an input device such as a keyboard, and presses the commentary edit button 307.

[0040] In this way, when an editing instruction is input to the commentary, the commentary is edited by the editing unit 117 of the video production device 10. Furthermore, audio data is generated in the audio data production unit 115 based on the edited content, and a video is generated based on the edited information. As a result, the video that can be viewed in the video playback area 305 is updated.

[0041] Next, the operation of the moving image generation device 10 in response to an input operation by the user on the above-mentioned moving image generation screen 300 will be described with reference to the drawings. Fig. 5 is a flowchart showing an example of a processing procedure of the moving image generation device 10 in response to an input operation by the user on the moving image generation screen 300.

[0042] First, when a user uploads material data to the material input area 301, the information acquisition unit 112 of the moving image production device 10 acquires the material data (S501).

[0043] Next, when the user inputs information about the video playback time into input area 302 and input area 303, information acquisition unit 112 acquires the information about the video playback time input by the user (S502). This acquires the total playback time of the video or the video playback time per page. The acquired information about the video generation time is saved as a variable in main memory 202, for example.

[0044] Next, when the user presses the moving image generation start button 304, a moving image generation start instruction is input to the moving image generation device 10. As a result, the moving image generation device 10 generates a moving image based on the input information (S503).

[0045] Next, information about the generated video is provided to the user by displaying it on the client terminal 100 (S504). Specifically, the video is displayed in the video playback area 305 of the video generation screen 300, and the commentary is displayed in the commentary confirmation area 306. The user checks the displayed commentary and edits it as necessary. When the user performs editing and inputs an editing instruction (S505: YES), the video creation device 10 edits the commentary based on the editing instruction (S506) and returns to step S503. As a result, a video is created based on the edited information.

[0046] On the other hand, if a certain period of time has passed without any instruction to edit the video being input on the video creation screen 300 (S505: NO), the process ends.

[0047] Next, each unit (each function) of the video production device 10 according to this embodiment will be described in more detail.

[0048] [Information acquisition unit 112] First, a detailed description will be given of the material acquisition process executed by the information acquisition unit 112. FIG.

[0049] When the information acquisition unit 112 acquires material data, it acquires from the page information attached to the material data how many pages the material data consists of, and stores the information in the storage unit 110 (S601). Next, the information acquisition unit 112 stores the material data as still images for each page in the material data storage unit 120 (S602), and when the still images of all pages have been saved, the material acquisition process ends.

[0050] [Material Analysis Section 113] Next, a detailed description will be given of the material analysis processing executed by the material analysis unit 113. Fig. 7 is a flowchart showing an example of the procedure of the material analysis processing executed by the material analysis unit 113.

[0051] The material analysis unit 113 reads the number of pages of material data stored in the storage unit 110 (S701), and then detects text information included in the still images saved in the material data storage unit 120 for each page (S702).

[0052] Here, the detection of character information will be described with reference to Fig. 8. Fig. 8 is a diagram showing an example of a still image of one page that constitutes a presentation data file, which is material data.

[0053] This still image contains character strings 401 to 404. Character string 401 is written in the largest character size on the page, and is displayed at the top of the page. Character strings 402 and 403 are written in a smaller character size than character string 401, have numbers at the beginning, and are written in the same starting position.

[0054] The character information detected in step S702 includes the character string contained in the still image and its associated information, which includes at least one of the following: character size, display position in the still image (top, center, bottom, etc.), layout, and color. As for the specific method for detecting the character information, any known technique may be appropriately adopted, and the specific method is not particularly limited. For example, image processing using OCR (Optical Character Recognition) technology on a page's worth of still images 400, or a method of directly obtaining text information by analyzing the data format of presentation materials, which are raw data, can be used.

[0055] After character information has been detected in this manner, the material analysis unit 113 then adds pre-registered attribute information to each character string based on the accompanying information (S703). Attribute information includes, for example, title, subtitle, heading, main text, bullet points, numbered lists, footnotes, captions, highlighted text, and link text.

[0056] For example, the material analysis unit 113 has attribute association information in which each piece of attribute information is associated with its feature information, and adds attribute information to a character string based on this attribute association information. For example, a title is written at the top of a still image and has a relatively large character size. Therefore, attribute information called a title is added to a character string 401 (see FIG. 8) having such characteristics. Also, if a number is written at the beginning of a character string, the character size is smaller than that of the title, and there are multiple character strings written in the same character size, and all are arranged in the same starting position, these character strings are likely to be headings of an itemized list. Therefore, attribute information of a heading of an itemized list is added to character strings 402 to 404 having such characteristics. In this way, by adding attribute information to character strings, meaning is assigned to each character string.

[0057] Next, the material analysis unit 113 stores the analysis results for each page in the storage unit 110 (S704). Then, it is determined whether or not analysis has been completed for all pages that make up the material data (S705). As a result, if analysis has not been completed for all pages (S705: NO), the process returns to step S702 and repeats the above-mentioned processing. On the other hand, if analysis has been completed for all pages (S705: YES), the process ends.

[0058] Next, a description will be given of an example of information stored in the storage unit 110. For example, the storage unit 110 stores a page information management table in which various information such as the analysis results of the material analysis unit 113 is associated with each page constituting the material data.

[0059] Fig. 9 is a diagram showing an example of a page information management table. As shown in Fig. 9, the page information management table registers, for each page constituting material data, a page number 1101, a page playback time 1102, path information 1103 of a still image, character string information 1104, character string accompanying information 1105, commentary text data 1106, path information 1107 of commentary audio data, and path information 1108 of a video, for example.

[0060] The page number 1101 is identification information for identifying each page, and the number of page numbers 1101 is generated for each page that constitutes the material data. The page playback time 1102 is the playback time for each page. The page playback time 1102 is a value based on information about the playback time of the video input in the input areas 302 and 303 of the video generation screen 300 (FIG. 4).

[0061] For example, a video playback table 1109 may be stored in the storage unit 110 in order to calculate the video playback time. Fig. 10 is a diagram showing an example of the video playback table 1109. As shown in Fig. 10, the video playback table 1109 registers, for the material data, the total video playback time, the video playback time per page, and the number of pages.

[0062] For example, if the user inputs the total playback time of the video (e.g., 180 seconds) on the video creation screen 300, the input value (e.g., 180 seconds) is stored in the video length field. The number of pages field stores the number of pages of the material data, in other words, the number of still images (e.g., 9). The length per page field stores the value obtained by dividing the length of the video by the number of pages (e.g., 20).

[0063] Furthermore, when the user inputs the playback time per page (e.g., 20 seconds) on the video creation screen 300, the input value (e.g., 20 seconds) is stored in the length per page field. The number of pages (e.g., 9) of the material data is stored in the number of pages field. The value obtained by multiplying the length per page by the number of pages (e.g., 180) is stored in the video length field.

[0064] Then, the playback time per page registered in the video playback table 1109 is registered as the playback time per page in the page information management table shown in FIG.

[0065] Still image path information 1103 is access information to the storage location where each still image is stored. The character string information 1104 is information about character strings detected by the material analysis unit 113 . The character string accompanying information 1105 is information accompanying each character string. Although not shown in the figure, the attribute information of each character string described above may be stored in association with each character string.

[0066] The comment text data 1106 is text data of the comment generated by the comment generation unit 114 . The path information 1107 for commentary audio data is information for accessing the storage location where the commentary audio data is stored. Video path information 1108 is information for accessing a storage location where video data generated by video generating unit 116 is stored. The process of generating the commentary and the process of generating the video data will be described in detail later.

[0067] [Explanatory sentence generation unit 114] Next, a detailed description will be given of the comment generation process executed by the comment generation unit 114. Fig. 11 is a flowchart showing an example of the procedure of the comment generation process executed by the comment generation unit 114.

[0068] First, the page information management table stored in the storage unit 110 is referenced to read the total number of pages of the material data (S801). Next, character information (character strings and their associated information or attribute information) is acquired for each page (S802). Then, a prompt for generating an explanatory text is generated based on the acquired information, and the generated prompt is used in a trained model to generate an explanatory text for each page (S803). An example of the trained model is a generation AI such as a large-scale language model (LLM).

[0069] A method for generating explanatory text when using LLM will be described with reference to Figures 12 and 13. In this embodiment, an example will be described in which an LLM service that provides an API (Application Programming Interface) is used. In other words, instructions to the LLM are executed via the API, and the explanatory text generated as a result is received.

[0070] Fig. 12 shows a prompt 1201 generated using only character string information detected from the still image of the material data shown in Fig. 4, and the resulting explanatory text 1202. The first half of the prompt 1201 indicates what to generate, and the second half shows the information that will be the basis for creating the explanatory text. Prompt 1201 instructs that a non-itemized sentence be generated in the form of a narration based on character strings extracted from one page of material data. The resulting explanatory text 1202 is written in a colloquial style suitable for a person to read aloud.

[0071] Fig. 13 also shows a prompt 1301 generated using character string information detected from the still image of the material data shown in Fig. 4 and its attribute information, and an explanatory text 1302 generated as a result. As shown in Fig. 13, the prompt includes attribute information of the character string, such as a specified title and the importance of the content to be explained. Comparing the commentary 1302 shown in FIG. 13 with the commentary 1202 shown in FIG. 12, it can be seen that the commentary tends to be generated in the indicated direction.

[0072] The commentary thus generated is registered in the page information management table in the storage unit 110 (S804). In other words, it is stored in the storage unit 110 in association with each page number. Next, it is determined whether the above process has been performed for all pages of the material data (S805), and if not (S805: NO), the process returns to step S802 and repeats. Then, when the generation and saving of the commentary for all pages is completed (S805: YES), the process ends.

[0073] In creating the commentary, the amount of text data (e.g., data capacity or number of characters) that constitutes the commentary may be adjusted based on the page playback time. For example, the number of characters contained in each page of material data varies. Therefore, if each page is set to the same playback time, the amount of commentary must be approximately uniform on each page. In this way, the amount of commentary per page can be determined from the playback time and the voice reading speed (described later), and the commentary may be generated to match that amount.

[0074] [Audio data generation unit 115] Next, a detailed description will be given of the audio data generation process executed by the audio data generation unit 115. Fig. 14 is a flowchart showing an example of the procedure of the audio data generation process executed by the audio data generation unit 115.

[0075] First, the total number of pages of the material data is read from the page information management table stored in the storage unit 110 (S901). Next, the explanatory text registered in the page information management table is read for each page (S902). Next, the playback time for each page is read from the page information management table (S903).

[0076] Next, audio data is generated from the commentary using the page playback time as a parameter (S904). Conversion to audio data can be performed using known techniques, such as VOICEVOX. Next, the generated voice data is stored as a voice file in the voice data storage unit 121 (S905). The file path to this storage destination is also registered in the page information management table stored in the storage unit 110 (S906). When converting to voice data, pre-registered default settings are used for parameters required for the voice conversion process (for example, the characteristics of the voice data to be generated). Next, it is determined whether the above process has been performed for all pages of the material data (S907), and if not (S907: NO), the process returns to step S902 and repeats. Then, when the generation and saving of audio data for all pages is completed (S907: YES), the process ends.

[0077] [Video Generation Unit 116] Next, a detailed description will be given of the moving image generation process executed by the moving image generation unit 116. Fig. 15 is a flowchart showing an example of the procedure of the moving image generation process executed by the moving image generation unit 116.

[0078] First, the total number of pages of the material data is read from the page information management table stored in the storage unit 110 (S1001). Next, the playback time per page registered in the page information management table is obtained (S1002), and still images that make up the material data are obtained from the material data storage unit 120 based on the storage destination link information (S1003).

[0079] Next, video data is generated from the still images that make up the material data based on the acquired playback time for each page (S1004). This process can be performed using a library such as ffmpeg, for example. Next, the audio data of the commentary for each page is acquired from the path information of the audio data registered in the page information management table (S1005), and the acquired audio data is synthesized with each page of the video data to generate video data with audio (S1006). The generated video data with audio is stored in the video data storage unit 122, and its file path is registered in the video path information 1108 of the page information management table.

[0080] In this way, video data with audio is created for all pages that make up the material data, and when the saving process is completed, the video generation unit 116 reads the video data from the video path information 1108 registered in the page information management table, combines the video data for all pages to complete the explanatory video, stores the completed explanatory video in the video data storage unit 122 (S1007), and terminates this process.

[0081] When a moving image with audio is generated in this way, the moving image is displayed in a moving image playback area 305 of a moving image generation screen 300, as shown in FIG.

[0082] As explained above, the video generation device 10 and video generation method according to this embodiment include an information acquisition unit 112 that acquires material data that is the basis for generating a video, a material analysis unit 113 that detects information for generating explanatory text from the material data, an explanatory text generation unit 114 that generates a prompt for generating explanatory text based on the detected information and uses the generated prompt in a trained model to generate explanatory text, and a video generation unit 116 that generates a video using the material data and explanatory text.

[0083] In this way, explanatory text is automatically generated from the material data, eliminating the need for users to prepare explanatory text in advance, as was previously the case, and thereby reducing the burden on users.

[0084] Furthermore, according to the video generation device 10 and video generation method of this embodiment, the information acquisition unit 112 acquires information regarding the video playback time input by the user, and the video generation unit 116 generates a video based on the information regarding the video playback time.

[0085] For example, with the conventional file generation method disclosed in Patent Document 1, it is difficult to control the video playback time in advance. This is because the playback time is calculated from the time it takes for the text written in the notes to be read aloud using electronic voice processing. With this method, it is not possible to know in advance how long the overall video playback time of the presentation materials will be. According to this embodiment, the playback time of the commentary for each page can be given in advance as input information, so that a video with a desired playback time can be automatically generated.

[0086] In this embodiment, the moving image is generated by generating audio data from the commentary and synthesizing the audio data with the still images, but the present invention is not limited to this. For example, the audio data generation process may be omitted, and the moving image may be generated by displaying the commentary corresponding to each still image as text information such as a caption.

[0087] [Variation 1] In the present embodiment, the editing work has been described as a configuration in which the commentary can be edited on the video generation screen 300 shown in Fig. 4, but the present invention is not limited to this. For example, the user may be able to edit the pronunciation of kanji characters included in the commentary.

[0088] For example, the video generation device 10 may cause the client terminal 100 to display an editing screen 1400 as shown in Fig. 16. This editing screen 1400 may be displayed in the commentary confirmation area 306 of the video generation screen 300 shown in Fig. 4, or may be displayed as a separate screen.

[0089] As shown in FIG. 16, the editing screen 1400 has an explanatory text display area 1401 where explanatory text is displayed, and a voice reading display area 1402 where audio data generated based on the explanatory text is displayed in hiragana. Here, kanji characters may be automatically highlighted in the commentary display area 1401, and the pronunciation of the kanji characters may be similarly highlighted in the voice reading display area 1402. On this editing screen 1400, the user can edit the pronunciation of the voice reading. When the user makes edits and presses edit button 1403, the editing unit 117 of the video production device 10 edits the commentary based on the edited pronunciation information. Then, the voice data production unit 115 converts the edited commentary into voice data and updates the voice data stored in the voice data storage unit 121. This makes it easy to correct mistakes in reading kanji characters, which are likely to occur when using text-to-speech synthesis technology. In addition, if the commentary contains multiple identical words or character strings, a batch conversion function that allows the pronunciation of all the words or character strings to be edited at once may be provided.

[0090] The editing unit 117 may also have a function of editing the timing of breathing when the commentary is read aloud. For example, the video generating device 10 may display an editing screen on the client terminal 100 for the user to edit the timing of breathing when an explanatory text is read aloud. Fig. 17 is a diagram showing an example of the editing screen for editing the timing of breathing. This editing screen may be displayed in the commentary confirmation area 306 of the video creation screen 300 shown in FIG. 4, or may be displayed as a separate screen.

[0091] As shown in Fig. 17, the editing screen has a commentary display area where commentary is displayed, and a voice reading display area where audio data generated based on the commentary is displayed in hiragana. In the voice reading display area, the text to be read is displayed so that breath positions can be edited. In this voice reading display area, breath positions are indicated by, for example, breath marks (V) 1602, 1603. These breath marks 1602, 1603 can be deleted, added, or moved. For example, the user can select the breath tool 1601 for editing breath timing on the editing screen and add a breath mark at the desired position by clicking on the part where the user wants to take a breath.

[0092] When the user edits the breath position and presses edit button 1403, editing unit 117 of moving image production device 10 generates commentary with information indicating the breath (for example, markers) 1701, 1072 input, as shown in Fig. 18. Then, audio data production unit 115 converts the edited commentary into audio data, and updates the audio data stored in audio data storage unit 121.

[0093] The editing unit 117 may also have a function of editing the intonation when the commentary is read aloud. For example, the moving image generating device 10 may display an editing screen on the client terminal 100 for the user to edit the intonation used when the commentary is read aloud. Fig. 19 is a diagram showing an example of an editing screen 1700 for editing the intonation. This editing screen 1700 may be displayed in the commentary confirmation area 306 of the video creation screen 300 shown in FIG. 4, or may be displayed as a separate screen.

[0094] 19, editing screen 1700 is provided with a commentary display area where commentary is displayed, and a voice reading display area where audio data generated based on the commentary is displayed in hiragana, similar to editing screen 1600. Editing screen 1700 also has intonation editing area 1703. Intonation editing area 1703, each phoneme of the audio data and its respective pitch are displayed in correspondence with the elapsed time since the start of audio playback, and the pitch of each phoneme is visually represented in an editable format. On this editing screen, the user can play back the commentary audio and check its contents. If the user wants to edit the intonation, they can change the intonation of the audio by selecting the target phoneme in the intonation editing area 1703 and adjusting the pitch of each phoneme up or down. When the user presses the save button 1704 after editing the pitch of each phoneme, the editing unit 117 of the video generation device 10 edits the intonation of the audio data based on the user's instructions and updates the corresponding audio data stored in the audio data storage unit 121.

[0095] In this way, the video generation device 10 is provided with various editing functions that allow the user to edit the automatically generated explanatory text and its audio data, allowing the user to make detailed edits and customize the video.

[0096] [Variation 2] In this embodiment, the playback time of each page constituting the material data is the same, but this is not limited to this. For example, the playback time of each page may be configured to be set individually. For example, the video generating device 10 presents the user with a playback time setting screen on the client terminal 100 for setting the playback time of each page.

[0097] Fig. 20 shows an example of a playback time setting screen. Fig. 20 has a setting field for setting the playback time for each page. The user can set or change the playback time for each page by inputting a desired value in this setting field. Furthermore, when each page is edited, the playback time of the entire video is calculated and displayed accordingly. The page advance interval may also be editable. The playback time set by the user on the playback time setting screen is reflected in the page information management table stored in the storage unit 110. Specifically, it is entered into the page playback time 1102 of the page information management table 1100. Then, audio data is generated based on the playback time entered into this page information management table 1100.

[0098] Furthermore, information regarding the playback time per page can be set not only as a specific numerical value, but also as a playback time "level." An example of the playback time setting screen in this case is shown in Figure 21. Figure 21 shows another form of the playback time setting screen, which is configured so that the playback time level can be selected for each page. The user can select or change the desired playback time for each page from three levels: "long," "medium," and "short."

[0099] The playback levels set by the user on this playback time setting screen are converted into corresponding specific time values ​​and reflected in the page information management table stored in the storage unit 110. Specifically, the playback times are recorded in the page playback time 1102 of the page information management table 1100. Then, based on the time information recorded in this page playback time 1102, corresponding audio data is generated. In the example of Figure 21, the playback time levels are divided into three levels (long, medium, and short), but this is not limited to this and it is possible to expand it to four or more levels or simplify it to two levels.

[0100] In this way, by allowing users to set the playback time for each page, it becomes possible to flexibly customize the audio reading speed and the playback time of the entire video.

[0101] In addition, the video playback time can be set by the user for each page, or it can be set automatically based on the number of characters on each page or the number of characters in the explanatory text. For example, the page playback time is not set in advance, but audio data is generated based on the explanatory text data. At this time, the reading speed is unified across pages. The playback time for each page can then be entered into the page information management table 1100. The result can then be presented to the user as the playback time setting screen described above, allowing the user to edit it. This allows the user to make detailed adjustments after being presented with a rough estimate of playback time based on the amount of text contained on the page.

[0102] Second Embodiment For example, you may want to re-create an explanatory video by changing or adding parts of the original presentation materials (material data) for an explanatory video you have already created. You may also want to create a new explanatory video by appropriately combining pages of multiple material data. To meet these needs, the moving image generating device 10a according to the third modification of the present invention further includes a material data generating section 118 as shown in FIG. Fig. 22 is a functional configuration diagram showing an example of functions provided in a moving image generating device 10a according to a second embodiment of the present invention. In Fig. 22, the same components as those in the first embodiment described above are assigned the same reference numerals. Below, differences from the first embodiment, namely, the material data generating unit 118, will be mainly described.

[0103] The material data generation unit 118 provides the user with a material data editing screen for generating new material data using one or more existing material data by displaying the screen on the client terminal 100. The new material data is then generated based on the editing instructions of the user on the material data editing screen.

[0104] 23 is a diagram showing an example of a material data editing screen. As shown in FIG. 23, the material data editing screen is provided with an original data display area 1801 and a replacement data area 1803. The original data display area 1801 displays the material data that will be used for editing, while the replacement data area 1803 displays the material data that will be used for editing. The material data displayed in either area is data selected by the user from among multiple pieces of material data stored in the material data storage unit 120. Here, an example is shown in which "PPT file name_AAA" is selected as the source material data for editing, and "PPT file name_BBB" and "PPT file name_CCC" are selected as the material data to be used for editing.

[0105] When the user selects any page in original data display area 1801, a still image of the selected page is displayed in preview area 1802, allowing the user to confirm the contents. In this state, when the user selects any page in replacement data area 1803, the selected page is displayed in edit area 1804, allowing the user to confirm the contents.

[0106] Here, when the user operates the replace button 1805, the page displayed in the preview area 1802 is replaced with the page displayed in the edit area 1804. Furthermore, when the user operates the add button 1806, the page displayed in the edit area 1804 can be added before or after the page displayed in the preview area 1802. After the user has performed the above-described editing and created new material data, he or she presses OK button 1807. As a result, the material data created by the user is accepted by information acquisition unit 112a of video creation device 10a and stored in material data storage unit 120. Furthermore, a series of processes related to the video creation described above are performed based on the new material data, thereby enabling a video to be created. In this way, by having the video generation device 10a equipped with the material data generation unit 118, the user can generate new material data on the material data editing screen based on one or more material data stored in the material data storage unit.

[0107] Third Embodiment Next, a moving image generating device and a moving image generating method according to a third embodiment of the present invention will be described with reference to the drawings. Fig. 24 is a functional configuration diagram showing an example of functions provided in a moving image generating device 10b according to a third embodiment of the present invention. In Fig. 24, the same components as those in the first embodiment are assigned the same reference numerals. Below, a description of the points in common with the first embodiment will be omitted, and the differences, namely, the decoration data generating unit 119, will be mainly described.

[0108] The decoration data generation unit 119 generates decoration data for decorating the material data. For example, the decoration data generation unit 119 generates at least one of an avatar and a background image as decoration to increase the interest of video viewers. For example, the decoration data generation unit 119 displays an element data input screen on the client terminal 100 to allow the user to input element data for generating decoration data. This element data input screen may be provided as one of the input screens constituting the above-mentioned video generation screen 300 (see FIG. 4), or may be provided as a supplementary screen when the user desires decoration.

[0109] FIG. 25 is a diagram showing an example of the element data input screen. 25, the element data input screen is provided with input areas for the user to input information about each element item. For example, the element data input screen is provided with an input area 2401 for inputting information about the atmosphere of the video, an input area 2402 for inputting information about the characteristics of the person, an input area 2403 for inputting information about the situation, an input area 2404 for inputting information about the person's clothing, etc.

[0110] The input for each of these element items may be freely written by the user, or may be in a selection format in which the user is given multiple options for each element item and must select one of them. On the element data input screen, when the user inputs the necessary information into the input areas 2401 to 2404 for each item and presses the generate button 2405, the input element data is acquired as input information by the information acquisition unit 112a. Then, the decoration data generation unit 119 generates decoration data based on this element data.

[0111] Specifically, the decorative data generation unit 119 generates a prompt including the acquired element data, and generates at least one of a background image and an avatar image by providing the generated prompt to a trained model (e.g., a generation AI).

[0112] Figure 26 shows an example of a prompt and an example of the result when the prompt is given to the generation AI. Figure 26 shows the prompt when, for example, a user enters "anime-style video" in input area 2401, "blonde male in his 30s" in input area 2402, "as a waiter in a coffee shop" in input area 2403, and "white shirt and green apron" in input area 2404 on the element data input screen.

[0113] As shown in FIG. 26, following the prompts, an avatar image and a landscape image are generated. When the decoration data generation unit 119 generates the decoration data, it causes the client terminal 100 to display a decoration data confirmation screen for the user to confirm the generated decoration data, and provides the user with the decoration data. Figures 27 and 28 show examples of the decoration data confirmation screen. Figure 27 shows an example of an avatar confirmation screen for confirming the avatar, and Figure 28 shows an example of a background image confirmation screen for confirming the background image.

[0114] The avatar confirmation screen shown in FIG. 27 is provided with an avatar image preview area 2602 that displays the avatar screen, and is configured to display the generated avatar image so that the details can be confirmed. Furthermore, in presentation material preview area 2603, a composite image is displayed that allows a preview of what the final explanatory video will look like when created using the avatar image.

[0115] The user can confirm the avatar on the avatar confirmation screen, and if there are no problems, press the OK button 2605 to confirm the avatar. If the user wishes to regenerate the avatar, he or she can press the Retry button 2604. In this case, the decoration data generation unit 119 causes the client terminal 100 to display the element data input screen shown in Fig. 25. This allows the user to start over by inputting the element data for generating decoration data.

[0116] The background image confirmation screen shown in FIG. 28 is provided with a background image preview area 2702 for displaying a background image, and is configured to display the generated background image so that the details can be confirmed. Furthermore, in the presentation material preview area 2703, a composite image is displayed that allows a preview of what the final explanatory video will look like when created using the background image. The user can confirm the background image on the background image confirmation screen, and if there are no problems, press the OK button 2705 to confirm the background image. If the user wishes to regenerate the background image, he or she can press the Retry button 2704. In this case, the decoration data generation unit 119 causes the client terminal 100 to display the element data input screen shown in Fig. 25. This allows the user to start over by inputting the element data for generating decoration data.

[0117] When the decoration data is confirmed by the user, decoration data generation unit 119 stores the confirmed decoration data in decoration data storage unit 123, for example, and registers the paths in the page information management table (see FIG. 9) or video playback table 2109. FIG. 29 shows an example of video playback table 1109 in which paths of avatar images and background images are registered. When generating a moving image, the moving image generating unit 116a refers to the moving image playback table 2109, reads out decoration data stored in the decoration data storage unit 123, and generates a moving image using the read out decoration data, thereby generating a decorated moving image. Fig. 30 is a diagram showing an example of the commentary video after decoration. As shown in Fig. 30, a background image is inserted and a video is generated in which the commentary is explained by an avatar.

[0118] Furthermore, element data input by the user when generating decorative data may be used by the voice data generating unit 115. For example, by generating voice data using, as a parameter, the "voice of a man in his 30s," it becomes possible to generate voice data that matches the attributes of the generated avatar.

[0119] Although the present invention has been described above using various embodiments and modifications, the technical scope of the present invention is not limited to the scope described in the above embodiments and modifications. Various modifications or improvements can be made to the above embodiments without departing from the gist of the invention, and such modifications or improvements are also included in the technical scope of the present invention. Furthermore, the various embodiments and modifications may be combined as appropriate. Furthermore, the processing flow described in the above embodiment is also an example, and unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged within the scope of the present invention. [Explanation of symbols]

[0120] 1: Video generation system 10: Video generation device 10a: Video generation device 10b: Video generation device 100: Client terminal 110: Storage section 112: Information acquisition department 112a: Information acquisition section 113: Material Analysis Department 114: Explanation generation section 115: Audio data generation unit 116: Video generation unit 116a: Video generation unit 117: Editorial Department 118: Material data generation unit 119: Decoration data generation unit 120: Material data storage unit 121: Audio data storage unit 122: Video data storage unit 123: Decoration data storage unit

Claims

1. an information acquisition means for acquiring material data that serves as a source for generating a moving image and includes one or more still images; a material analysis means for detecting information for generating a description from the material data; An explanatory sentence generation means for generating a prompt for generating an explanatory sentence based on the detected information and generating an explanatory sentence by using the generated prompt in a trained model; a video generation means for generating a video using the material data and the commentary; Equipped with the material analysis means detects character information including character strings and accompanying information of the character strings from each still image constituting the material data, and adds pre-registered attribute information to each character string based on the accompanying information; the comment generation means generates the comment by creating a prompt using attribute information of the character string; The accompanying information includes at least one of the size of characters, the display position in the still image, the layout, and the color.

2. The video generation device described in Claim 1, wherein the material analysis means has attribute association information that associates each of the attribute information with its feature information, and adds attribute information to each of the character strings using the attribute association information and accompanying information of the character strings.

3. 2. The video generating device according to claim 1, further comprising an editing unit that provides the commentary in a state that allows the commentary to be edited by a user and edits the commentary based on an editing instruction from the user.

4. a voice data generating means for generating voice data of the commentary; The moving image generating device according to claim 1 , wherein the moving image generating means generates the moving image using the audio data.

5. providing a reading of the kanji included in the commentary in a state that the reading can be edited by a user, and editing means for editing the reading of the commentary based on an editing instruction from the user; The moving image generating device according to claim 4 , wherein the audio data generating means generates audio data of the commentary text with the pronunciation edited.

6. The video generation device of claim 4 is provided with an editing means that presents an editing screen to the user for editing the timing of breathing when the commentary is read aloud, and edits the timing of breathing in the audio data of the commentary based on the user's editing instructions.

7. The video generation device of claim 4 further comprises an editing means for presenting an editing screen to a user for editing the intonation of the explanatory text when read aloud, and editing the intonation of the audio data of the explanatory text based on the user's editing instructions.

8. the information acquisition means acquires information regarding a video playback time input by a user, The moving image generating device according to claim 1 , wherein the moving image generating means generates the moving image based on information about the moving image playback time.

9. The moving image generating device according to claim 8 , wherein the information regarding the moving image playback time is information regarding the playback time of the entire moving image or information regarding the playback time per page.

10. 10. The moving image generating device according to claim 9, wherein the information regarding the playback time per page can be set for each page.

11. 10. The moving image generating device according to claim 9, wherein the information regarding the playback time per page is given as a plurality of playback length levels, and the user is allowed to select any one of the levels.

12. a material data editing means for providing a user with a material data editing screen for generating new material data using one or more of the material data, and for generating the new material data based on an editing instruction from the user on the material data editing screen; 2. The video generating device according to claim 1, wherein said material data editing screen allows input of an editing instruction for arbitrarily combining one or more pages included in one or more of said material data.

13. a decoration data generating means for generating decoration data for decorating the material data, The moving image generating device according to claim 1 , wherein the moving image generating means generates a moving image using the decoration data.

14. the information acquisition means acquires element data relating to the decoration data from a user; The video generation device according to claim 13 , wherein the decoration data generation means generates at least one of a background image and an avatar image by providing a prompt including the acquired element data to a trained model.

15. an information acquisition step of acquiring material data that serves as a basis for generating a moving image and includes one or more still images; a material analysis step of detecting information for generating a description from the material data; an explanatory sentence generation step of generating a prompt for generating an explanatory sentence based on the detected information and generating an explanatory sentence by using the generated prompt in a trained model; a video generation step of generating a video using the material data and the commentary; The computer executes the material analysis step detects character information including character strings and accompanying information of the character strings from each still image constituting the material data, and adds pre-registered attribute information to each character string based on the accompanying information; the comment generation step generates the comment by creating a prompt using attribute information of the character string; The accompanying information includes at least one of the size of characters, the display position in the still image, the layout, and the color.

16. A program for causing a computer to function as the video generation device according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Voice electronic book

    JP1995129620A

  • Program, file generation method, information processing device, and information processing system

    JP2023100149A

  • Moving picture data creation system, and program for moving picture data creation system

    JP2025023363A

  • Artificial intelligence driven presenter

    US20240371089A1

  • Slide playback program, slide playback device, and slide playback method

    WO2023002300A1