Program, information processing system, and information processing method

The system uses semantic vectors from frame images and text to identify key frames in videos, addressing the challenge of conventional methods by accurately extracting important images based on content relevance.

JP2025169664APending Publication Date: 2025-11-14KONICA MINOLTA INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024074592
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-02
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Conventional methods fail to accurately extract important frame images from videos, as these images often have minimal pixel value changes, especially at scene boundaries.

Method used

A system that utilizes semantic vectors derived from both frame images and text, such as audio or user-entered text, to calculate similarities and identify key frame images based on predetermined conditions.

Benefits of technology

Effectively extracts important frame images relevant to the video content, ensuring accurate representation of the video's key moments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025169664000001_ABST
    Figure 2025169664000001_ABST
Patent Text Reader

Abstract

To provide a program, an information processing system, and an information processing method which can appropriately extract an important frame image in a moving image.SOLUTION: The program causes a computer to function as acquisition means which acquires a plurality of first semantic vectors generated on the basis of a plurality of frame images in a moving image and at least one second semantic vector generated on the basis of a text representing contents of the moving image, similarity calculation means which calculates a similarity between each of the plurality of first semantic vectors and the at least one second semantic vector, and extraction means which specifies a first semantic vector for which a similarity satisfying a prescribed condition has been calculated from among the plurality of first semantic vectors and extracts, from among the plurality of frame images, a frame image used for generation of the specified first semantic vector.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a program, an information processing system, and an information processing method. [Background technology]

[0002] Conventionally, there are known techniques for extracting thumbnail images that represent a video from among multiple frame images that make up the video, or extracting representative parts of the video to generate a shortened video (for example, see Patent Document 1). In such techniques, frame images in which pixel values ​​have changed significantly are detected as frame images that correspond to scene divisions in the video, and these are used as thumbnail images or to determine division positions in the video. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-33417 Summary of the Invention [Problem to be solved by the invention]

[0004] However, important frame images that represent a video are often included in the middle of each scene, where pixel values ​​change little. Therefore, frame images that correspond to scene boundaries are not necessarily important frame images in the video. As such, the above-mentioned conventional technology has the problem of being unable to properly extract important frame images in a video.

[0005] An object of the present invention is to provide a program, an information processing system, and an information processing method that can appropriately extract important frame images from a moving image. [Means for solving the problem]

[0006] In order to achieve the above object, the invention of the program described in claim 1 is: Computer, an acquisition means for acquiring a plurality of first semantic vectors generated based on a plurality of frame images of a video, and at least one second semantic vector generated based on text representing the content of the video; a similarity calculation means for calculating a similarity between each of the plurality of first semantic vectors and each of the at least one second semantic vector; an extraction means for identifying a first semantic vector from among the plurality of first semantic vectors for which the degree of similarity has been calculated and which satisfies a predetermined condition, and extracting a frame image used to generate the identified first semantic vector from among the plurality of frame images; Function as.

[0007] The invention described in claim 2 is the program described in claim 1, The text is audio text obtained by converting the audio of the video.

[0008] The invention described in claim 3 is the program described in claim 1, The text is either text entered by a user, text obtained by converting audio separate from the audio of the video, or text obtained by performing a predetermined analysis process on the video.

[0009] The invention described in claim 4 is the program described in claim 1, The obtaining means obtains, for each sentence included in the text, the second semantic vector generated based on the sentence.

[0010] The invention described in claim 5 is the program described in claim 1, The acquisition means acquires the plurality of first semantic vectors generated based on a plurality of image texts representing the contents of each of the plurality of frame images.

[0011] The invention described in claim 6 is the program described in claim 1, the acquiring means acquires at least one type of additional semantic vector generated based on information that represents the content of the video, the information being different from both the plurality of frame images and the text; the similarity calculation means calculates a similarity between each combination of n types of semantic vectors consisting of the first semantic vector, the second semantic vector, and the at least one type of additional semantic vector to generate an n-dimensional similarity map; The extraction means identifies the first semantic vector for which the similarity that satisfies the predetermined condition is calculated from among the multiple similarities in the n-dimensional similarity map.

[0012] The invention described in claim 7 is the program described in claim 1, The predetermined condition is satisfied when the calculated similarities are arranged in descending order and the similarity is within a predetermined number of positions from the top.

[0013] The invention described in claim 8 is the program described in claim 1, The extraction means identifies the first semantic vector for which the degree of similarity that satisfies the predetermined condition is calculated within each part of the moving image divided by a predetermined method.

[0014] The invention described in claim 9 is the program described in claim 8, The extraction means acquires a segment position of the video identified based on the content of the text, and identifies the portion of the video based on the segment position.

[0015] In order to achieve the above object, the invention of the information processing system described in claim 10 is as follows: an acquisition means for acquiring a plurality of first semantic vectors generated based on a plurality of frame images of a video and at least one second semantic vector generated based on text representing the content of the video; a similarity calculation means for calculating a similarity between each of the plurality of first semantic vectors and each of the at least one second semantic vector; an extraction means for identifying a first semantic vector from among the plurality of first semantic vectors for which the degree of similarity has been calculated and which satisfies a predetermined condition, and extracting a frame image used to generate the identified first semantic vector from among the plurality of frame images; Equipped with.

[0016] In order to achieve the above object, the invention of the information processing method described in claim 11 is as follows: 1. A computer-implemented information processing method, comprising: an acquisition step of acquiring a plurality of first semantic vectors generated based on a plurality of frame images of the video and at least one second semantic vector generated based on text representing the content of the video; a similarity calculation step of calculating a similarity between each of the plurality of first semantic vectors and each of the at least one second semantic vector; an extraction step of identifying a first semantic vector from among the plurality of first semantic vectors for which the degree of similarity has been calculated and which satisfies a predetermined condition, and extracting a frame image used to generate the identified first semantic vector from among the plurality of frame images; Includes. [Effects of the Invention]

[0017] According to the present invention, important frame images in a moving image can be appropriately extracted. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 is a block diagram showing a configuration of a document generation system. [Figure 2] 10 is a flowchart of a document generation process. [Figure 3] FIG. 10 is a diagram showing a document generation screen. [Figure 4] FIG. 10 is a diagram illustrating a process for generating image text data. [Figure 5] FIG. 10 is a diagram illustrating a process for generating voice text data. [Figure 6] FIG. 10 is a diagram showing a document generation screen on which chapter headings are displayed. [Figure 7] FIG. 10 is a diagram showing a document generation screen on which a main text is displayed. [Figure 8] 10 is a flowchart showing a control procedure for illustration extraction processing. [Figure 9] FIG. 10 is a diagram illustrating a process of converting into a first semantic vector. [Figure 10] FIG. 10 is a diagram illustrating a process of converting into a second semantic vector. [Figure 11] FIG. 10 is a diagram showing a similarity map. [Figure 12] FIG. 10 is a diagram illustrating a method for extracting a first semantic vector. DETAILED DESCRIPTION OF THE INVENTION

[0019] Hereinafter, embodiments of the present invention will be described with reference to the drawings, but the scope of the invention is not limited to the illustrated examples.

[0020] (Document Generation System Configuration) FIG. 1 is a block diagram showing the configuration of a document generation system 1 (information processing system) according to an embodiment of the present invention. The document generation system 1 includes a terminal device 10 and a cloud computing system 100. The terminal device 10 and the cloud computing system 100 are communicably connected via a communication network such as the Internet. The document generation system 1 generates electronic documents (hereinafter simply referred to as "documents") for users of the terminal device 10, and provides a service for saving and viewing the documents. Hereinafter, this service will be referred to as a "document generation service." The documents may be, for example, manuals, instruction manuals, or documents containing know-how, but are not limited to these. In this embodiment, an example will be described in which a manual for a coffee machine is generated using the document generation system 1.

[0021] The terminal device 10 is, for example, a notebook PC, a desktop PC, a tablet terminal, a smartphone, etc. The terminal device 10 includes a CPU (Central Processing Unit) 11, a memory 12, a storage unit 13, a display unit 14, an operation unit 15, and a communication unit 16. The components of the terminal device 10 are connected via a data transmission path such as a bus.

[0022] The CPU 11 is a processor that controls the operation of each part of the terminal device 10 by executing various processes in accordance with a program 131 stored in the storage unit 13. The memory 12 is, for example, a random access memory (RAM), which provides a working memory space for the CPU 11 and stores temporary data. The storage unit 13 is configured with a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or the like. The storage unit 13 stores the program 131 and video data 132 used to generate manuals. The video data 132 may be generated by an imaging unit (not shown) provided in the terminal device 10, or may be obtained from a source external to the terminal device 10. The program 131 includes a web browser. The CPU 11 displays various information and documents on the web browser on the display unit 14 based on data received from the cloud computing system 100.

[0023] The display unit 14 includes a display device such as a liquid crystal display. The display unit 14 displays various information and documents in accordance with control signals and image data input from the CPU 11. The operation unit 15 includes input means such as a mouse, keyboard, touch panel, and operation buttons. When an operation is performed on the input means, the operation unit 15 outputs an operation signal corresponding to the operation to the CPU 11. The communication unit 16 performs communication operations in accordance with a predetermined communication standard. Through this communication operation, the communication unit 16 transmits and receives data to and from the service providing server 20 of the cloud computing system 100.

[0024] The cloud computing system 100 includes a service providing server 20, a document generation server 30, a video analysis module 40, and a large-scale language model 50. Hereinafter, the large-scale language model 50 will be abbreviated as "LLM (Large Language Models) 50." The service providing server 20 and the document generation server 30 are virtual servers. Specifically, the cloud computing system 100 includes multiple physical servers (not shown) that are interconnected for communication. In the cloud computing system 100, these multiple physical servers create a virtual environment in which multiple virtual servers can be logically constructed. The service providing server 20 and the document generation server 30 are virtual servers constructed in such a virtual environment. The virtual CPU, virtual memory, and virtual storage unit of the virtual server are realized by logically dividing or integrating the CPU, memory, storage unit, etc. that constitute the physical server, respectively.

[0025] The service providing server 20 includes a virtual CPU 21, a virtual memory 22, and a virtual storage unit 23. The virtual CPU 21 executes various processes related to the provision of the document generation service in accordance with a program 231 stored in the virtual storage unit 23. The virtual memory 22 provides a working memory space for the virtual CPU 21 and stores temporary data. The virtual storage unit 23 stores the program 231, document data 232 generated in the document generation service, and the like.

[0026] In response to a request from the terminal device 10, the virtual CPU 21 executes various processes related to the provision of a document generation service, and transmits the process results, generated document data 232, and the like to the terminal device 10. The processes executed by the virtual CPU 21 include a process of receiving information specifying the specifications and content of the document to be generated from the terminal device 10, a process of causing the document generation server 30 to generate document data 232 based on the received information, a process of displaying a document related to the generated document data 232 on the display unit 14 of the terminal device 10, and a process of managing the generated document data 232. As described above, the information specifying the specifications and content of the document that the service providing server 20 receives from the terminal device 10 includes video data 132.

[0027] The document generation server 30 includes a virtual CPU 31, a virtual memory 32, and a virtual storage unit 33. The virtual CPU 31 executes various processes related to the generation of document data 232 according to a program 331 stored in the virtual storage unit 33. The virtual CPU 31 executes various processes according to the program 331, thereby functioning as an acquisition unit, a similarity calculation unit, and an extraction unit. The virtual CPU 31 as an acquisition unit acquires a first semantic vector 3351 and a second semantic vector 3352 (described later). The virtual CPU 31 as a similarity calculation unit calculates the similarity between the first semantic vector 3351 and the second semantic vector 3352 to generate a similarity map 336. The virtual CPU 31 as an extraction unit extracts frame images appropriate as illustrations for the manual based on the similarity calculation results. The details of these processes performed by the virtual CPU 31 will be described in detail later.

[0028] The virtual memory 32 provides a working memory space for the virtual CPU 31 and stores temporary data. The virtual storage unit 33 stores a program 331 and various data used to generate the document data 232. Specifically, the virtual storage unit 33 stores video data 332, image text data 333, audio text data 334, semantic vector data 335, a similarity map 336, and the like. Of these, the video data 332 is data transmitted from the terminal device 10 via the service providing server 20, and its content is the same as the video data 132. The video data 332 includes frame image data 3321 containing image data of multiple frame images of the video, and audio data 3322 related to the audio of the video. The contents of the image text data 333, audio text data 334, semantic vector data 335, and the similarity map 336 will be described later.

[0029] The process for generating the document data 232 executed by the virtual CPU 31 includes a process for causing the video analysis module 40 to analyze the video data 332, and a process for causing the LLM 50 to generate the chapters and body of the document.

[0030] The video analysis module 40 executes an analysis process on the video data and outputs the results. The analysis process by the video analysis module 40 can be invoked and executed from any virtual server in the cloud computing system 100. Like a virtual server, the video analysis module 40 includes a virtual CPU, virtual memory, and virtual storage unit (not shown), which together constitute an AI (artificial intelligence) that analyzes the video data. This AI has a machine learning model trained to extract and output analytical information from the video data. For example, the video analysis module 40 recognizes and analyzes the audio contained in the audio data 3322 of the input video data 332, converts it into text, and outputs it. This process is referred to as "transcription" of the video audio. Furthermore, in this specification, the text obtained by transcribing the video audio is referred to as "audio text." Furthermore, the video analysis module 40 analyzes each frame image contained in the frame image data 3321 of the video data 332 and outputs text representing the content of the frame image. In this specification, the text representing the content of the frame image is referred to as "image text." Image text is also called caption.

[0031] The LLM50 is a language model pre-trained using massive amounts of data and deep learning technology to assign probabilities to word sequences. In the pre-training using deep learning technology, the neural network model parameters are adjusted so that appropriate probabilities are assigned to word sequences. When a prompt, which is an input sentence that instructs the LLM50's operation, is input, the LLM50 estimates and outputs the word sequence that follows the prompt, i.e., a response sentence. Specifically, the LLM50 divides the input prompt into the smallest units called tokens and extracts the features of the tokens. The LLM50 constructs a response sentence by repeatedly deriving the probability of the tokens that follow the prompt based on the extracted features. Through this operation, the LLM50 can execute various tasks requested by the prompt. Tasks executed by the LLM50 in this embodiment include generating a document's chapter structure and body text based on input titles, speech text, etc. Hereinafter, determining the structure of a document consisting of multiple chapters and generating chapter titles for each chapter will be referred to as "organizing into chapters."

[0032] (Operation of document generation system) Next, the operation of the document generation system 1 will be described. Fig. 2 is a flowchart of the document generation process executed by each device of the document generation system 1 when the document generation system 1 generates a document to provide a document generation service. Fig. 2 shows the processes executed by the CPU 11 of the terminal device 10, the virtual CPU 21 of the service providing server 20, the virtual CPU 31 of the document generation server 30, the video analysis module 40, and the LLM 50, as well as the flow of data transmitted and received between each device. The document generation process can be broadly divided into a process for generating chapters for the manual (steps S1 to S12, Fig. 6), a process for generating the main text of the manual (steps S13 to S20, Fig. 7), and a process for extracting illustrations to be inserted into the manual from the frame image data 3321 (steps S21 to S27, Fig. 8).

[0033] When the document generation process is started, the CPU 11 of the terminal device 10 causes the display unit 14 to display the document generation screen 140 shown in Fig. 3 (step S1). More specifically, when a user performs an input operation on the operation unit 15 of the terminal device 10 to instruct the start of document generation, the CPU 11 transmits a request to start the document generation process to the service providing server 20. The virtual CPU 21 of the service providing server 20, which has received the start request, transmits data to the terminal device 10 for displaying the document generation screen 140 of Fig. 3 on the display unit 14. Based on the data, the CPU 11 causes the display unit 14 to display the document generation screen 140 of Fig. 3 on the web browser.

[0034] 3 displays an upload button 141 for registering video data 332 to be used in generating a manual, a text box 142 for inputting the title of the manual, and a configuration creation button 143 for instructing the creation of a manual configuration. When the upload button 141 is selected, a window (not shown) for selecting a video to register is displayed. By selecting video data 1322 in the window and then selecting a registration button (not shown), the video data 132 to be used in generating a manual can be registered. The CPU 11 transmits the registered video data 132 to the service providing server 20. The video data 132 in this embodiment is video data that demonstrates and explains how to make coffee using a coffee machine, accompanied by audio commentary.

[0035] The user also inputs the title of the manual to be generated in a text box 142. In FIG. 3, the title "How to brew coffee using an automatic coffee machine" is input.

[0036] When the video data 132 is registered using the upload button 141 and the configuration creation button 143 is selected with a title entered in the text box 142, the processing of steps S2 to S12 in Fig. 2 is executed to generate a chapter structure for the manual. First, the CPU 11 transmits the video data 132 and title data to the service providing server 20 (step S2). The virtual CPU 21 of the service providing server 20 transfers the received data to the document generation server 30 (step S3) and instructs the document generation server 30 to generate a chapter structure for the manual.

[0037] The virtual CPU 31 of the document generation server 30 stores the received video data 132 in the virtual memory unit 33 as video data 332, and transmits this video data 332 to the video analysis module 40 (step S4). The video analysis module 40 executes an analysis process of the received video data 332 (step S5). This analysis process includes a process of generating image text data 333 based on frame image data 3321 of the video data 332, and a process of transcribing audio data 3322 of the video data 332 to generate audio text data 334.

[0038] FIG. 4 is a diagram illustrating the process of generating image text data 333. The left side of FIG. 4 shows one frame image from the frame image data 3321 of the video data 332. The video analysis module 40 performs a predetermined image recognition process on this frame image to identify the type of object, the human action, and the like contained in the frame image. Based on the identification results, the video analysis module 40 generates image text representing the content of the frame image. In the example shown in FIG. 4, the frame image shows a person standing in front of a coffee machine placing a cup on the coffee machine. The video analysis module 40 analyzes this frame image and generates image text with the content "A person is placing a cup." The video analysis module 40 performs this process for each frame image and generates multiple image texts representing the content of each of the multiple frame images included in the frame image data 3321. The video analysis module 40 generates image text data 333 including these multiple image texts. In the image text data 333, each of the multiple image texts is registered in association with its position in the video related to the video data 332. The position in the video is represented by, for example, a frame number or the elapsed time from the start of the video. When the image text is the same across two or more consecutive frame images, these image texts may be combined into one in the image text data 333. In the image text data 333, the range of the video corresponding to the combined video text, i.e., the start point and end point of the range, may be registered in association with each other.

[0039] FIG. 5 is a diagram illustrating the process of generating audio-text data 334. The left side of FIG. 5 shows a video related to video data 332. The video analysis module 40 performs a predetermined speech recognition process on audio data 3322 of the video to identify the audio content, i.e., the words being spoken by the person. Based on the identification results, the video analysis module 40 converts the audio of the video into audio text representing the audio content. In the example shown in FIG. 4, a person in the video utters the sentence, "I hear a voice saying, 'Please put down the cup,' so I will make sure I have placed the cup." The video analysis module 40 converts this sentence into audio text. The video analysis module 40 performs this process across the entire video to generate audio text for each of the multiple sentences included in audio data 3322. The video analysis module 40 generates audio-text data 334 including the audio text of the multiple sentences. In the audio-text data 334, each of the audio texts of the multiple sentences is registered in association with a position in the video related to video data 332, such as the start position of the sentence. The position in the video is represented by, for example, a frame number or the elapsed time from the start of the video. In Fig. 5, the position of each voice text is represented by the elapsed time from the start of the video. The voice text of the voice text data 334 is one aspect of "text representing the content of the video."

[0040] The video analysis module 40 transmits the generated image text data 333 and audio text data 334 to the document generation server 30 (step S6 in FIG. 2). Note that the video analysis module that generates the image text data 333 and the video analysis module that generates the audio text data 334 may be provided separately.

[0041] The virtual CPU 31 of the document generation server 30 inputs the received voice text data 334 and the title data already received from the service providing server 20 into the LLM 50, causing the LLM 50 to generate a chapter structure for the manual (step S7). For example, the virtual CPU 31 inputs a prompt such as "Please organize the following text into chapters," along with the title and voice text data 334, to the LLM 50. In response, the LLM 50 divides the content of the voice text data 334 into multiple chapters and generates a chapter title for each chapter (step S8). The LLM 50 transmits information about the chapter structure to the document generation server 30 (step S9). The information about the chapter structure includes the text of the chapter title for each chapter. This information about the chapter structure is transmitted to the terminal device 10 via the document generation server 30 and the service providing server 20 (steps S10 and S11).

[0042] Based on the received chapter information, the CPU 11 of the terminal device 10 displays the chapter structure of the manual in the left half of the document generation screen 140, as shown in Fig. 6 (step S12). In Fig. 6, a chapter structure consisting of chapters 1 to 5 of the manual is generated, and the titles of each chapter are displayed in text boxes 144. The user can modify the chapter titles as necessary. The CPU 11 also displays a main text creation button 145 on the document generation screen 140 together with the chapter structure.

[0043] 6, when the main text creation button 145 is selected, an instruction to create the main text of the manual is transmitted from the terminal device 10 to the service providing server 20 (step S13). The virtual CPU 21 of the service providing server 20, which has received the instruction to create the main text, transmits the instruction to create the main text to the document generation server 30 (step S14). Furthermore, if the chapter structure has been changed in the text box 144, the changed, confirmed chapter information is also transmitted to the service providing server 20 and the document generation server 30.

[0044] The virtual CPU 31 of the document generation server 30 inputs the voice text data 334, the title, and the determined chapter information to the LLM 50, causing the LLM 50 to generate the main text of the manual (step S15). For example, the virtual CPU 31 inputs a prompt such as "Please create the main text of the manual based on the text below" along with the voice text data 334, the title, and the determined chapter information to the LLM 50. In response to this, the LLM 50 generates the main text of the manual (step S16). Note that the voice text data 334 may be omitted, and the LLM 50 may generate the main text based on the title and the determined chapter information. The LLM 50 transmits the generated main text information to the document generation server 30 (step S17). The main text information is transmitted to the terminal device 10 via the document generation server 30 and the service providing server 20 (steps S18 and S19).

[0045] Based on the received information about the main text, the CPU 11 of the terminal device 10 displays the main text 146 of the manual in the right half of the document generation screen 140, as shown in FIG. 7 (step S20). The CPU 11 also displays an illustration setting button 147 for each chapter in the main text 146. The user can set an illustration to be inserted into a desired chapter by selecting the illustration setting button 147 of the desired chapter. In FIG. 7, the illustration setting buttons 147 for chapters 2, 3, and 4 have been selected. In response to the selection of the illustration setting button 147, the CPU 11 displays a frame for an illustration area 148, into which the illustration will be inserted, at the right end of the corresponding chapter of the main text 146. Furthermore, when one or more illustration setting buttons 147 are selected, the CPU 11 displays an illustration creation button 149 below the main text 146.

[0046] 7, when the illustration creation button 149 is selected, steps S21 to S25 are executed to extract an appropriate illustration from the frame image data 3321. First, the CPU 11 of the terminal device 10 transmits an illustration extraction instruction to the service providing server 20 (step S21). In response to this, the virtual CPU 21 of the service providing server 20 transmits the illustration extraction instruction to the document generation server 30 (step S22). Having received the illustration extraction instruction, the virtual CPU 31 of the document generation server 30 executes illustration extraction processing (step S23).

[0047] FIG. 8 is a flowchart showing the control procedure for the illustration extraction process. When the illustration extraction process starts, the virtual CPU 31 converts each image text in the image text data 333 into a first semantic vector 3351 (step S231). FIG. 9 is a diagram illustrating the conversion process into the first semantic vector 3351. As shown in FIG. 9, the virtual CPU 31 converts each image text in the image text data 333 into a first semantic vector 3351 having X vector elements according to a predetermined conversion rule. The number of vector elements X of the first semantic vector 3351 is, for example, several tens to several hundreds, but may be several thousand or more. The conversion rule for the first semantic vector 3351 can be determined arbitrarily as long as the content of the image text is reflected in the first semantic vector 3351. For example, the conversion process into the first semantic vector 3351 may include a process of converting each word contained in the image text, such as "person," "cup," and "installation," into a vector having X elements according to a predetermined conversion rule, and a process of adding up the elements of the resulting vectors.

[0048] Next, the virtual CPU 31 converts each speech text of the speech text data 334 into a second semantic vector 3352 (step S232). FIG. 10 is a diagram illustrating the conversion process into the second semantic vector 3352. As shown in FIG. 10, the virtual CPU 31 converts each speech text of the speech text data 334 into a second semantic vector 3352 having the same number of elements X as the first semantic vector 3351 in accordance with a predetermined conversion rule. The conversion into the second semantic vector 3352 is performed in accordance with the same conversion rule as the conversion rule into the first semantic vector 3351.

[0049] The process of converting image text into first semantic vector 3351 is one aspect of the process of acquiring first semantic vector 3351. The process of converting voice text into second semantic vector 3352 is one aspect of the process of acquiring second semantic vector 3352. Steps S231 and S232 correspond to an "acquisition step." Note that virtual CPU 31 may input image text data 333 to a predetermined vector conversion module provided outside document generation server 30, thereby converting image text into first semantic vector 3351 and acquiring first semantic vector 3351. Furthermore, virtual CPU 31 may input voice text data 334 to the above-mentioned vector conversion module, thereby converting voice text into second semantic vector 3352 and acquiring second semantic vector 3352.

[0050] Next, the virtual CPU 31 calculates the similarity between each of the multiple first semantic vectors 3351 and each of the multiple second semantic vectors 3352 to generate a similarity map 336 (step S233). Step S233 corresponds to a "similarity calculation step." FIG. 11 is a diagram showing the similarity map 336. In FIG. 11, each image text in the image text data 333 is written in multiple columns. Each of these image texts corresponds to one first semantic vector 3351. In FIG. 11, each first semantic vector 3351 is represented as "VA1" to "VAn." t1 to tn shown next to each first semantic vector 3351 indicate the position (time point) of each image text in the video. Also, in FIG. 11, the audio text of each sentence included in the audio text data 334 is written in multiple rows. Each of these audio texts corresponds to one second semantic vector 3352. 11, the secondary semantic vectors 3352 are denoted as "VB1" to "VBm." t1 to tm shown next to each secondary semantic vector 3352 indicate the position (time point) of each audio text in the video.

[0051] The numerical value written in the cell where the column of the image text intersects with the row of the audio text represents the similarity between the first semantic vector 3351 corresponding to the image text and the second semantic vector 3352 corresponding to the audio text. Here, the similarity is normalized so that the minimum value is 0 and the maximum value is 100. The higher the similarity, the greater the similarity between the first semantic vector 3351 and the second semantic vector 3352, i.e., the greater the semantic similarity between the image text and the audio text. In the example shown in FIG. 11 , for example, for the audio text "Placing a cup in front of the coffee machine," the similarity to the image text with the similar semantic content "A person is placing a cup" is high at "80." On the other hand, the similarity to this audio text with the image text that does not contain the word "cup," such as the image text with the content "A person is pressing a button on the coffee machine" or "A person is throwing something in the trash," is low.

[0052] The similarity is calculated based on, for example, the product (inner product) of the first semantic vector 3351 and the second semantic vector 3352, the Euclidean distance, the cosine distance, the angle between the vectors, or the maximum difference between the components of the first semantic vector 3351 and the second semantic vector 3352, with the smaller the value, the greater the similarity. For example, the similarity may be calculated by normalizing the reciprocal of the above value. The actual data of the similarity map 336 only needs to associate the similarity with any combination of the first semantic vector 3351 and the second semantic vector 3352, and the data of the audio text and the image text may be omitted.

[0053] Returning to FIG. 8, when the generation of the similarity map 336 is completed, the virtual CPU 31 identifies the segmentation positions of the video corresponding to the chapter boundaries of the manual (step S234). The bars shown in the upper part of FIG. 12 represent the start and end points of the video related to the video data 332. Furthermore, P1 to P5 represent portions (partial video) of the video corresponding to chapters 1 to 5 of the manual shown in FIG. 7, respectively. Hereinafter, any one of the partial videos P1 to P5 will be referred to as a "partial video P." The segmentation positions T2 to T5 are the start points of the partial videos P2 to P5, respectively, and correspond to the segmentation positions when segmenting the video into partial videos P1 to P5. In step S234, the virtual CPU 31 identifies the segmentation positions T2 to T5 based on the positions of the audio text in the video corresponding to the respective chapter boundaries when the audio text data 334 is organized into chapters by the LLM 50, for example.

[0054] Returning to FIG. 8, the virtual CPU 31 selects one chapter for which an instruction to extract an illustration has been given (step S235). In the example shown in FIG. 7, the instruction to extract an illustration has been given for chapters 2 to 4, so the virtual CPU 31 selects one of these chapters. Next, the virtual CPU 31 extracts a first semantic vector 3351, the similarity of which satisfies a predetermined condition, from a portion of the similarity map 336 corresponding to the selected chapter (hereinafter referred to as a "partial map") (step S236). Step S236 corresponds to an "extraction step."

[0055] FIG. 12 is a diagram illustrating a method for extracting the first semantic vector 3351. In step S236, the virtual CPU 31 refers to the partial map of the similarity map 336 corresponding to the chapter selected in step S235. This partial map is a portion of the similarity map 336 in which the position of the first semantic vector 3351 in the video and the position of the second semantic vector 3352 in the video are both included in the time range of the partial video P corresponding to the selected chapter. For example, FIG. 12 shows a partial map 3362 corresponding to chapter 2 and a partial map 3363 corresponding to chapter 3. In the partial map 3362, the times tn1 to tn3 of the first semantic vector 3351 and the times tm1 to tm3 of the second semantic vector 3352 belong to the time range T2 to T3 of the partial video P2. In other words, the partial map 3362 is a portion of the similarity map 336 that represents the similarity between the image text and the audio text belonging to the partial video P2 corresponding to chapter 2. Furthermore, times tn4 to tn6 of the first semantic vector 3351 and times tm4 to tm6 of the second semantic vector 3352 in the partial map 3363 belong to the time range T3 to T4 of the partial video P3. In other words, the partial map 3363 is a part of the similarity map 336 that represents the similarity between the image text and audio text that belong to the partial video P3 corresponding to chapter 3.

[0056] If the virtual CPU 31 selects Chapter 2 in step S235, in step S236, the virtual CPU 31 identifies, from the partial map 3362, a first semantic vector 3351 whose similarity satisfies a predetermined condition. Here, the predetermined condition is satisfied if the first semantic vector 3351 is within a predetermined number of positions from the top when the multiple similarities in the partial map 3362 are arranged in descending order. For example, if the predetermined number is set to "1," the virtual CPU 31 identifies the first semantic vector 3351 corresponding to the highest similarity in the partial map 3362. If the predetermined number is set to "2 or more," the virtual CPU 31 identifies a predetermined number of first semantic vectors 3351 corresponding to the predetermined number of similarities in the partial map 3362, starting from the highest. In this way, by using a method for identifying first semantic vectors 3351 with high similarity in the partial map 3362, it is possible to identify first semantic vectors 3351 corresponding to frame images that are highly relevant to the audio content in the partial video P2.

[0057] Next, the virtual CPU 31 determines an illustration for the selected chapter from among the frame images corresponding to the extracted first semantic vector (step S237). For example, if one first semantic vector 3351 is identified for chapter 2 in step S236, the virtual CPU 31 extracts a frame image used to generate the first semantic vector 3351 and determines it as the illustration for chapter 2. If two or more first semantic vectors 3351 are identified for chapter 2, the virtual CPU 31 extracts two or more frame images used to generate each of the two or more first semantic vectors 3351. The virtual CPU 31 then selects one frame image from the two or more extracted frame images using a predetermined method and determines it as the illustration for chapter 2. A method for selecting one frame image may, for example, display the two or more extracted frame images on the display unit 14 of the terminal device 10 and allow the user to select a desired frame image.

[0058] In the partial map, the range of the first semantic vector 3351 corresponds to the time range of the partial video P, and the second semantic vector 3352 may include the second semantic vector 3352 for the entire range of the video. In other words, the partial map may be a similarity map 336 in which the range of the first semantic vector 3351 is narrowed. By using such a partial map, it is possible to extract, as illustrations, frame images from the partial video P that are highly relevant to the audio content of the entire video.

[0059] Next, the virtual CPU 31 determines whether all chapters for which an illustration extraction instruction has been issued have been selected (step S238). If it determines that any chapter has not been selected ("NO" in step S238), the virtual CPU 31 returns the process to step S235 and selects the next chapter. If it determines that all chapters for which an illustration extraction instruction has been issued have been selected ("YES" in step S238), the virtual CPU 31 ends the illustration extraction process and returns the process to the document generation process of FIG. 2.

[0060] When the illustration extraction process is completed, the virtual CPU 31 transmits illustration information relating to the extracted illustrations to the service providing server 20 (step S24). The virtual CPU 21 of the service providing server 20 transmits the received illustration information to the terminal device 10 (step S25). Here, the illustration information includes, for example, the frame number of the extracted frame image for each chapter for which extraction of an illustration was instructed. Alternatively, the illustration information may include the image data itself of the extracted frame image.

[0061] Based on the received illustration information, the CPU 11 of the terminal device 10 displays the frame image extracted as an illustration in the illustration area 148 of each chapter in FIG. 7. As a result, the completed manual is displayed on the document generation screen 140 (step S26). Furthermore, the virtual CPU 21 of the service providing server 20 stores the document data 232 of the completed manual in the virtual memory unit 23 (step S27). When step S27 is completed, each device of the document generation system 1 ends the document generation process.

[0062] (Variation 1) Next, a first modification of the above embodiment will be described. Below, differences from the above embodiment will be described, and a description of commonalities with the above embodiment will be omitted.

[0063] In the above embodiment, audio text transcribed from audio data 3322 is used as text representing the content of the video, and this audio text is converted into second semantic vector 3352. However, the text representing the content of the video is not limited to audio text. For example, the text representing the content of the video may be text explaining the content of the video that is input by the user into terminal device 10. Furthermore, the text representing the content of the video may be text obtained by transcribing audio explaining the content of the video, separate from the audio of the video. Furthermore, the text representing the content of the video may be text obtained by performing a predetermined analysis process on the video, for example, text summarizing the content of the video using AI including LLM.

[0064] (Variation 2) Next, a second modification of the above embodiment will be described. Differences from the above embodiment will be described below, and a description of commonalities with the above embodiment will be omitted. The second modification may be combined with the first modification.

[0065] In the above embodiment, as shown in FIG. 11, a two-dimensional similarity map 336 of two types of semantic vectors is used. However, instead, an n-dimensional similarity map representing the similarity of n types of semantic vectors may be used. Here, n is a natural number greater than or equal to 3. The n types of semantic vectors consist of a first semantic vector 3351, a second semantic vector 3352, and at least one additional semantic vector. The additional semantic vector is information representing the content of the video related to the video data 332, and is generated based on information (hereinafter referred to as "additional information") that is different from both the frame image data 3321 and the audio data 3322. The additional information may be, for example, the title of a manual entered in the text box 142 of FIG. 3. The additional information may also be various texts representing the content of the video, as exemplified in the first modification. The additional information may also be additional video information relating to additional video data obtained by capturing the same subject as the video data 332 from a different angle when the video data 332 was captured. The additional video information may be frame image data and / or audio data of the additional video data. When there is such additional information, the virtual CPU 31 converts the additional information into an additional semantic vector with the number of elements X in the same manner as the first semantic vector 3351 and the second semantic vector 3352.

[0066] In step S233 of FIG. 8, the virtual CPU 31 calculates the similarity of each combination of n types of semantic vectors, including the first semantic vector 3351, the second semantic vector 3352, and at least one type of additional semantic vector, to generate an n-dimensional similarity map 336. The n-dimensional similarity map 336 is a map in which the similarity of the n types of semantic vectors corresponding to each position in an n-dimensional space is registered. The method for calculating the similarity in the n-dimensional similarity map 336 is not particularly limited as long as the method increases the similarity as the semantic content of each piece of information corresponding to the n types of semantic vectors becomes closer. For example, a method may be used in which a set of semantic vectors is selected from the n types of semantic vectors, a process of calculating the similarity is performed for all sets (e.g., three sets for three types of semantic vectors), and the obtained similarities are averaged. Alternatively, a method may be used in which a value such as the product (inner product), Euclidean distance, or cosine distance of a set of semantic vectors selected from the n types of semantic vectors is calculated for all sets, and the reciprocal of a representative value (e.g., average) of the obtained values ​​is normalized. Furthermore, the elements of n types of semantic vectors may be multiplied together, and the resulting X products may be added together to calculate the similarity, similar to the inner product.

[0067] In step S236 of FIG. 8 , the virtual CPU 31 extracts a first semantic vector 3351 whose similarity satisfies a predetermined condition from among multiple similarities in an n-dimensional partial map, which is a part of the n-dimensional similarity map 336. Note that if the additional information corresponding to the additional semantic vector is information on additional video data, i.e., if there are two pieces of video data that can be referenced in generating a manual, frame images can be extracted from each of the two pieces of video data for one similarity that satisfies the predetermined condition. In this case, it is sufficient to determine in advance which frame image of the video data will be used as an illustration, or a method for determining this. For example, the video data 332 and the additional video data to be used as an illustration may be determined in advance. Alternatively, both frame images of the video data 332 and the additional video data may be used as illustrations. Alternatively, frame images of the video data 332 and the additional video data may be displayed on the display unit 14 of the terminal device 10, allowing the user to select a desired frame image.

[0068] (effect) As described above, the program 331 according to this embodiment causes the virtual CPU 31 of the document generation server 30 (as a computer) to function as an acquisition unit, a similarity calculation unit, and an extraction unit. The virtual CPU 31 as an acquisition unit generates and acquires a plurality of first semantic vectors 3351 generated based on a plurality of frame images of a video and a plurality of second semantic vectors 3352 generated based on the audio text of the audio-text data 334 representing the content of the video. The virtual CPU 31 as a similarity calculation unit calculates the similarity between each of the plurality of first semantic vectors 3351 and each of the plurality of second semantic vectors 3352. The virtual CPU 31 as an extraction unit identifies a first semantic vector 3351 from the plurality of first semantic vectors 3351 whose calculated similarity satisfies a predetermined condition. Furthermore, the virtual CPU 31 as an extraction unit extracts the frame image used to generate the identified first semantic vector 3351 from the plurality of frame images. When the similarity between the first semantic vector 3351 and the second semantic vector 3352 satisfies a predetermined condition, the frame image corresponding to the first semantic vector 3351 and the audio text corresponding to the second semantic vector 3352 are highly related. Therefore, according to the method of this embodiment, it is possible to appropriately extract important frame images that are highly related to the audio content of a video. In other words, it is possible to appropriately extract frame images of scenes that correspond to the audio content of a video. Conventional methods for extracting frame images that correspond to scene boundaries were unable to extract important frame images in the middle of a scene, but according to the method of this embodiment, it is possible to appropriately extract frame images in such positions.

[0069] Furthermore, the virtual CPU 31 as an acquisition unit acquires a plurality of second semantic vectors 3352 generated based on audio text obtained by converting the audio of a video. The audio of a video represents the content of the video. Therefore, by using the similarity with the second semantic vectors 3352 obtained by converting the audio text, it is possible to appropriately extract important frame images that are highly relevant to the content of the video.

[0070] Furthermore, in the first modification, the virtual CPU 31 serving as an acquisition unit acquires a plurality of second semantic vectors 3352 generated based on any of text entered by a user, text obtained by converting a sound separate from the sound of the video data 332, or text obtained by performing a predetermined analysis process on the video data 332. Such text also represents the content of the video. Therefore, by using the similarity between the second semantic vectors 3352 obtained by converting such text, it is possible to appropriately extract important frame images that are highly relevant to the content of the video.

[0071] Furthermore, the virtual CPU 31 as an acquisition unit acquires, for each sentence of the speech text, a second semantic vector 3352 generated based on the sentence. This allows the content of one sentence of the speech text to be appropriately reflected in the second semantic vector 3352. By calculating the similarity between such second semantic vector 3352 and the first semantic vector 3351, it is possible to appropriately evaluate the degree of relevance between one sentence of the speech text and each frame image.

[0072] Furthermore, the virtual CPU 31 as an acquisition unit acquires a plurality of first semantic vectors 3351 generated based on a plurality of image texts that represent the content of each of a plurality of frame images. This allows the content of the frame images to be appropriately reflected in the first semantic vectors 3351.

[0073] In the second modification, the virtual CPU 31, which functions as an acquisition unit, acquires at least one additional semantic vector. The additional semantic vector is additional information representing the content of the video, and is generated based on additional information that is different from both the multiple frame images in the frame image data 3321 and the audio text in the audio text data 334. The virtual CPU 31, which functions as a similarity calculation unit, calculates the similarity between each combination of n types of semantic vectors, each consisting of a first semantic vector 3351, a second semantic vector 3352, and at least one additional semantic vector, to generate an n-dimensional similarity map 336. The virtual CPU 31, which functions as an extraction unit, identifies the first semantic vector 3351 for which a similarity that satisfies a predetermined condition is calculated, from among the multiple similarities in the n-dimensional similarity map 336. This makes it possible to extract frame images that are highly related to both the content of the audio text and the content of the additional information. This makes it possible to more appropriately extract important frame images.

[0074] Furthermore, the predetermined condition is met if the calculated similarities are arranged in descending order and the result is within a predetermined number of positions from the top. This allows a predetermined number of important frame images to be extracted. Furthermore, by setting the predetermined number to "1", the most important single frame image can be extracted.

[0075] Furthermore, the virtual CPU 31 as extraction means identifies, for each partial video P obtained by dividing the video data 332 using a predetermined method, a first semantic vector 3351 for which a similarity that satisfies a predetermined condition has been calculated within each partial video P. This makes it possible to extract important frame images for each partial video P. This makes it possible to perform processing such as extracting frame images suitable for illustrations for each of multiple chapters of a manual.

[0076] Furthermore, the virtual CPU 31 as an extraction means acquires the segmentation positions of the moving image identified based on the content of the audio text of the audio text data 334, and identifies the partial moving image P based on the segmentation positions. This makes it possible to appropriately segment the moving image based on the audio text data 334 and identify the partial moving image P.

[0077] The document generation system 1 according to this embodiment also includes a virtual CPU 31 functioning as an acquisition unit, a similarity calculation unit, and an extraction unit. The virtual CPU 31 as an acquisition unit generates and acquires a plurality of first semantic vectors 3351 generated based on a plurality of frame images of a video, and a plurality of second semantic vectors 3352 generated based on the audio text of audio-text data 334 representing the content of the video. The virtual CPU 31 as a similarity calculation unit calculates the similarity between each of the plurality of first semantic vectors 3351 and each of the plurality of second semantic vectors 3352. The virtual CPU 31 as an extraction unit identifies a first semantic vector 3351 among the plurality of first semantic vectors 3351 whose calculated similarity satisfies a predetermined condition. The virtual CPU 31 as an extraction unit also extracts, from the plurality of frame images, a frame image used to generate the identified first semantic vector 3351. This allows for appropriate extraction of important frame images highly relevant to the audio content of the video.

[0078] The information processing method according to this embodiment also includes an acquisition step, a similarity calculation step, and an extraction step. In the acquisition step, the virtual CPU 31 generates and acquires a plurality of first semantic vectors 3351 generated based on a plurality of frame images of the video, and a plurality of second semantic vectors 3352 generated based on the audio text of audio-text data 334 representing the content of the video. In the similarity calculation step, the virtual CPU 31 calculates the similarity between each of the plurality of first semantic vectors 3351 and each of the plurality of second semantic vectors 3352. In the extraction step, the virtual CPU 31 identifies a first semantic vector 3351 from the plurality of first semantic vectors 3351 whose calculated similarity satisfies a predetermined condition. In the extraction step, the virtual CPU 31 extracts, from the plurality of frame images, a frame image used to generate the identified first semantic vector 3351. This allows important frame images highly relevant to the audio content of the video to be appropriately extracted.

[0079] The present invention is not limited to the above-described embodiment, and various modifications are possible. For example, although the embodiment in which the service providing server 20 and the document generation server 30 are virtual servers has been exemplified, the present invention is not limited to this. The service providing server 20 and the document generation server 30 may be physical servers, i.e., independent servers that actually exist.

[0080] Although the example has been given in which the document generation server 30 is provided with a virtual CPU 31 that functions as all of the acquisition means, similarity calculation means, and extraction means, the present invention is not limited to this. Some or all of the acquisition means, similarity calculation means, and extraction means may be provided in separate virtual servers or separate physical servers.

[0081] Furthermore, the document generation server 30 may execute the processing that was executed by at least one of the video analysis module 40 and the LLM 50.

[0082] The service providing server 20 and the document generation server 30 may also be integrated.

[0083] Furthermore, in the above embodiment, speech text corresponding to one sentence in the speech text data 334 is converted into one second semantic vector 3352, but this is not limiting. For example, a group of speech for each predetermined time unit in the speech text data 334 may be converted into one second semantic vector 3352. Also, a portion of the speech text data 334 corresponding to one chapter may be converted into one second semantic vector 3352. Also, the entire speech text data 334 may be converted into one second semantic vector 3352. Therefore, at least one second semantic vector 3352 is sufficient.

[0084] Furthermore, although an example has been described in which the manual is organized into chapters by LLM50 using the audio text data 334, the method for generating the chapters of the manual is not limited to this. For example, the chapters of the manual may be determined by a method in which the points at which the pixel values ​​of frame images change significantly are used as dividing scenes in the video, and a chapter is provided for each scene.

[0085] In the above embodiment, the position for inserting an illustration is specified for each chapter of the manual, but this is not intended to be limiting. For example, the speech text of a certain sentence may be specified, and an illustration appropriate for this sentence may be extracted. In this case, image text corresponding to the first semantic vector 3351 whose similarity satisfies a predetermined condition may be extracted from the row corresponding to the specified speech text in the similarity map 336 of FIG. 11 or the partial map of FIG. 12. Alternatively, one illustration may be extracted for the entire manual. In this case, image text corresponding to the first semantic vector 3351 whose similarity satisfies a predetermined condition may be extracted using the entire similarity map 336 of FIG. 11, without using the partial map.

[0086] 2 is an example and can be modified as appropriate. For example, after generating the chapters, when instructing the generation of the main text, the insertion position of the illustration may be specified, and the illustration may be extracted at the same time as the generation of the main text. Also, although an example in which the main text is generated after the chapters have been generated has been described, instead, the chapters and the main text may be generated and displayed at the same time.

[0087] Although several embodiments of the present invention have been described, the scope of the present invention is not limited to the above-described embodiments, but includes the scope of the invention described in the claims and its equivalents. [Explanation of symbols]

[0088] 1 Document generation system (information processing system) 10 Terminal Equipment 11 CPU 20 Service provider servers 21 virtual CPUs 30 Document Generation Servers 31 Virtual CPU (acquisition means, similarity calculation means, extraction means) 332 video data 3321 Frame image data 3322 Audio Data 333 Image Text Data 334 Audio-Text Data 335 Semantic Vector Data 3351 First Semantic Vector 3352 Second Semantic Vector 336 Similarity Map 3362, 3363 Partial Map 40 Video Analysis Module 50 Large-Scale Language Models (LLM) 100 Cloud Computing Systems 140 Document Generation Screen

Claims

1. Computer, an acquisition means for acquiring a plurality of first semantic vectors generated based on a plurality of frame images of a video, and at least one second semantic vector generated based on text representing the content of the video; a similarity calculation means for calculating a similarity between each of the plurality of first semantic vectors and each of the at least one second semantic vector; an extraction means for identifying a first semantic vector from among the plurality of first semantic vectors for which the degree of similarity has been calculated and which satisfies a predetermined condition, and extracting a frame image used to generate the identified first semantic vector from among the plurality of frame images; A program that functions as a

2. The program according to claim 1 , wherein the text is audio text obtained by converting audio from the video.

3. The program of claim 1, wherein the text is either text entered by a user, text obtained by converting audio separate from the audio of the video, or text obtained by a predetermined analysis process of the video.

4. The program according to claim 1 , wherein the acquiring means acquires, for each sentence included in the text, the second semantic vector generated based on the sentence.

5. 2. The program according to claim 1, wherein said acquiring means acquires said plurality of first semantic vectors generated based on a plurality of image texts representing the contents of each of said plurality of frame images.

6. the acquiring means acquires at least one type of additional semantic vector generated based on information that represents the content of the video, the information being different from both the plurality of frame images and the text; the similarity calculation means calculates a similarity between each combination of n types of semantic vectors consisting of the first semantic vector, the second semantic vector, and the at least one type of additional semantic vector to generate an n-dimensional similarity map; the extraction means identifies the first semantic vector for which the similarity that satisfies the predetermined condition is calculated from among the plurality of similarities in the n-dimensional similarity map. The program according to claim 1.

7. 2. The program according to claim 1, wherein the predetermined condition is satisfied when the calculated similarities are arranged in descending order and the first similarity is within a predetermined number of similarities from the top.

8. The program according to claim 1 , wherein the extraction means identifies the first semantic vector for each part of the video divided in a predetermined manner, for which the similarity satisfying the predetermined condition has been calculated within each part.

9. The program according to claim 8 , wherein the extraction means acquires a segment position of the video identified based on the content of the text, and identifies the portion of the video based on the segment position.

10. an acquisition means for acquiring a plurality of first semantic vectors generated based on a plurality of frame images of a video and at least one second semantic vector generated based on text representing the content of the video; a similarity calculation means for calculating a similarity between each of the plurality of first semantic vectors and each of the at least one second semantic vector; an extraction means for identifying a first semantic vector from among the plurality of first semantic vectors for which the degree of similarity has been calculated and which satisfies a predetermined condition, and extracting a frame image used to generate the identified first semantic vector from among the plurality of frame images; An information processing system comprising:

11. 1. A computer-implemented information processing method, comprising: an acquisition step of acquiring a plurality of first semantic vectors generated based on a plurality of frame images of the video and at least one second semantic vector generated based on text representing the content of the video; a similarity calculation step of calculating a similarity between each of the plurality of first semantic vectors and each of the at least one second semantic vector; an extraction step of identifying a first semantic vector from among the plurality of first semantic vectors for which the degree of similarity has been calculated and which satisfies a predetermined condition, and extracting a frame image used to generate the identified first semantic vector from among the plurality of frame images; An information processing method including:

Citation Information

Patent Citations

  • Video processing device and program

    JP2014033417A