Information processing device, method for controlling the information processing device, and control program for the information processing device

JP7912103B1Active Publication Date: 2026-08-27SOFTBANK CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025046337
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2026-08-27
Estimated Expiration
2045-03-21

Smart Images

  • Figure 0007912103000001_ABST
    Figure 0007912103000001_ABST
Patent Text Reader

Abstract

To provide an information processing device that generates documents from videos, reducing the effort required from the user. [Solution] An information processing device according to one embodiment of the present invention includes an instruction unit that inputs instructions to a generation AI model, instructions to convert the content of a video into text, instructions to identify a target frame which is an image frame of an important scene in the video from a plurality of image frames constituting the video, and instructions to convert the content of the important scene into text; an acquisition unit that acquires a first output result which is the output result of the generation AI model in response to a first prompt based on the plurality of instructions; an extraction unit that extracts a target frame from the video based on timestamp information of the target frame acquired from the first output result; and a document generation unit that generates a document based on the text of the content of the video and the text of the content of the important scene acquired from the first output result, and the target frame extracted by the extraction unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, a control method for the information processing apparatus, and a control program for the information processing apparatus.

Background Art

[0002] Conventionally, an electronic manual creation apparatus that creates work image data indicating work content based on text data describing the work content and outputs the text data and the work image data in association with each other is known (for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Means for Solving the Problems

[0004] An information processing apparatus according to an embodiment of the present invention includes: an instruction unit that inputs an instruction to textify the content of a video, an instruction to identify a target frame that is an image frame of an important scene in the video from a plurality of image frames constituting the video, and an instruction to textify the content of the important scene to a generation AI model; an acquisition unit that acquires a first output result that is an output result of the generation AI model according to a first prompt based on the plurality of instructions; an extraction unit that extracts the target frame from the video based on the time stamp information of the target frame acquired from the first output result; and a document generation unit that generates a document based on the text of the content of the video and the text of the content of the important scene acquired from the first output result and the target frame extracted by the extraction unit.

[0005] In the information processing apparatus according to an embodiment of the present invention, the document generation unit may generate a document in which an image of an important scene based on the target frame is associated with the text of the content of the important scene.

[0006] In an information processing device according to one embodiment of the present invention, the generating AI model processes image frames sampled at predetermined intervals from among a plurality of image frames constituting a video when identifying a target frame in response to a first prompt, the extraction unit extracts a predetermined number of additional image frames that are consecutive in time series before and after the extracted target frame from among a plurality of image frames constituting a video when the target frame extracted by the extraction unit does not meet predetermined criteria, the instruction unit inputs the additional image frames and a second prompt based on an instruction to identify the image frame among the additional image frames that meet predetermined criteria as a new target frame to the generating AI model, and the document generation unit may use the new target frame obtained from the second output result, which is the output result in response to the second prompt, when generating a document.

[0007] An information processing device according to one embodiment of the present invention may further include an output unit that stores documents generated by a document generation unit in a predetermined storage unit and outputs information that enables outputting the documents stored in the predetermined storage unit to a user's communication terminal.

[0008] In an information processing device according to one embodiment of the present invention, the first prompt and the second prompt may be realized by a single prompt.

[0009] In an information processing device according to one embodiment of the present invention, the generated AI model may be a large-scale multimodal model.

[0010] A control method for an information processing device according to one embodiment of the present invention involves the information processing device inputting an instruction to transcribe the content of a video into a generating AI model, an instruction to identify a target frame which is an image frame of an important scene in the video from among a plurality of image frames constituting the video, and an instruction to transcribe the content of the important scene into text; obtaining a first output result which is the output result of the generating AI model in response to a first prompt based on the plurality of instructions; extracting a target frame from the video based on the timestamp information of the target frame obtained from the first output result; and generating a document based on the text of the video content and the text of the content of the important scene obtained from the first output result, and the target frame extracted in the extraction step.

[0011] A control program for an information processing device according to one embodiment of the present invention provides the information processing device with the following functions: inputting instructions to a generation AI model to transcribe the content of a video into text, to identify a target frame which is an important scene in the video from among a plurality of image frames constituting the video, and to transcribe the content of the important scene into text; acquiring a first output result which is the output result of the generation AI model in response to a first prompt based on the plurality of instructions; extracting a target frame from the video based on the timestamp information of the target frame acquired from the first output result; and generating a document based on the text of the video content and the text of the important scene acquired from the first output result, and the target frame extracted by the extraction function. [Brief explanation of the drawing]

[0012] [Figure 1] Figure 1 is a schematic diagram of the information processing system configuration according to one embodiment of the present invention. [Figure 2] Figure 2 shows an example of a sequence between a user terminal, a server, and a generation AI system in one embodiment of the present invention. [Figure 3] Figures 3(a) and 3(b) are schematic diagrams illustrating one embodiment of the present invention. [Figure 4]Figures 4(a) and 4(b) are schematic diagrams illustrating one embodiment of the present invention. [Figure 5] Figure 5 shows an example of a sequence between a user terminal, a server, and a generation AI system in one embodiment of the present invention. [Figure 6] Figure 6 is an example of a functional block diagram of a server (information processing device) according to one embodiment of the present invention. [Figure 7] Figure 7 shows an example of the control flow of a server according to one embodiment of the present invention. [Modes for carrying out the invention]

[0013] Hereafter, an embodiment of the invention described herein (also referred to as the present invention) will be explained using the figures. Note that the figures are examples only, and the present invention is not limited to those shown in the figures. For example, the number of user terminals (communication terminals), servers (information processing devices), database servers, artificial intelligence (AI) systems, sequence diagrams, flowcharts, images, and documents shown are examples only, and the present invention is not limited to these.

[0014] In recent years, with the widespread use of smartphones and tablets, the use of photos, videos, and audio recordings taken with these devices in business operations has increased. For example, using video and audio manuals can clearly and easily demonstrate work procedures. However, when the operation procedure is complex and requires a long time to complete, watching the video every time for confirmation is inefficient. Also, different users may miss or overlook certain points, and video alone cannot guarantee the accuracy of the viewer's understanding. For these reasons, there is a certain demand for manuals that convert video data into text. Here, as described in Patent Document 1 above, there is a known electronic manual creation device that creates work image data showing the work content based on text data describing the work content, and outputs the text data and work image data in association. However, the technology in Patent Document 1 requires the user to input text data, which is time-consuming.

[0015] In contrast, according to one embodiment of the present invention, a document consisting of accurate text that conforms to the content of a video and images of points of interest can be generated based on video data using generative AI, which has seen remarkable development in recent years. In this case, the user is not required to perform tasks such as text input or image placement, thus realizing a highly usable document generation system.

[0016] <System Configuration> Figure 1 shows an example of the configuration of an information processing system according to one embodiment of the present invention. The information processing system 600 may be a system that converts the content of a video into text and generates a document. Furthermore, according to one embodiment of the present invention, a document including images contained in the video may be generated. Hereafter, one embodiment of the present invention will be described using as an example a case in which the video is a recording of a processing procedure, such as machine operation, equipment assembly, or cooking procedure, and a manual is generated as a document. However, the present invention is not limited to this and may be applicable to videos with various contents, such as story summaries or sports commentary.

[0017] The information processing system 600 may include a server (information processing device) 100, a user communication terminal (user terminal) 200, a generation AI model system 300, and a database server 400. These may be connected to each other via a network 500. The network 500 may include a wireless network or a wired network. Specifically, for example, the network 500 may be a wireless LAN (WLAN), a wide area network (WAN), LTE (Long Term Evolution), 4th generation communication (4G), 5th generation communication (5G), and 6th generation communication (6G) or later mobile communication systems. However, the network 500 is not limited to these examples and may also be, for example, Bluetooth (registered trademark), an optical fiber line, etc. Furthermore, the network 500 may be a combination of these.

[0018] Server 100 may be able to execute various processes related to the document generation service realized by the information processing system 600. The various processes related to the document generation service may include a process of generating a document conforming to the content of the video from the video. In FIG. 1, only one server 100 is shown, but it is not limited thereto. That is, each function described as being provided in server 100 may be realized by a plurality of servers. Also, server 100 may be, for example, a distributed server system that cooperates by communicating via a network, or a so-called cloud server. That is, server 100 is not limited to a physical server and may include a virtual server by software. Further, each function described as being performed by server 100 may be provided by a cloud-based platform operating via network 500.

[0019] User terminal 200 may be a communication terminal of a user who uses the document generation service. The user may transmit a video to be documented to server 100 from user terminal 200. A user interface for transmitting the video to server 100 may be provided to user terminal 200 by an application or program for using the document generation service. The video may include audio or may not include audio. In FIG. 1, a notebook computer is shown as user terminal 200, but user terminal 200 may be any terminal that can realize the functions described in each embodiment.

[0020] The generation AI model system 300 may be a system that provides the function of a generation AI model via the network 500. In one embodiment of the present invention, the generation AI model provided by the generation AI model system 300 may be a multimodal generation AI model. The multimodal generation AI model is an artificial intelligence system constructed by integratively processing multiple different types of information such as text, images, and voices through deep learning. Examples of the multimodal generation AI model include "Gemini (registered trademark)" by Google (registered trademark) and "ChatGPT GTP-4 (registered trademark)" by OpenAI. Note that the server 100 may be provided with the execution infrastructure of the generation AI model.

[0021] The generation AI model system 300 may execute processing based on a prompt transmitted from the server � 00. The prompt may be text for inputting instructions or questions to the AI.

[0022] The database server 400 may store (hold) various types of information (data) used in the information processing system 600. In FIG. 1, only one database server 400 is shown separately from the server 100, but it may be integrated with the server 100. The database server 400 may, for example, temporarily store videos, images, documents, etc. when generating a document by the generation AI model system 300, or store the generated document.

[0023] <Document generation process> With reference to FIGS. 2 and 3, a document generation process according to an embodiment of the present invention will be described. FIG. 2 is an example of a sequence among the user terminal 200, the server <00, and the generation AI system 300. FIGS. 3(a) and (b) are diagrams for explaining the outline of the document generation process.

[0024] The user terminal 200 may send the video to be documented to the server 100 (step S10). The server 100 may input the following to the generating AI model: an instruction to convert the content of the video received from the user terminal 200 into text, an instruction to identify a target frame that is an important scene in the video from among multiple image frames that make up the video, and an instruction to convert the content of the important scene into text (step S11). Input to the generating AI model may mean sending it to the generating AI model system 300. A prompt (first prompt) based on these multiple instructions may be input to the generating AI model. Here, the term "first prompt" is not limited to the input of multiple instructions to the generating AI model as a single prompt. That is, the above multiple instructions may be combined into a single prompt and input to the generating AI model, or they may be input to the generating AI model in multiple steps. Furthermore, the input of the video to be documented to the generating AI model may be at any time, as long as the generating AI model can recognize that it is the video to be processed by the prompts. Moreover, the input of the video to the generating AI model may be performed in any manner. For example, video data transmitted from the user terminal 200 may be provided to the generating AI model, or information about the location where the video is stored (such as a URL (Uniform Resource Locator) or a website address) may be provided to the generating AI model. Subsequently, the generating AI model system 300 may use the generating AI model to generate text of the video content, identify important scenes within the video, and generate text (captions) of the identified important scenes (step S12).

[0025] Furthermore, "important moments" may be specified in the first prompt as, for example, points in the operating procedure that are prone to errors, parts where the product's functions are particularly prominent, or points where a warning is needed. In addition, the first prompt may instruct users to determine the important moments from the context, rather than just using words such as "caution," "important," or "point."

[0026] The generation AI system 300 may send an output result (first output result), which is the output result of the generation AI model in response to the first prompt, to the server 100 (step S13). The first output result may include timestamp information of the target frame. The server 100 may extract the target frame from the video based on the timestamp information of the target frame obtained from the first output result. That is, the server 100 may generate a capture of an important scene (step S14).

[0027] Here, using Figure 3(a), we will explain how to identify important scenes in a video and how to generate captures. In Figure 3, a video showing the procedure for opening a delivery box is used as an example. The generating AI model may identify target frames, which are image frames of important scenes in the video, from among multiple image frames (image frames 11 to 15) that make up the video. Here, the generating AI model may transcribe the content of the video into text in response to the first prompt, and from the context, image frame 11 may be identified as the target frame. Furthermore, image frame 13, which relates to the unlocking operation, and image frame 15, which shows the package inside the opened delivery box, may also be identified as target frames. The generating AI model may also transcribe the content of important scenes into text in response to the first prompt. That is, the generating AI model may generate captions for important scenes. In the example in Figure 3(a), captions have been generated for target frames 11, 13, and 15, respectively. Server 100 may extract the target frames that have been identified as important scenes from the video. That is, Server 100 may acquire captures of important scenes. Furthermore, the first output result of the AI ​​model generated by the first prompt may include timestamp information of the target frame as information about the target frame. That is, the first output result may include the timestamps of the target frame being "00:01", "00:03", and "00:05". Therefore, in the example in Figure 3(a), target frames 11, 13, and 15 may be captured from the video.

[0028] Returning to Figure 2, the server 100 may generate a document (step S15). That is, the server 100 may generate a document based on the text of the video content and the text of important scenes (captions) obtained from the first output result, and the capture of the target frame. Then, the server 100 receives a request from the user terminal 200 (step S16) and may provide the generated document to the user terminal 200 (step S17).

[0029] Figure 3(b) shows an example of a generated document. Document 21 is a manual titled "How to Open a Delivery Box" and may include text of the video content. Note that the figure is just an example and the present invention is not limited thereto. For example, in the example figure, the procedure is explained by text in the first half and images are included in the second half. However, the procedure may be explained by combining text and images from the beginning.

[0030] Thus, according to one embodiment of the present invention, a document containing images of important scenes in a video may be generated from the video. In this case, the user does not need to perform operations such as entering text or selecting images to include in the document, thus providing a highly usable document generation system that reduces the user's workload.

[0031] The documents generated by server 100 may be stored in a designated storage unit, for example, in a database server 400. Server 100 may then output information to the user terminal 200 that allows the document stored in the designated storage unit to be accessed, in response to a request from the user terminal 200. This information may include, for example, a URL from which the document can be downloaded or viewed.

[0032] As shown in Figure 3(b), in one embodiment of the present invention, document 21 may consist of a capture of an important scene and its caption. That is, server 100 may generate a document that associates an image of an important scene based on the target frame with text describing the content of that important scene. This makes the document easier to understand.

[0033] Furthermore, some generating AI models, in identifying the target frame in response to the first prompt described above, process image frames sampled at predetermined intervals from among the multiple image frames that make up the video. This will be explained using Figure 4.

[0034] Figures 4(a) and (b) are schematic diagrams representing the image frames in a video in chronological order. That is, in Figures 4(a) and (b), the video consists of image frames T1, P1~P5, T2, P6... Here, it is assumed that the generating AI model can only process image frames T1, T2, T3, and T4, which are sampled at predetermined intervals and shown by thick lines. Therefore, when identifying the target frame of an important scene based on the first prompt, the determination will be made only based on image frames T1, T2, T3, and T4, which are shown by thick lines. In this case, although image frame T2, shown by the dashed line 31, is identified as the target frame, phenomena may occur that make it unsuitable for inclusion in the document, such as the product being obscured by a hand.

[0035] In contrast, according to one embodiment of the present invention, among the multiple image frames constituting a video, image frames that are not subject to processing by the generating AI model may be included as targets for determining important scenes. That is, as explained using Figure 4(b), if the target frame T2 does not meet a predetermined criterion, a target frame to replace the target frame T2 may be extracted from a predetermined number of additional image frames P4, P5, P6, P7, which are consecutive in the time series before and after the target frame T2, among the multiple image frames T1, P1~P5, T2, P6... constituting the video. In the example of Figure 4(b), the image frame P7 shown by the dashed line 32 may be identified as a new target frame to replace the image frame T2.

[0036] This will be explained using the sequence diagram in Figure 5. Note that Figure 5 shows the steps that follow steps S10 to S14 in Figure 4, and Figure 5 begins with step S14.

[0037] The server 100 may send a prompt to the generating AI model system 300 to determine whether the capture of the important scene (i.e., the target frame) acquired in step S14 meets predetermined criteria (step S21). The generating AI model system 300 may determine whether the target frame meets predetermined criteria, i.e., the quality of the capture (step S22), and output the determination result to the server 100 (step S23). The predetermined criteria may be conditions that allow determination of whether the part of interest in the target frame is clearly visible. For example, if the part of interest is hidden by an obstacle, the text is illegible, or the image is blurred, it may be determined that the predetermined criteria are not met. The determination of whether the target frame meets predetermined criteria may be performed by the server 100 using image recognition processing.

[0038] If the server 100 determines that the quality of the target frame is good, that is, that the target frame meets a predetermined standard (OK in step S24), it may proceed to step S29 and generate the document. Note that the document generation is the same as in step S15 in Figure 2, so the explanation is omitted. If the server 100 determines that the target frame does not meet a predetermined standard (NG in step S24), it may extract a predetermined number of additional image frames that are consecutive in time series before and after the target frame from among the multiple image frames that make up the video (step S25). The server 100 may then input the additional image frames and a prompt (second prompt) based on an instruction to identify the image frame among the additional image frames that meet the predetermined standard as a new target frame (step S26). The generation AI model system 300 may, in response to the second prompt, identify a new target frame from the additional image frames and generate a caption for that target frame (step S27). Note that the processing in step S27 may be the same as in step S12 in Figure 2. Subsequently, the generating AI system 300 may send the output result corresponding to the second prompt (second output result) to the server 100 (step S28).

[0039] Server 100 may generate a document using the new target frame (step S29). Note that the document generation is the same as in step S15 in Figure 2, so the explanation is omitted. Then, Server 100 receives a request from user terminal 200 (step S30) and may provide the generated document to user terminal 200 (step S31).

[0040] Thus, according to one embodiment of the present invention, even when there are limitations in the generating AI model, images more suitable for the document are selected. Therefore, it is possible to generate documents with higher accuracy.

[0041] In the above, we described the case in step S27 where a caption for a new target frame is generated. However, generating a caption is not mandatory, and in the document, the one generated in step S12 in Figure 2 may be used as the caption for the new target frame.

[0042] <Structure> Figure 6 will be used to explain the hardware and functional configuration of server 100.

[0043] <server> (1) Server hardware configuration The server 100 may include a control unit 110, a communication unit 120, an input / output unit 130, and a storage unit 170.

[0044] The control unit 110 is typically a processor, and may include a central processing unit (CPU), a microprocessing unit (GPU), a microprocessor, etc., and may be implemented by logic circuits (hardware) or dedicated circuits formed on an integrated circuit chip (IC (Integrated Circuit) chip, LSI (Large Scale Integration)), etc.

[0045] The communication unit 120 may be implemented as hardware such as a NIC or network adapter, communication software, or a combination thereof. The communication unit 120 may send and receive various types of data to and from the receiving terminal 20 and the user terminal 200 via the network 500.

[0046] The input / output unit 130 may include an input device for inputting various operations to the server 100, and an output device for outputting processing results processed by the server 100. The input device may include, for example, hardware keys such as a touch panel, touch display, or keyboard, a pointing device such as a mouse, a camera, or a microphone. The output device may output processing results processed by the control unit 110. The output device may include, for example, a display, touch panel, or speaker.

[0047] The storage unit 170 stores various programs and data necessary for the server 100 to operate. The storage unit 170 may include, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), flash memory, etc. The storage unit 170 may also include memory that provides a workspace for the control unit 110.

[0048] (2) Server Functional Configuration The server 100 may include, as functions implemented by the control unit 110, an instruction unit 111, an acquisition unit 112, an extraction unit 113, a document generation unit 114, and an output unit 115. Note that, among the functional units shown in Figure 6, those not essential in each embodiment may be omitted. Furthermore, the functions or processing of each functional unit may be implemented by machine learning or AI to the extent feasible.

[0049] The instruction unit 111 may input to the generating AI model an instruction to transcribe the content of the video into text, an instruction to identify a target frame that is an important scene in the video from among multiple image frames that make up the video, and a first prompt based on the instruction to transcribe the content of the important scene into text. This is as explained in steps S11 and S12 of Figure 2. The instruction unit 111 may also input to the generating AI model an additional image frame and a second prompt based on an instruction to identify an image frame from the additional image frame that meets predetermined criteria as a new target frame. This is as explained in steps S24 to S27 of Figure 5.

[0050] The acquisition unit 112 may acquire the output results (first output result, second output result) of the generated AI model corresponding to the prompts (first prompt, second prompt).

[0051] The extraction unit 113 may extract a target frame from the video based on the timestamp information of the target frame obtained from the first output result. Furthermore, if the target frame does not meet predetermined criteria, the extraction unit 113 may extract a predetermined number of additional image frames from among the multiple image frames constituting the video that are consecutive in time series before and after the extracted target frame. This is explained using Figure 4.

[0052] The document generation unit 114 may generate a document based on the text of the video content and the text of important scenes (captions) obtained from the first output result, and the target frame extracted by the extraction unit 113. In this case, the document generation unit 114 may generate the document by associating the image of the important scene based on the target frame with the text of the important scene. This is explained using Figure 3. In addition, when generating the document, the document generation unit 114 may use a new target frame obtained from the second output result, which is the output result corresponding to the second prompt. This is explained in steps S26 to S29 in Figure 5.

[0053] The output unit 115 may output information that allows the document generated by the document generation unit 114 and stored in a predetermined storage unit to the user's communication terminal 200.

[0054] <Server control flowchart> The control method for the server 100 described above will be explained using the flowchart in Figure 7. First, the instruction unit 111 of the server 100 inputs a video, a first prompt instructing the generation AI model to transcribe the content of the video into text, to identify a target frame which is an important scene in the video from among multiple image frames that make up the video, and to transcribe the content of the important scene into text (step S11). Next, the acquisition unit 112 may acquire a first output result which is the output result of the generation AI model corresponding to the first prompt (step S12). The extraction unit 113 may extract the target frame from the video based on the timestamp information of the target frame acquired from the first output result (step S13). The document generation unit 114 may generate a document based on the text of the video content and the text of the important scene content acquired from the first output result, and the extracted target frame (step S14).

[0055] The present invention has been described based on various drawings and embodiments, but it should be noted that those skilled in the art will find it easy to make various modifications and alterations based on this disclosure. Therefore, it should be noted that these modifications and alterations are within the scope of the present invention. For example, the functions included in each component, step, etc., can be rearranged in a logically consistent manner, and multiple components or steps, etc., can be combined into one or divided. Furthermore, the configurations shown in the above embodiments may be combined as appropriate. For example, each component described as being provided by server 100 may be implemented by multiple servers in a distributed manner. Furthermore, the functions described as being performed by server 100 may be performed by the generation AI model system 300. Also, the functions described as being performed by the generation AI model system 300 may be performed by server 100.

[0056] For example, in the above description, the first prompt and the second prompt were described as separate prompts. However, these prompts may be input to the generative AI model as a single prompt, or they may be input to the generative AI model as three or more prompts.

[0057] The programs of each embodiment of this disclosure may be provided stored in a storage medium readable by the information processing device. The storage medium is a “non-temporary tangible medium” capable of storing programs. The programs include, for example, software programs and information processing device programs. When each functional unit of the server 100 as an information processing device is implemented by software, the server 100 functions as an instruction unit 111, an acquisition unit 112, an extraction unit 113, a document generation unit 114, and an output unit 115 by having the processor execute a program loaded into memory.

[0058] The storage medium may, where appropriate, include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable storage medium, or two or more suitable combinations thereof. The storage medium may, where appropriate, be volatile, non-volatile, or a combination of volatile and non-volatile.

[0059] Furthermore, each embodiment of this disclosure can also be realized in the form of a data signal embedded in a carrier wave, where the program is embodied by electronic transmission. The programs of this disclosure may be implemented using, for example, scripting languages ​​such as JavaScript® and Python®, or languages ​​such as C, Go, Swift®, Koltin®, and Java®.

[0060] According to each aspect of this disclosure described above, by providing users with a more user-friendly work environment, we can contribute to achieving Sustainable Development Goal (SDG) 11, "Make cities and human settlements inclusive, safe, resilient and inclusive, safe [Explanation of Symbols]

[0061] 100 Servers (Information Processing Devices) 110 Control Unit 111 Instruction section 112 Acquisition Department 113 Extraction part 114 Document Generation Unit 115 Output section 120 Communications Department 130 Input / output section 170 Storage section 200 User terminals (communication terminals) 300 Generative AI Model Systems 400 Database Servers 500 Networks 600 Information Processing Systems

Claims

1. An instruction unit inputs instructions to a generating AI model, including instructions to transcribe the content of a video into text, instructions to identify a target frame that is an important scene in the video from among multiple image frames that make up the video, based on the context of the video's content, and instructions to transcribe the content of the important scene into text. An acquisition unit that acquires a first output result, which is the output result of the generated AI model in response to a first prompt based on the plurality of instructions, An extraction unit that extracts the target frame from the video based on the timestamp information of the target frame obtained from the first output result, A document generation unit generates a document based on the text of the video content and the text of the important scenes obtained from the first output result, and the target frame extracted by the extraction unit. Equipped with, The generation AI model, in identifying the target frame in response to the first prompt, processes image frames sampled at predetermined intervals from among the plurality of image frames constituting the video, If the target frame extracted by the extraction unit does not meet a predetermined criterion that the important part of the scene is clearly visible, the extraction unit extracts a predetermined number of additional image frames from among the multiple image frames constituting the video that are not used in the processing by the generation AI model in response to the first prompt, and that are consecutive in time series before and after the extracted target frame. The instruction unit inputs the additional image frames and a second prompt based on an instruction to identify the additional image frames that satisfy the predetermined criteria as new target frames to the generating AI model. The document generation unit generates a document in which, for each of the multiple important scenes, an image of the important scene based on the target frame and text of the content of the important scene generated by the generation AI model are arranged in association, and in generating the document, the new target frame obtained from the second output result, which is the output result in response to the second prompt, is used.

2. The document generated by the document generation unit is stored in a predetermined storage unit. The aforementioned information processing device is The system further includes an output unit that outputs information capable of outputting the document stored in the predetermined storage unit to the user's communication terminal. The information processing apparatus according to claim 1.

3. The first prompt and the second prompt are realized by one prompt. The information processing apparatus according to claim 1.

4. The aforementioned generative AI model is a large-scale multimodal model. The information processing apparatus according to claim 1.

5. Information processing device, The steps include inputting instructions to a generating AI model, such as instructions to transcribe the content of a video into text, instructions to identify a target frame that is an important scene in the video from among multiple image frames that make up the video, based on the context of the video's content, and instructions to transcribe the content of the important scene into text. The steps include obtaining a first output result, which is the output result of the generated AI model in response to a first prompt based on the plurality of instructions, The steps include: extracting the target frame from the video based on the timestamp information of the target frame obtained from the first output result; A step of generating a document based on the text of the video content and the text of the important scenes obtained from the first output result, and the target frame extracted in the extraction step, Execute, The generation AI model, in identifying the target frame in response to the first prompt, processes image frames sampled at predetermined intervals from among the plurality of image frames constituting the video, In the extraction step, if the extracted target frame does not meet a predetermined criterion that the important part of the scene is clearly visible, then a predetermined number of additional image frames that are consecutive in time series before and after the extracted target frame, and which are not used in the processing by the generation AI model in response to the first prompt, are extracted from among the multiple image frames that make up the video. In the input step described above, the additional image frames and a second prompt based on an instruction to identify the image frames among the additional image frames that satisfy the predetermined criteria as new target frames are input to the generating AI model. The step of generating the document is a control method for an information processing device, in which, for each of the multiple important scenes, a document is generated in which an image of the important scene based on the target frame and text of the content of the important scene generated by the generation AI model are arranged in association, and in generating the document, the new target frame obtained from the second output result, which is the output result in response to the second prompt.

6. In an information processing device, A function to input instructions to a generating AI model, including instructions to transcribe the content of a video into text, instructions to identify a target frame that is an important scene in the video from among multiple image frames that make up the video, based on the context of the video's content, and instructions to transcribe the content of the important scene into text. A function to acquire a first output result, which is the output result of the generated AI model in response to a first prompt based on the plurality of instructions, A function to extract the target frame from the video based on the timestamp information of the target frame obtained from the first output result, A function to generate a document based on the text of the video content and the text of the important scenes obtained from the first output result, and the target frame extracted by the extraction function, To make it happen, The generation AI model, in identifying the target frame in response to the first prompt, processes image frames sampled at predetermined intervals from among the plurality of image frames constituting the video, The extraction function, if the extracted target frame does not meet a predetermined criterion that the important part of the scene is clearly visible, extracts a predetermined number of additional image frames from among the multiple image frames constituting the video that are not used in the processing by the generation AI model in response to the first prompt, and that are consecutive in time series before and after the extracted target frame. The input function inputs the additional image frames and a second prompt based on an instruction to identify the image frames among the additional image frames that satisfy the predetermined criteria as new target frames into the generating AI model. The function for generating the document generates a document in which, for each of the multiple important scenes, an image of the important scene based on the target frame and text of the content of the important scene generated by the generation AI model are arranged in association, and in generating the document, the new target frame obtained from the second output result, which is the output result in response to the second prompt, is used. This is a control program for an information processing device.

Citation Information

Patent Citations

  • Video understanding method and device

    CN116935287A

  • Electronic manual preparation device, electronic manual preparation method, and electronic manual preparation program

    JP2006215986A

  • Thumbnail display method and information recording and reproducing device

    JP2007329732A

  • Information processing device, information processing method and information processing program

    JP2023043782A

  • JPP7550949B