Information processing device, method, program, and system
The system uses a large-scale language model to automatically generate explanatory texts and summaries for video frames, addressing the challenges of summarizing video content and enhancing information retrieval.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- OPTIM
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-07
AI Technical Summary
Existing technologies face challenges in efficiently summarizing video content and associating motion and sound information with images, leading to increased capacity and difficulty in information retrieval.
A system utilizing a large-scale language model to generate explanatory texts for video frames and summarize videos based on user prompts, enabling easy grasping of video content by automatically selecting frames and generating summaries.
Enables users to easily understand video content through generated summaries, reducing the need to watch the entire video and improving information retrieval efficiency.
Smart Images

Figure JP2025037718_07052026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus, Method, Program, and System
[0001] The present disclosure relates to an information processing apparatus, method, program, and system.
[0002] In Patent Document 1, a technique for providing an image that facilitates grasping the motion and sound of an object represented in the image is described. For example, in Patent Document 1, an information processing apparatus receives moving image data composed of a plurality of frames and representing one or more objects (persons or things), and selects a characteristic frame from the moving image data based on the motion of the one or more objects or the sound from the one or more objects. The information processing apparatus selects a characteristic object from the one or more objects based on the motion of the one or more objects represented in the moving image data or the sound from the one or more objects. The information processing apparatus associates text information indicating at least one of the motion of the characteristic object and the sound from the characteristic object with the characteristic object and displays it on the image of the characteristic frame.
[0003] Japanese Patent Application Laid-Open No. 2015-073198
[0004] In Patent Document 1, text information indicating at least one of the motion of the characteristic object and the sound from the characteristic object is associated with the characteristic object and displayed on the image of the characteristic frame. In Patent Document 1, by checking the image, it is possible to grasp the motion and sound of the object represented in the image, but there are problems such as an increase in capacity and difficulty in searching for desired information in order to associate information with the image.
[0005] An object of the present disclosure is to make it possible to easily grasp the summary of a video.
[0006] A program for operating a computer comprising a processor and memory, the program causing the processor to perform the following steps: acquiring a video consisting of multiple frames; inputting a video or predetermined frames comprising a video and prompts for explaining the content of the frames comprising the video into a large-scale language model and causing the large-scale language model to output explanatory text for the frames; inputting multiple explanatory texts and prompts for outputting a summary of the video into the large-scale language model and causing the large-scale language model to output a summary of the video; and presenting the summary output by the large-scale language model to the user.
[0007] The video summary can be easily grasped.
[0008] This is a block diagram showing the overall configuration of System 1. This is a block diagram showing a functional configuration example of terminal device 10. This is a block diagram showing a functional configuration example of server 20. This is a diagram showing the data structure of a table. This is a diagram showing the data structure of a table. This is a diagram showing an example of the processing flow in System 1. This is a diagram showing an example of the screen of this disclosure. This is a block diagram showing the basic hardware configuration of computer 90.
[0009] The embodiments of this disclosure will be described below with reference to the drawings. In all the drawings illustrating the embodiments, common components are denoted by the same reference numerals, and repeated explanations are omitted. The following embodiments are not intended to unduly limit the content of this disclosure as described in the claims. Not all components shown in the embodiments are necessarily essential components of this disclosure. Also, each drawing is a schematic diagram and is not necessarily a strict illustration.
[0010] Furthermore, in the following description, "processor" refers to one or more processors. At least one processor is typically a microprocessor such as a CPU (Central Processing Unit), but may be another type of processor such as a GPU (Graphics Processing Unit). At least one processor may be single-core or multi-core.
[0011] Furthermore, at least one processor may be a broad-sense processor, such as a hardware circuit that performs some or all of the processing (e.g., an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit)).
[0012] Furthermore, in the following explanation, we may use expressions such as "xxx table" to describe information from which an output is obtained for a given input. This information can be data of any structure, or it can be a learning model such as a neural network that generates an output for a given input. Therefore, "xxx table" can be referred to as "xxx information."
[0013] Furthermore, in the following explanation, the configuration of each table is just an example; one table may be divided into two or more tables, or all or part of two or more tables may be a single table.
[0014] Furthermore, in the following explanation, the subject of the process may sometimes be "program," but since a program is executed by a processor and performs defined processes using the memory and / or interface as appropriate, the subject of the process may also be the processor (or a device such as a controller that has that processor).
[0015] The program may be installed on a device such as a computer, or it may reside on a program distribution server or a computer-readable (e.g., non-temporary) recording medium. Furthermore, in the following description, two or more programs may be implemented as a single program, or one program may be implemented as two or more programs.
[0016] Furthermore, in the following explanation, identification numbers are used as identification information for various objects, but other types of identification information (for example, identifiers including letters or symbols) may also be used.
[0017] Furthermore, in the following explanations, when describing similar elements without distinction, a reference code (or a common code among reference codes) may be used, and when describing similar elements with distinction, the element's identification number (or reference code) may be used.
[0018] Furthermore, in the following explanation, only control lines and information lines deemed necessary for the explanation are shown, and not all control lines and information lines in the product are necessarily shown. All components may be interconnected.
[0019] Each information processing device consists of a computer equipped with an arithmetic unit and a memory device. The basic hardware configuration of the computer and the basic functional configuration of the computer realized by said hardware configuration will be described later. For each of the server 20 and terminal device 10, explanations that overlap with the basic hardware configuration and basic functional configuration of the computer described later will be omitted.
[0020] <Outline> The system according to this embodiment generates descriptive text for multiple frames (e.g., video) in a time-series sequence according to predetermined rules using a large language model (LLM), and generates a summary of the multiple frames (e.g., a video summary) using the LLM based on the generated descriptive texts.
[0021] <1. Overall System Configuration Diagram> Figure 1 is a block diagram showing an example of the overall configuration of System 1. System 1 shown in Figure 1 includes, for example, a terminal device 10, a server 20, and an LLM system 30. The terminal device 10, the server 20, and the LLM system 30 are connected to each other via, for example, a network 80.
[0022] Figure 1 shows an example where System 1 includes two terminal devices 10, but the number of terminal devices 10 included in System 1 is not limited to two. System 1 may include one terminal device 10, or it may include three or more terminal devices 10.
[0023] Figure 1 shows an example where System 1 includes one LLM system 30, but the number of LLM systems 30 included in System 1 is not limited to one. System 1 may include two or more LLM systems 30.
[0024] Figure 1 shows an example where server 20 is independent of the LLM system 30, but server 20 may include the functions of the LLM system 30. In other words, server 20 may store the LLM.
[0025] In this embodiment, a collection of multiple devices may be treated as a single server. The method of allocating the multiple functions required to implement the server 20 according to this embodiment to one or more hardware can be appropriately determined in view of the processing capacity of each hardware and / or the specifications required for the server 20.
[0026] The terminal device 10 shown in Figure 1 is an information processing device used by users who utilize services provided by the server 20. For example, the terminal device 10 is an information processing device operated by a user who uses the document output service provided by the server 20 to create documents. The terminal device 10 can be implemented as, for example, a stationary PC (Personal Computer), a laptop PC, a head-mounted display, etc. Alternatively, the terminal device 10 may be a portable computer such as a smartphone or a tablet device.
[0027] The terminal device 10 includes a communication interface 12, an input device 13, an output device 14, a memory 15, storage 16, and a processor 19. The input device 13 is a device for receiving input operations from the user (e.g., a touch panel, touchpad, mouse or other pointing device, keyboard, etc.). The output device 14 is a device for presenting information to the user (display, speaker, etc.).
[0028] Server 20 is an information processing device that provides a service for generating video summaries. Server 20 is an information processing device implemented by, for example, a computer connected to network 80. As shown in Figure 1, Server 20 includes a communication IF 22, an input / output IF 23, a memory 25, a storage 26, and a processor 29. The input / output IF 23 functions as an interface to an input device for receiving input operations from the user and an output device for outputting information to the user.
[0029] The LLM system 30 is a system in which a large-scale artificial intelligence model (LLM) used, for example, in the field of natural language processing (NLP) has been constructed. By learning from large amounts of text data (web pages, books, articles, etc.), the LLM can understand patterns in human language and effectively perform natural language generation (NLG) tasks.
[0030] LLM is used in many NLP tasks, such as generating responses to specific questions, automatically generating text, summarizing text, translation, and sentiment analysis. It can also be used in a variety of applications, including education, entertainment, customer service, and product development. LLM is, for example, multimodal LLM. Examples of LLM include: • GPT-4 (registered trademark) (OpenAI Inc.) • Gemini (registered trademark) (Google Inc.) • StableLM (StableAI Inc.) • Llama3.2 (Meta Inc.)
[0031] The LLM system 30 inputs text data and prompts received from the server 20 into the LLM, and causes the LLM to output a response based on the input text data and prompts. The LLM system 30 then sends the response output from the LLM to the server 20.
[0032] The imaging device 31 is a device that receives light using a light-receiving element and outputs it as an image signal. The imaging device 31 is installed so that it can photograph the situation in a space (hereinafter referred to as the site) that the user wants to understand. Specifically, the imaging device 31 is installed in a position that allows it to see the entire site without any obstructions. For example, the imaging device 31 is installed at a high place in a corner of the site. If one imaging device 31 is not enough to capture the situation of the entire site, multiple devices are installed. The imaging device 31 is, for example, a surveillance camera installed in a busy shopping district, park, train station, office, school, etc. The imaging device 31 is, for example, a camera used to film surgery in a hospital.
[0033] Each information processing device consists of a computer equipped with an arithmetic unit and a memory device. The basic hardware configuration of the computer and the basic functional configuration of the computer realized by said hardware configuration will be described later. For each of the terminal device 10 and the server 20, explanations that overlap with the basic hardware configuration and basic functional configuration of the computer described later will be omitted.
[0034] <2. Configuration of the Terminal Device> Figure 2 is a block diagram showing an example of the functional configuration of the terminal device 10. As shown in Figure 2, the terminal device 10 includes a communication unit 120, an input device 13, an output device 14, an audio processing unit 17, a microphone 171, a speaker 172, a camera 160, a position information sensor 150, a storage unit 180, and a control unit 190. Each block included in the terminal device 10 is electrically connected, for example, by a bus.
[0035] The communication unit 120 performs processing such as modulation and demodulation processing for the terminal device 10 to communicate with other devices. The communication unit 120 performs transmission processing on the signal generated by the control unit 190 and transmits it to an external source (for example, the server 20). The communication unit 120 performs reception processing on the signal received from the external source and outputs it to the control unit 190.
[0036] The input device 13 is a device for a user operating the terminal device 10 to input instructions or information. The input device 13 can be implemented, for example, by a touch-sensitive device 131 on which instructions are input by touching the operating surface. If the terminal device 10 is a PC, the input device 13 may be implemented by a reader, keyboard, mouse, etc. The input device 13 converts the instructions input by the user into electrical signals and outputs the electrical signals to the control unit 190. The input device 13 may also include, for example, a receiving port that accepts electrical signals input from an external input device.
[0037] The output device 14 is a device for presenting information to the user operating the terminal device 10. The output device 14 is implemented, for example, by a display 141. The display 141 displays data according to the control of the control unit 190. The display 141 is implemented, for example, by an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display.
[0038] The audio processing unit 17 performs, for example, digital-to-analog conversion processing of the audio signal. The audio processing unit 17 converts the signal provided from the microphone 171 into a digital signal and provides the converted signal to the control unit 190. The audio processing unit 17 also provides the audio signal to the speaker 172. The audio processing unit 17 is implemented, for example, by an audio processing processor. The microphone 171 receives an audio input and provides the audio signal corresponding to the audio input to the audio processing unit 17. The speaker 172 converts the audio signal provided from the audio processing unit 17 into audio and outputs the audio to the outside of the terminal device 10.
[0039] Camera 160 is a device that receives light using a photodetector and outputs it as a shooting signal.
[0040] The position information sensor 150 is a sensor that detects the position of the terminal device 10, and is, for example, a GPS (Global Positioning System) module. The GPS module is a receiving device used in a satellite positioning system. In the satellite positioning system, signals from at least three or four satellites are received, and based on the received signals, the current position of the terminal device 10 on which the GPS module is mounted is detected. The position information sensor 150 may detect the current position of the terminal device 10 from the position of the wireless base station to which the terminal device 10 is connected.
[0041] The storage unit 180 is realized by, for example, the memory 15, the storage 16, etc., and stores data and programs used by the terminal device 10. The storage unit 180 stores, for example, user information 181.
[0042] The user information 181 includes, for example, information about the user who uses the terminal device 10. Information about the user includes, for example, the user's name, age, address, date of birth, contact information, etc.
[0043] The control unit 190 is realized by the processor 19 reading the program stored in the storage unit 180 and executing the instructions included in the program. The control unit 190 controls the operation of the terminal device 10. By operating according to the program, the control unit 190 exhibits the functions as the operation reception unit 191, the transmission / reception unit 192, and the presentation control unit 193.
[0044] The operation reception unit 191 performs processing for receiving an instruction or information input from the input device 13. Specifically, for example, the operation reception unit 191 receives an instruction or information input from the touch-sensitive device 131 or the like.
[0045] In addition, the operation reception unit 191 receives a voice instruction input from the microphone 171. Specifically, for example, the operation reception unit 191 receives a voice signal input from the microphone 171 and converted into a digital signal by the voice processing unit 17. The operation reception unit 191 obtains an instruction from the user, for example, by analyzing the received voice signal and extracting a predetermined noun.
[0046] The transmission / reception unit 192 performs processes for the terminal device 10 to transmit and receive data to and from an external device such as the server 20 according to a communication protocol. Specifically, for example, the transmission / reception unit 192 transmits information input by the user or an instruction from the user to the server 20. Also, the transmission / reception unit 192 receives information provided from the server 20.
[0047] The presentation control unit 193 controls the output device 14 to present the information provided from the server 20 to the user. Specifically, for example, the presentation control unit 193 causes the display 141 to display information regarding a document transmitted from the server 20. Also, the presentation control unit 193 causes the speaker 172 to output the information transmitted from the server 20.
[0048] <3. Functional Configuration of Server> Figure 3 is a diagram showing a functional configuration example of the server 20. As shown in Figure 3, the server 20 functions as a communication unit 201, a storage unit 202, and a control unit 203.
[0049] The communication unit 201 performs processes for the server 20 to communicate with an external device.
[0050] The storage unit 202 has, for example, a user information table 2021, a video log table 2022, etc. The tables stored in the storage unit 202 are not limited to these.
[0051] The user information table 2021 is a table that stores information about the user. Details will be described later.
[0052] The video log table 2022 is a table that stores information about the frames constituting a video. Details will be described later.
[0053] The control unit 203 is realized when the processor 29 reads a program stored in the memory unit 202 and executes instructions contained in the program. The program includes applications such as a web browser application. The program includes a programming language such as JavaScript® that is executed on the web browser application stored in the terminal device 10. By operating according to the program, the control unit 203 performs the functions indicated as the receive control module 2031, the transmit control module 2032, the service processing module 2033, and the presentation control module 2034.
[0054] The reception control module 2031 controls the process by which the server 20 receives signals from an external device according to a communication protocol.
[0055] The transmission control module 2032 controls the process by which the server 20 transmits signals to an external device according to a communication protocol.
[0056] The service processing module 2033 controls the process of prompting the LLM and generating a video summary.
[0057] The presentation control module 2034 controls the process of presenting information to the user.
[0058] The image analysis module 2035 detects a predetermined object, a predetermined action, or a combination thereof by analyzing a video.
[0059] <4. Data Structure> The data structure of the tables stored by the server 20 is described below. Note that the data structure described is just one example and does not exclude any data not listed. Also, even if data is listed in the same table, it may be stored in separate memory areas in the storage unit 202.
[0060] Figure 4 shows the data structure of the user information table 2021. The user information table 2021 shown in Figure 4 is a table that uses User ID as the key and has columns for Name, Age, Gender, Date of Birth, and Contact Information.
[0061] User ID is an item that stores an identifier to uniquely identify the user. Name is an item that stores the user's name. Age is an item that stores the user's age. Gender is an item that stores the user's gender. Date of birth is an item that stores the user's date of birth. Contact information is an item that stores the contact information (e.g., telephone number, email address, etc.) of the terminal device 10 that the user possesses.
[0062] Figure 5 shows the data structure of the video log table 2022. As shown in Figure 5, the video log table 2022 has items such as video ID, frame ID, time, and description. The video log table 2022 lists the frame IDs under the video ID. The video log table 2022 associates the time and description with the frame ID as the key.
[0063] The video ID indicates identification information for identifying a video recorded by the recording device 31. For example, a different video ID is assigned to each video recorded by the recording device 31. Even if videos are recorded by the same recording device 31, they may be recognized as different videos at different times (for example, every day, every three hours, etc.) and assigned different video IDs.
[0064] The frame ID is an identifier used to identify a frame that makes up a video. A frame ID is assigned to a specific frame within a video. Details will be described later. The time indicates the time the frame was extracted.
[0065] The description text explains the content of the frame. Further details will be provided later.
[0066] <5. Operation> An example of the processing flow in System 1 is described below.
[0067] Figure 6 is a flowchart illustrating an example of the operation when server 20 causes LLM to generate a video summary.
[0068] In step S1001, the server 20 receives video from the camera 31. Specifically, the receiving control module 2031 receives video of the scene captured by the camera 31 from the camera 31. The video is assigned a video ID based on, for example, the identification information of the camera 31 that captured it. The server 20 stores the received video in the storage unit 202, for example. The camera 31 may send the video to the server 20 all at once as a single file, or it may send the video to the server 20 in packets.
[0069] The receiving control module 2031 may, for example, receive a predetermined frame from the recording device 31, instead of a video, from a plurality of frames that make up a video. In this case, the recording device 31 transmits a frame to the server 20 at a predetermined interval (for example, every 30 seconds). Alternatively, the recording device 31 may transmit a frame at random timings to the server 20.
[0070] In step S1002, the server 20 selects a frame. Specifically, the service processing module 2033 selects a predetermined frame from among a plurality of frames that make up the received video, according to a predetermined rule.
[0071] For example, the service processing module 2033 selects frames in the video at predetermined intervals (for example, every 30 seconds).
[0072] Alternatively, for example, the service processing module 2033 may select frames at random timings in the video.
[0073] Alternatively, for example, the service processing module 2033 may select a frame in which the image analysis module 2035 has detected a predetermined object or predetermined action. In this case, for example, the image analysis module 2035 analyzes the video received in step S1001. A predetermined object could be, for example, a vehicle, an animal, or a weapon. A predetermined action could be, for example, entering or leaving a room, falling, or colliding. By analyzing the video, the image analysis module 2035 detects a predetermined object, a predetermined action, or a combination thereof.
[0074] The service processing module 2033 assigns a frame ID to the selected frame.
[0075] In step S1003, the server 20 sends a frame and prompt A to the LLM system 30. Specifically, the service processing module 2033 sends a selected frame and prompt A to the LLM system 30 to describe the content of the frame. Prompt A is set by default by, for example, the video summarization service provider. The text of prompt A is, for example, "Please output a sentence that describes the content of the input frame." The LLM system 30 inputs the transmitted frame and prompt A into the LLM.
[0076] The service processing module 2033 may also send a video to the LLM system 30 instead of a frame. In this case, the text of prompt A may be, for example, "Please output a statement describing the content of the frame to which the frame ID has been assigned." The text of prompt A may also be a statement describing the content of a frame at a predetermined timing. Specifically, for example, the text of prompt A may be, "Please output a statement describing the content of the frame to which the tag has been assigned." Alternatively, the text of prompt A may be, for example, "Please output a statement describing the content of a frame at a predetermined period."
[0077] In step S1004, the server 20 causes the LLM system 30 to output an explanatory text. Specifically, the LLM outputs an explanatory text in response to the input prompt A. The receiving control module 2031 receives the outputted explanatory text from the LLM system 30. The server 20 stores the received explanatory text, as well as the corresponding video ID, frame ID, and time in the video log table 2022. In other words, the service processing module 2033 causes the LLM to output a frame explanatory text and stores it in the video log table 2022, even without instructions from the user.
[0078] In step S1005, the server 20 presents information related to the explanatory text to the terminal device 10. Specifically, the transmission control module 2032 transmits the video data, video ID, frame ID, time, and explanatory text to the terminal device 10. The presentation control module 2034 transmits the arrangement information for displaying the video data, video ID, frame ID, time, and explanatory text on the terminal device 10 to the terminal device 10.
[0079] In step S1006, the terminal device 10 presents the user with information relating to the explanatory text. Specifically, the presentation control unit 193 displays the video data, video ID, frame ID, time, and explanatory text on the display 141 according to the arrangement information transmitted from the server 20. At this time, the presentation control unit 193 may display the explanatory text of a frame in which a predetermined object or predetermined action is shown in a manner that makes it distinguishable from other explanatory texts. For example, the presentation control unit 193 may mark the explanatory text of a frame in which a predetermined object or predetermined action is shown with a mark that will attract the user's attention. A mark that will attract the user's attention may be, for example, a star. A predetermined object may be, for example, a vehicle, an animal, a weapon, etc. A predetermined action may be, for example, entering or leaving a room, falling, colliding, etc. The presentation control unit 193 does not have to display the information relating to the explanatory text on the display 141.
[0080] In step S1007, the terminal device 10 receives a command from the user to output a summary. Specifically, the operation reception unit 191 receives prompt B for outputting a video summary. For example, the user refers to a screen displaying explanatory text for multiple frames and enters prompt B on the display 141. The user may also enter prompt B on a screen where no explanatory text is displayed. The text of prompt B may be, for example, "If a dangerous situation occurs in the video, please tell me the situation and the time." Alternatively, the text of prompt B may be, for example, "Please tell me what happened from HH:MM1 to MM2." The server 20 receives prompt B from the terminal device 10. The user may or may not specify the time range (start time to end time) of the video to be summarized in prompt B. If the user does not specify the time range of the video to be summarized in prompt B, the service processing module 2033 assumes, for example, that the entire video, i.e., the entire time from the start time to the end time of the video, has been specified as the time range.
[0081] In step S1008, the server 20 sends prompt B and an explanatory text to the LLM system 30. Specifically, the service processing module 2033 sends prompt B, which was sent from the terminal device 10, and an explanatory text for the time range necessary to create the summary to the LLM system 30. If the time range of the target video is specified in prompt B, the service processing module 2033 sends the explanatory text for the time range specified in prompt B to the LLM system 30.
[0082] In step S1009, the server 20 causes the LLM system 30 to output a video summary. Specifically, the LLM outputs a video summary based on the prompt B transmitted in step S1008 and the explanatory text. The video summary output from the LLM is, for example, based on the content of multiple explanatory texts. The receiving control module 2031 receives the video summary output from the LLM system 30.
[0083] In step S1010, the server 20 presents information related to the video summary to the terminal device 10. Specifically, the presentation control module 2034 presents information related to the video summary to the terminal device 10. The transmission control module 2032 transmits to the terminal device 10 arrangement information for displaying the video summary on the terminal device 10.
[0084] In step S1011, the terminal device 10 displays information related to the video summary to the user. Specifically, the display control unit 193 displays the video summary on the display 141 based on the arrangement information transmitted from the server 20. The display control unit 193 may display words, phrases, and sentences that indicate noteworthy events in the video summary in a way that makes them distinguishable from other words, phrases, and sentences. Specifically, for example, the display control unit 193 may add decorations that attract the user's attention. Examples of decorations that attract the user's attention include bold text and colored text.
[0085] Note that the LLM system 30 that outputs the explanatory text in step S1004 and the LLM system 30 that outputs the video summary in step S1009 may be separate systems. For example, LLM system 30A may output the explanatory text in response to prompt A, and LLM system 30B may output the video summary in response to prompt B.
[0086] <6. Screen Examples> An example of the screen of the display 141 of the terminal device 10 in this disclosure will be described.
[0087] Figure 7 shows an example of a screen where the explanatory text and summary are displayed, and prompt B is entered.
[0088] Region 501 is the area where the video is displayed in step 1006. In Figure 7, a frame showing a person lying on the floor is displayed in region 501. Information about the video is displayed in the corner of region 501. This information includes, for example, the video ID and the time of the displayed frame.
[0089] Bar 502 is a bar that represents the playback position of the video on the timeline. The user can adjust the playback position by adjusting the dots on bar 502.
[0090] Icon group 503 consists of icons for controlling the video. Icon group 503 is not limited to the three icons shown in Figure 7.
[0091] Table 504 is a table used in step 1006 to display the frame ID, time, and description. When the user clicks on the frame ID, time, or description in the table, the corresponding frame is displayed in table area 501.
[0092] Box 505 is a box in step 1007 for the user to input prompt B. Prompt B entered by the user may remain displayed in box 505 until, for example, the user terminates the video summary service (for example, until the screen in Figure 7 is closed by the user).
[0093] Box 506 is a box for displaying a summary in step 1011. The summary output as a response to prompt B may remain displayed in box 506 until, for example, the user terminates the video summary service (for example, until the screen in Figure 7 is closed by the user). Although Figure 7 illustrates a case where prompt B and the response as a summary are displayed in different boxes, the display format is not limited to this. Prompt B and the response as a summary may be displayed in a chat format.
[0094] <7. Summary> As described above, in the above embodiment, the server 20 acquires a video consisting of multiple frames. The server 20 inputs the video or predetermined frames that make up the video, and prompts for explaining the content of the frames that make up the video, to the large-scale language model, and causes the large-scale language model to output explanatory text for the frames. The server 20 inputs multiple explanatory texts and prompts for outputting a summary of the video to the large-scale language model, and causes the large-scale language model to output a summary of the video. The server 20 presents the summary output by the large-scale language model to the user. As a result, the user can easily grasp the outline of the events occurring in the video through the video summary without having to watch the entire video.
[0095] Furthermore, in the above embodiment, in the step of outputting the explanatory text, the server 20 inputs frames from among multiple frames at predetermined intervals into the large-scale language model. This makes it possible to select frames based on time conditions. In addition, frames are automatically input into the large-scale language model.
[0096] Furthermore, in the above embodiment, in the step of outputting the explanatory text, the server 20 inputs a frame at a random timing from among multiple frames into the large-scale language model. This makes it possible to select a random frame. Also, frames are automatically input into the large-scale language model.
[0097] Furthermore, in the above embodiment, in the step of outputting the explanatory text, the server 20 inputs frames in which a predetermined object or predetermined action is detected into the large-scale language model. This makes it possible to select frames based on the object or action that the user is interested in. In addition, frames are automatically input into the large-scale language model.
[0098] Furthermore, in the above embodiment, the server 20 presents the user with a list of explanatory texts and receives a prompt from the user to output a video summary. This improves the user's ability to easily view the explanatory texts and makes it easier for them to grasp the information needed to make a decision on how to generate a summary.
[0099] Furthermore, in the above embodiment, during the receiving step, the server 20 presents to the user a descriptive text for a frame in which a predetermined object or predetermined action is displayed, in a way that allows it to be distinguished from other descriptive texts. This allows the user to instantly identify the frame in which the object or action they are interested in is displayed.
[0100] Furthermore, in the above embodiment, in the step of outputting a summary, the server 20 inputs descriptive texts for multiple frames included in a predetermined time width and a prompt to output a summary of the video of the predetermined time width into the large-scale language model, causing the large-scale language model to output a summary of the video of the predetermined time width. As a result, the user can obtain a video summary of only the time width they desire.
[0101] Furthermore, in the above embodiment, the video is a video captured by a surveillance camera. Alternatively, the video may be a video captured during surgery. This allows for the summarization of videos from a variety of locations and situations.
[0102] <8. Modifications> Modifications of the above embodiment will now be described. <8.1. Modification 1> In the above embodiment, in step S1007, the user specifies the time range to be summarized in the video using prompt B, that is, in a linguistic way. However, the user may specify the time range to be summarized in a non-linguistic way. For example, the terminal device 10 may accept the specification of the time range from the user by dragging the bar 502 in Figure 7 from the start time point to the end time point. Alternatively, the terminal device 10 may accept the specification of the time range from the user by dragging the table 504 in Figure 7 from the start time row to the end time row.
[0103] In these cases, the user does not specify a time range in Prompt B. The text of Prompt B might be, for example, "If a dangerous situation occurs in the video, please tell me the situation and the time."
[0104] <8.2. Modification 2> In the above embodiment, the server 20 caused the LLM system 30 to output the descriptive text for the entire area within the selected frame. However, the server 20 may also cause the LLM system 30 to output the descriptive text for only a portion of the area within the selected frame. For example, a portion of the area 501 on the display 141 where the video in Figure 7 is displayed is selected by the user by dragging, and a prompt is input requesting a description of the image represented in that area. The prompt at this time may be, for example, "How many XX are there in this?" or "Please tell me the situation in this." The server 20 receives information identifying the specified area and the prompt from the terminal device 10. The server 20 extracts an image from the frame based on the specified area. The server 20 inputs the extracted image and the prompt to the LLM system 30, causing the LLM system 30 to output the descriptive text for the specified portion of the selected frame.
[0105] Furthermore, if there is a scene in a designated area that should be processed, the server 20 may process that scene. Scenes to be processed are, for example, scenes related to privacy. The processing is a process that reduces the identifiableness of the subject, such as masking or mosaic processing. For example, if the server 20 detects a scene that should be considered for privacy (for example, a person's face) within the designated area, it will apply a mosaic to that subject.
[0106] <8.3. Modification 3> The server 20 may refer to the frame descriptions and create a video by stitching together important frames, i.e., a highlight video. For example, an important frame is a frame with a long description (the description is five sentences or more, the description is 50 characters or more, etc.). Alternatively, an important frame may be a frame with a description that refers to a specific object or a specific action. A specific object could be, for example, a vehicle, an animal, a weapon, etc. A specific action could be, for example, entering or leaving a room, falling, colliding, etc.
[0107] Server 20 receives, for example, a user's instruction to create a highlight video. Upon receiving the instruction, Server 20 analyzes, for example, the stored explanatory text and extracts important frames. The analysis may be performed using, for example, natural language processing, a pre-trained model trained to extract important frames, or LLM. Based on the important frames, Server 20 extracts frames that can constitute a highlight video. Server 20 may use the important frames as frames that can constitute a highlight video, or it may use multiple frames based on the important frames as frames that can constitute a highlight video. Server 20 connects the extracted frames and creates a highlight video. This makes it easy for the user to obtain the highlight video and reduces the burden of reviewing the video.
[0108] <8.4. Modification 4> In the above embodiment, in step S1007, prompt B is a prompt to output a summary of the video, and the text of prompt B is, for example, "If a dangerous situation occurs in the video, please tell me the situation and time." However, prompt B may also be a prompt to analyze events in the video. That is, the text of prompt B may be, for example, "Please tell me the age distribution of customers in the video." In this case, in step S1009, LLM may output the results in an intuitive format. That is, for example, the text of prompt B may be, "Please show me the age distribution of customers in the video in a graph." For example, in step S1009, LLM outputs a graph of the age distribution of customers based on the explanatory text referring to the age of customers in each frame output in step S1004.
[0109] <Basic Hardware Configuration of Computer> Figure 8 is a block diagram showing the basic hardware configuration of computer 90. Computer 90 includes at least a processor 901, main memory 902, auxiliary storage 903, and a communication IF 991 (interface). These are electrically connected to each other by a communication bus 921.
[0110] The processor 901 is hardware for executing the instruction set described in the program. The processor 901 consists of an arithmetic unit, registers, peripheral circuits, etc.
[0111] The main memory 902 is for temporarily storing programs and data processed by programs, etc. For example, it is a volatile memory such as DRAM (Dynamic Random Access Memory).
[0112] The auxiliary storage device 903 is a storage device for storing data and programs. Examples include flash memory, HDD (Hard Disc Drive), magneto-optical disk, CD-ROM, DVD-ROM, semiconductor memory, etc.
[0113] A communication interface (IF991) is an interface for inputting and outputting signals for communication with other computers via a network using wired or wireless communication standards. The network consists of various mobile communication systems, such as the Internet, LANs, and wireless base stations. For example, networks include 3G, 4G, and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks (e.g., Wi-Fi®) that can connect to the Internet via designated access points. When connecting wirelessly, communication protocols include, for example, Z-Wave®, ZigBee®, and Bluetooth®. When connecting via wired connections, the network also includes connections made directly via USB (Universal Serial Bus) cables, etc.
[0114] Furthermore, by distributing all or part of each hardware configuration across multiple computers 90 and connecting them to each other via a network, a computer 90 can be virtually realized. Thus, the concept of computer 90 includes not only a computer 90 housed in a single enclosure or case, but also a virtualized computer system.
[0115] <Basic Functional Configuration of Computer 90> The functional configuration of the computer realized by the basic hardware configuration of computer 90 (Figure 8) is described below. The computer comprises at least one functional unit: a control unit, a memory unit, and a communication unit.
[0116] Furthermore, the functional units of computer 90 can also be realized by distributing all or part of each functional unit across multiple computers 90 interconnected via a network. The concept of computer 90 includes not only a single computer 90 but also a virtualized computer system.
[0117] The control unit is realized when the processor 901 reads various programs stored in the auxiliary storage device 903, loads them into the main memory device 902, and executes processing according to those programs. The control unit can realize various functional units that perform information processing depending on the type of program. In this way, the computer is realized as an information processing device that performs information processing.
[0118] The memory unit is implemented by a main memory 902 and an auxiliary memory 903. The memory unit stores data, various programs, and various databases. The processor 901 can also reserve memory areas corresponding to the memory unit in the main memory 902 or the auxiliary memory 903 according to the program. The control unit can also cause the processor 901 to perform addition, update, and deletion operations on data stored in the memory unit according to the various programs.
[0119] The term "database" refers to a relational database, which is used to manage and associate data sets called tables and masters, which are structured in a tabular format defined by rows and columns. In a database, tables are called tables, masters are called masters, the columns of tables are called columns, and the rows of tables are called records. In a relational database, relationships can be established and linked between tables and masters. Typically, each table and each master has a primary key column to uniquely identify a record, but setting a primary key for a column is not mandatory. The control unit can cause the processor 901 to add, delete, and update records in specific tables and masters stored in the storage unit according to various programs. Furthermore, by storing data, various programs, and various databases in the storage unit, the information processing device and information processing system described in this disclosure can be considered manufactured.
[0120] Furthermore, the databases and masters in this disclosure may include any data structures (lists, dictionaries, associative arrays, objects, etc.) in which information is structurally defined. Data structures also include data that can be considered as data structures by combining data with functions, classes, methods, etc., written in any programming language.
[0121] The communication unit is implemented by the communication IF 991. The communication unit implements the function of communicating with other computers 90 via the network. The communication unit can receive information transmitted from other computers 90 and input it to the control unit. The control unit can cause the processor 901 to perform information processing on the received information according to various programs. The communication unit can also transmit information output from the control unit to other computers 90.
[0122] Furthermore, each of the above-mentioned configurations, functions, processing units, processing means, etc., may be implemented in hardware, in whole or in part, for example, by designing them as integrated circuits. The present invention can also be implemented by software program code that realizes the functions of the embodiment. In this case, a storage medium on which the program code is recorded is provided to a computer, and the processor of that computer reads the program code stored in the storage medium. In this case, the program code read from the storage medium itself realizes the functions of the embodiment described above, and the program code itself and the storage medium on which it is stored constitute the present invention. Examples of storage media used to supply such program code include flexible disks, CD-ROMs, DVD-ROMs, hard disks, SSDs, optical disks, magneto-optical disks, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, and the like.
[0123] Furthermore, the program code that implements the functions described in this embodiment can be implemented in a wide range of programming or scripting languages, such as assembler, C / C++, Perl, Shell, PHP, and Java®.
[0124] Furthermore, the program code for the software that implements the functions of the embodiment may be distributed via a network and stored in a storage means such as a computer's hard disk or memory, or in a storage medium such as a CD-RW or CD-R, and the computer's processor may read and execute the program code stored in the storage means or storage medium.
[0125] The functions realized by the components described herein may be implemented in a circuit or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), CPUs (a Central Processing Unit), conventional circuits, and / or combinations thereof, programmed to realize the described functions. A processor, including transistors and other circuits, is considered a circuit or processing circuitry. A processor may be a programmed processor that executes a program stored in memory. In this specification, circuitry, unit, and means are hardware programmed to realize or perform the described functions. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to realize or perform the described functions. If such hardware is a processor that is considered a type of circuitry, then such circuitry, means, or unit is a combination of hardware and software used to constitute such hardware and / or processor.
[0126] While several embodiments of this disclosure have been described above, these embodiments can be implemented in a variety of other forms, and various omissions, substitutions, and modifications are permitted without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims and their equivalents.
[0127] (Note) The matters described in each of the above embodiments are noted below.
[0128] (Note 1) A program for operating a computer comprising a processor and memory, the program causing the processor to perform the following steps: acquiring a video consisting of multiple frames; inputting a video or predetermined frames comprising the video and prompts for explaining the content of the frames comprising the video into a large-scale language model and outputting explanatory text for the frames into the large-scale language model; inputting multiple explanatory texts and prompts for outputting a summary of the video into the large-scale language model and outputting a summary of the video into the large-scale language model; and presenting the summary output by the large-scale language model to the user. (Note 2) The program described in (Note 1), wherein, in the step of outputting explanatory text, frames at predetermined intervals from among the multiple frames are input into the large-scale language model. (Note 3) The program described in (Note 1) or (Note 2), wherein, in the step of outputting explanatory text, frames at random timings from among the multiple frames are input into the large-scale language model. (Note 4) A program according to any one of (Note 1) to (Note 3), wherein in the step of outputting a description, the program inputs frames in which a predetermined object or predetermined action is detected from among multiple frames into a large-scale language model. (Note 5) A program according to any one of (Note 1) to (Note 4) that causes the processor to execute a step of presenting a list of description sentences to the user and receiving a prompt from the user to output a summary of the video. (Note 6) A program according to (Note 5) that, in the receiving step, presents to the user the description sentence of the frame in which a predetermined object or predetermined action appears in a manner that allows it to be distinguished from other description sentences. (Note 7) A program according to any one of (Note 1) to (Note 6) that, in the step of outputting a summary, inputs the description sentences of multiple frames contained within a predetermined time width and a prompt to output a summary of the video of the predetermined time width into a large-scale language model, and causes the large-scale language model to output a summary of the video of the predetermined time width. (Note 8) In the acquisition step, the video is a video captured by a surveillance camera, as described in any of (Note 1) to (Note 7).(Note 9) A program according to any one of (Note 1) to (Note 8), wherein in the step of acquiring the video, the video is a video taken during surgery. (Note 10) A method to be executed on a computer comprising a processor and memory, wherein the processor executes all the steps performed in the invention according to any one of (Note 1) to (Note 9). (Note 11) An information processing apparatus comprising a control unit and a storage unit, wherein the control unit executes all the steps performed in the invention according to any one of (Note 1) to (Note 9). (Note 12) A system comprising means for executing all the steps performed in the invention according to any one of (Note 1) to (Note 9).
[0129] 1...System 10...Terminal device 12...Communication IF 13...Input device 14...Output device 15...Memory 16...Storage 19...Processor 20...Server 22...Communication IF 23...Input / Output IF 25...Memory 26...Storage 29...Processor 30...LLM system 31...Imaging device 80...Network
Claims
1. A program for operating a computer comprising a processor and memory, wherein the program causes the processor to perform the following steps: acquiring a video consisting of multiple frames; inputting the video or predetermined frames comprising the video and prompts for explaining the content of the frames comprising the video to a large-scale language model and causing the large-scale language model to output explanatory text for the frames; inputting a plurality of explanatory texts and prompts for outputting a summary of the video to the large-scale language model and causing the large-scale language model to output a summary of the video; and presenting the summary output by the large-scale language model to a user.
2. The program according to claim 1, wherein, in the step of outputting the explanatory text, frames from among multiple frames, at predetermined intervals, are input to the large-scale language model.
3. The program according to claim 1, wherein, in the step of outputting the explanatory text, frames at random timings from among multiple frames are input to the large-scale language model.
4. The program according to claim 1, wherein, in the step of outputting the explanatory text, frames in which a predetermined object or predetermined action is detected from among multiple frames are input to the large-scale language model.
5. The program according to claim 1, wherein the processor is instructed to perform the steps of presenting the user with a list of the explanatory texts and receiving a prompt from the user to output a summary of the video.
6. The program according to claim 5, wherein, in the receiving step, the program presents to the user, in a manner distinguishable from other descriptive texts, a descriptive text of a frame in which a predetermined object or predetermined action is shown.
7. The program according to claim 1, wherein, in the step of outputting the summary, the program inputs a description of multiple frames contained within a predetermined time width and a prompt for outputting a summary of the video of the predetermined time width to the large-scale language model, and causes the large-scale language model to output a summary of the video of the predetermined time width.
8. The program according to claim 1, wherein in the step of acquiring the video, the video is a video captured by a surveillance camera.
9. The program according to claim 1, wherein in the step of acquiring, the video is a video taken during surgery.
10. A method to be performed on a computer comprising a processor and memory, wherein the processor performs all steps performed in any of the inventions according to claims 1 to 9.
11. An information processing apparatus comprising a control unit and a storage unit, wherein the control unit performs all steps performed in the invention according to any one of claims 1 to 9.
12. A system comprising means for performing all steps performed in the invention according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method, device and equipment for generating video abstract, and computer readable storage medium
CN113542910A
Interest point image generation method and device, electronic equipment and storage medium
CN117237606A
Video title generation method and training method of video title generation model
CN117609550A
Cboth document generation model training method, copywriting generation method and copywriting generation device
CN117746279A
Method, server and computer program
JP7385204B1