Information Processing Apparatus and Information Processing Program

The information processing apparatus generates video descriptions that reflect scene context by using an image caption model with past frame descriptions, addressing limitations of conventional technologies and enabling efficient processing of long videos.

JP7709709B1Active Publication Date: 2025-07-17SOFTBANK CORPORATION +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024040463
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-03-14
Publication Date
2025-07-17
Estimated Expiration
2044-03-14

AI Technical Summary

Technical Problem

Conventional video captioning technologies struggle to generate descriptions that reflect the context formed by the stacking of scenes in a video, often requiring large model sizes and having limitations on the number of frames that can be processed, making them unsuitable for long videos.

Method used

An information processing apparatus that uses an image caption model to generate video descriptions by inputting each frame of a video along with a past description text, allowing it to capture the context between frames and generate descriptions that reflect the stacking of scenes, while utilizing a smaller model size and avoiding limitations on the number of frames.

Benefits of technology

The apparatus effectively generates video descriptions that reflect the context of each scene, overcoming limitations of previous technologies by using a smaller model size and enabling the processing of long videos without upper frame limits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007709709000001_ABST
    Figure 0007709709000001_ABST
Patent Text Reader

Abstract

It is possible to generate a video description text that reflects the context formed by stacking each scene in a video. 【Solution means】 The information processing apparatus according to the present application is an information processing apparatus including a machine learning model that generates a video description text, which is a sentence that describes the content included in the video to be processed, from the video to be processed. The information processing apparatus includes an acquisition unit that acquires the video to be processed, and inputs, to the machine learning model, a current frame that is a frame to be a target for generating a description text among the frames constituting the video to be processed, and a past description text that describes the content of a frame before the current frame. By doing so, a current frame description text that describes the content of the current frame is generated, and a generation unit that generates a video description text including the past description text and the current frame description text is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus and an information processing program.

Background Art

[0002] Conventionally, a technique for generating a caption of a video (also referred to as a video caption. Hereinafter, it will be described as a "video description"). For example, a technique is known in which a video captured by a surveillance camera is input to a multi-layer neural network that outputs elements included in an image as words, and a video description is generated.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the above conventional technology, since a video captured by a surveillance camera is only input to a multi-layer neural network that outputs elements included in an image as words and a video description is generated, it is not always possible to generate a video description that reflects the context formed by the stacking of each scene in the video.

[0005] An object of the present application is to provide an information processing apparatus and an information processing program capable of generating a video description that reflects the context formed by the stacking of each scene in a video.

Means for Solving the Problems

[0006] The information processing apparatus according to the present application is an information processing apparatus including a machine learning model that generates a video description text, which is a text describing the content included in the video to be processed, from the video to be processed. The information processing apparatus includes an acquisition unit that acquires the video to be processed, and inputs, to the machine learning model, a current frame, which is a frame to be a target for generating a description text, among the frames constituting the video to be processed, and a past description text that describes the content of a frame previous to the current frame, to generate a current frame description text that describes the content of the current frame, and a generation unit that generates the video description text including the past description text and the current frame description text.

Effect of the Invention

[0007] According to one aspect of the embodiment, it is possible to generate a video description text that reflects the context formed by stacking each scene in the video.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Mode for Carrying Out the Invention

[0009] Hereinafter, embodiments for implementing the information processing apparatus and the information processing program according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the information processing apparatus and the information processing program according to the present application are not limited by this embodiment. Also, in each of the following embodiments, the same parts are denoted by the same reference numerals, and redundant descriptions are omitted.

[0010] (Embodiment) [1. Introduction] Conventionally, a technique for generating a video description text (also referred to as a video caption. Hereinafter, it will be described as "video description text") that is a text explaining the content of a video has been known. For example, in recent years, techniques related to a machine learning model for generating a video description text from a video (hereinafter, may be described as a "video caption model") have been known, but these video caption models are known to have a very large model size (Reference 1; Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, Yu Qiao, "VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking", [online], 2023, CVPR, [searched on February 1, 2025], Internet <URL:https: / / openaccess.thecvf.com / content / CVPR2023 / papers / Wang_VideoMAE_V2_Scaling_Video_Masked_Autoencoders_With_Dual_Masking_CVPR_2023_paper.pdf>). Thus, the fact that the model size of the video caption model becomes large means that the amount of computation by the information processing apparatus becomes large. That is, it means that there is a limit to the scaling up of the video caption model. Also, among these video caption models, there is originally an upper limit to the number of frames that can be received (for example, it can only receive a video with a sequence length learned during training. Also, the number of frames is 6 to 16, etc.).)Therefore, there are also some that are difficult to apply to long videos (Reference 2; Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, Haibin Ling, "Expanding Language-Image Pretrained Models for General Video Recognition", [online], 4 Aug 2022, ECCV, [searched on February 1, 2024], Internet <URL:https: / / arxiv.org / abs / 2208.02816>).

[0011] Conventionally, there has also been known a technique for generating an image description text (also referred to as an image caption. Hereinafter, it will be referred to as "image description text") which is a text describing the content of an image from an image (still image). For example, a technique related to a machine learning model (hereinafter, may be referred to as an "image caption model") for generating an image description text which is a text describing the content of an image from an image is known. Therefore, it is conceivable to input a predetermined frame constituting a video into the image caption model to generate an image description text corresponding to the predetermined frame constituting the video (hereinafter, may be referred to as a "method using an image caption model"). However, the method using an image caption model only generates a description text corresponding to one scene corresponding to one frame in the video for each frame (that is, for each scene), so there is a limit in applying it to a video. For example, it is difficult for the method using an image caption model to capture the relationship between the frames constituting the video. For this reason, the method using an image caption model becomes more difficult to generate an appropriate video description text as the video becomes longer and the story of the video becomes more prominent, because it becomes necessary to capture the relationship between the frames constituting the video. Therefore, it is difficult for the method using an image caption model to generate a video description text that reflects the context formed by the stacking of each scene in the video.

[0012] In contrast, the information processing apparatus according to the embodiment generates a caption (video description text) for the entire video by repeatedly applying an image caption model to each frame constituting the video. Specifically, the information processing apparatus repeatedly inputs each of a plurality of frames constituting the video into the image caption model to generate a plurality of frame description texts for explaining the content of each of the plurality of frames, and generates a text including the plurality of frame description texts as the video description text. More specifically, when generating each of the plurality of frame description texts, the information processing apparatus inputs a past description text for explaining the content of a frame before the current frame (hereinafter sometimes referred to as a "past frame") into the image caption model to generate a current frame description text for explaining the content of the current frame. Here, the frame before the current frame (past frame) refers to one or more frames displayed before the current frame in the video to be processed. In other words, the frame before the current frame refers to one or more frames displayed at a time earlier than the time when the current frame is displayed in the video to be processed. In other words, the frame before the current frame refers to one or more frames played at a time earlier than the playback time of the current frame in the playback time of the video. In other words, the frame before the current frame refers to one or more frames played at a time earlier than the playback time of the current frame in the playback time of the video. For example, the past frame according to the present embodiment includes one frame before the current frame. Further, the past frame according to the present embodiment may include a plurality of frames (constituting a video) before the current frame among the frames constituting the video to be processed. Further, the past description text according to the present embodiment includes a description text for explaining the content of one frame before the current frame. Further, the past description text according to the present embodiment includes a description text for explaining the content of a plurality of frames (constituting a video) before the current frame among the frames constituting the video to be processed.

[0013] In this way, the information processing device inputs, as past context information formed by stacking scenes prior to the scene of the present frame, a past description text that describes the content of a frame prior to the present frame into an image caption model. Thereby, the information processing device can generate a video description text that reflects the context formed by stacking each scene in the video. Also, in this way, the information processing device can suppress the amount of calculation when generating a video description text by using an image caption model with a smaller model size compared to a conventional video caption model. Further, since there is no upper limit to the number of frames that the image caption model can receive by repeatedly using the image caption model, the information processing device can generate a video description text corresponding to a long video.

[0014] [2. Configuration of Information Processing Device] With reference to FIG. 1, a configuration example of an information processing device 100 according to an embodiment will be described. FIG. 1 is a diagram showing a configuration example of the information processing device 100 according to the embodiment. The information processing device 100 includes a machine learning model that generates a video description text, which is a text that describes the content included in the video to be processed, from the video to be processed. Also, the information processing device 100 includes a communication unit 110, a storage unit 120, and a control unit 130.

[0015] (Communication Unit 110) The communication unit 110 is realized by a NIC (Network Interface Card), an antenna, or the like. The communication unit 110 is connected to various networks by wire or wirelessly, and performs transmission and reception of information, for example, with other information processing devices other than the information processing device 100.

[0016] (Storage Unit 120) The storage unit 120 is implemented by, for example, a semiconductor memory device such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. Specifically, the storage unit 120 stores various data. For example, the storage unit 120 stores various programs. For example, the storage unit 120 stores the information processing program according to the embodiment. Further, the storage unit 120 may store information regarding a machine learning model that generates an image description text, which is a text describing the content of an image from the image. For example, the storage unit 120 may store information regarding a Visual Language Model (VLM) as a machine learning model that generates an image description text from an image. For example, the storage unit 120 may store information regarding CoCa (Contrastive Captioners are Image-Text Foundation Models), BLIP (Bootstrapping Language-Image Pre-training), BLIP2, GIT (Generative Image to Text Transformer), etc. as a visual language model. Further, the storage unit 120 may store information regarding the video to be processed acquired by the acquisition unit 131. Further, the storage unit 120 may store information regarding an encoder that generates feature information indicating the features of an image from the image. Further, the storage unit 120 may store information regarding the image description text and the video description text generated by the generation unit 133.

[0017] (Control unit 130) The control unit 130 is a controller, which is realized, for example, by executing various programs stored in the storage device inside the information processing apparatus 100 with the RAM as a work area by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or the like. Further, the control unit 130 is a controller, which is realized, for example, by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0018] The control unit 130 has an acquisition unit 131, a selection unit 132, and a generation unit 133 as functional units, and may realize or execute the operations of information processing described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in FIG. 1, and may be any other configuration as long as it can perform the information processing described later. Further, each functional unit represents the function of the control unit 130, and does not necessarily have to be physically distinguished.

[0019] (Acquisition Unit 131) The acquisition unit 131 acquires a video to be processed. Specifically, the acquisition unit 131 may acquire the video to be processed from an external information processing apparatus used by a user via the communication unit 110. Further, when the acquisition unit 131 acquires the video to be processed, it may store information regarding the video to be processed in the storage unit 120. Further, when the acquisition unit 131 acquires the video to be processed, it may acquire a machine learning model that generates an image description text, which is a text describing the content of an image from the image, by referring to the storage unit 120. For example, the acquisition unit 131 may acquire a vision language model as the machine learning model that generates an image description text from an image. Further, when the acquisition unit 131 acquires the video to be processed, it may output the video to be processed to the selection unit 132. Further, when the acquisition unit 131 acquires the machine learning model, it may output the machine learning model to the generation unit 133.

[0020] (Selection Unit 132) The selection unit 132 may acquire the video to be processed output from the acquisition unit 131. When the selection unit 132 acquires the video to be processed, the selection unit 132 may select this frame, which is the frame to be the target for generating the image description text, from among the plurality of frames constituting the video to be processed. The selection unit 132 selects this frame from among the plurality of frames constituting the video to be processed.

[0021] For example, the selection unit 132 may select the first frame, which is the frame to be the target for generating the image description text first, from among the frames constituting the video to be processed. For example, the selection unit 132 may select the first frame as the first frame from among the frames constituting the video to be processed. Here, the first frame refers to the frame corresponding to the start time of the video. Note that the selection unit 132 may select, as the first frame, the frame displayed after a predetermined time (for example, 10 seconds, etc.) has elapsed from the start time of the video from among the frames constituting the video to be processed. Further, when the selection unit 132 selects the first frame, the selection unit 132 may output the first frame to the generation unit 133. Further, the generation unit 133 acquires the first frame from the selection unit 132. Further, when the generation unit 133 acquires the first frame, the generation unit 133 generates the first frame description text for describing the content of the first frame.

[0022] Further, the selection unit 132 may select a second reference frame, which is a frame to be the target for generating an image description text next after the first reference frame, from among the frames constituting the video to be processed. For example, the selection unit 132 may calculate the frame similarity between the first reference frame and each of a plurality of frames after the first reference frame. For example, the selection unit 132 may calculate, as the frame similarity, the similarity between the feature information indicating the features of the first reference frame and the feature information indicating the features of each of the frames after the first reference frame. Subsequently, the selection unit 132 may determine whether there is a frame in which the frame similarity is lower than a predetermined threshold among the frames after the first reference frame. For example, the selection unit 132 may determine whether there is a frame in which the similarity between the feature information indicating the features of the first reference frame and the feature information indicating the features of each of the frames after the first reference frame is lower than a predetermined threshold among the frames after the first reference frame. Here, the frames after the first reference frame refer to the frames that are displayed after the first reference frame in the video to be processed. In other words, the frames after the first reference frame refer to the frames that are displayed at a time later than the time when the first reference frame is displayed in the video to be processed. In other words, the frames after the first reference frame refer to the frames that are reproduced at a time later than the reproduction time of the first reference frame in the reproduction time of the video to be processed. That is, the frames after the first reference frame refer to the frames that are reproduced at a time later than the reproduction time of the first reference frame in the reproduction time of the video to be processed. In this way, the selection unit 132 calculates the frame similarity between the past frame (for example, the first reference frame), which is the frame for which a past description text (for example, the first reference frame description text) for explaining the content of the frame before the current frame has been generated, and a plurality of frames after the past frame, and selects, as the current frame (for example, the second reference frame), a frame in which the frame similarity is lower than a predetermined threshold. Further, the selection unit 132 calculates the similarity between the feature information indicating the features of the past frame (for example, the first reference frame) and the feature information indicating the features of the frames after the past frame.

[0023] For example, the selection unit 132 may obtain an encoder that generates feature information indicating the features of an image with reference to the storage unit 120. The encoder is a machine learning model that outputs, as output information, a feature amount indicating the features of an image when the image is input as input information. For example, the encoder may be a convolutional neural network (CNN). Further, the encoder may be a ResNet (Residual Network), AlexNet, VGGNet, GoogLeNet, SENet (Squeeze-and-Excitation Networks), EfficientNet, or ZFNet developed for object recognition. Further, the encoder may be Faster R-CNN, YOLO (You Look Only Once), or SSD (Single Shot MultiBox Detector) developed for object detection. Further, when the selection unit 132 obtains the encoder, the selection unit 132 may input the first frame to the encoder to generate feature information indicating the features of the first frame. Further, the selection unit 132 may input each of the frames after the first frame to the encoder to generate feature information indicating the features of each of the frames after the first frame. Further, when the selection unit 132 generates the feature information indicating the features of the first frame and the feature information indicating the features of each of the frames after the first frame, the selection unit 132 may calculate the degree of similarity between the feature information indicating the features of the first frame and the feature information indicating the features of each of the frames after the first frame. Further, the selection unit 132 may determine whether there is a frame in which the calculated degree of similarity is less than a predetermined threshold among the frames after the first frame.

[0024] Further, when the selection unit 132 determines that there is a frame in the frames after the first frame whose calculated similarity is lower than a predetermined threshold, the selection unit 132 may determine that this frame exists among the frames after the first frame. Further, when the selection unit 132 determines that this frame exists, the selection unit 132 may select a second frame from among the frames whose similarity is lower than the predetermined threshold. For example, when there is only one frame whose similarity is lower than the predetermined threshold, the selection unit 132 may select that frame whose similarity is lower than the predetermined threshold as the second frame. Further, when there are two or more frames whose similarity is lower than the predetermined threshold, the selection unit 132 may select, as the second frame, a frame that is displayed at a time closer to the time when the first frame is displayed from among the two or more frames whose similarity is lower than the predetermined threshold. For example, the selection unit 132 may calculate a value obtained by subtracting the time when the first frame is displayed from the time when each of the two or more frames whose similarity is lower than the predetermined threshold is displayed. Subsequently, the selection unit 132 may select, as the second frame, a frame having the smallest calculated value from among the two or more frames whose similarity is lower than the predetermined threshold. Further, when the selection unit 132 selects the second frame, the selection unit 132 may output the second frame to the generation unit 133. Further, the generation unit 133 acquires the second frame from the selection unit 132. Further, when the generation unit 133 acquires the second frame, the generation unit 133 generates a second frame description text for explaining the content of the second frame.

[0025] Similarly, the selection unit 132 may select the (N + 1)-th frame, which is the frame to generate an image description text next after the N-th (N is a natural number) frame, from among the frames constituting the video to be processed. For example, the selection unit 132 may calculate the frame similarity between the N-th frame and each of the frames after the N-th frame. For example, the selection unit 132 may calculate, as the frame similarity, the similarity between the feature information indicating the features of the N-th frame and the feature information indicating the features of each of the frames after the N-th frame. Subsequently, the selection unit 132 may determine whether there is a frame whose frame similarity is lower than a predetermined threshold among the frames after the N-th frame. For example, the selection unit 132 may determine whether there is a frame whose similarity between the feature information indicating the features of the N-th frame and the feature information indicating the features of each of the frames after the N-th frame is lower than a predetermined threshold among the frames after the N-th frame. Here, the frames after the N-th frame refer to the frames displayed after the N-th frame in the video to be processed. In other words, the frames after the N-th frame refer to the frames displayed at a time later than the time when the N-th frame is displayed in the video to be processed. In other words, the frames after the N-th frame refer to the frames reproduced at a time later than the reproduction time of the N-th frame in the reproduction time of the video to be processed. That is, the frames after the N-th frame refer to the frames reproduced at a time later than the reproduction time of the N-th frame in the reproduction time of the video to be processed. Thus, the selection unit 132 calculates the frame similarity between the past frame (for example, the N-th frame), which is the frame for which a past description text (for example, the N-th frame description text) explaining the content of the frame before the current frame is generated, and a plurality of frames after the past frame, and selects, as the current frame (for example, the (N + 1)-th frame), the frame whose frame similarity is lower than a predetermined threshold. Also, the selection unit 132 calculates the similarity between the feature information indicating the features of the past frame (for example, the N-th frame) and the feature information indicating the features of the frames after the past frame.Here, in addition to the frame description text that describes the content of a specific frame before this frame in the past description text, the past description text may include video description text that describes the content of the video composed of the frames up to this frame among the frames constituting the video to be processed.

[0026] For example, the selection unit 132 may acquire an encoder. Further, when the selection unit 132 acquires the encoder, the selection unit 132 may input the N-th main frame to the encoder to generate feature information indicating the features of the N-th main frame. Further, the selection unit 132 may input each of the frames after the N-th main frame to the encoder to generate feature information indicating the features of each of the frames after the N-th main frame. Further, when the selection unit 132 generates the feature information indicating the features of the N-th main frame and the feature information indicating the features of each of the frames after the N-th main frame, the selection unit 132 may calculate the similarity between the feature information indicating the features of the N-th main frame and the feature information indicating the features of each of the frames after the N-th main frame. Further, the selection unit 132 may determine whether there is a frame among the frames after the N-th main frame whose calculated similarity is less than a predetermined threshold.

[0027] Also, when the selection unit 132 determines that there is a frame in the frames after the N-th frame whose calculated similarity is less than a predetermined threshold, the selection unit 132 may determine that this frame exists among the frames after the N-th frame. Further, when the selection unit 132 determines that this frame exists, the selection unit 132 may select the (N + 1)-th frame from among the frames whose similarity is less than the predetermined threshold. For example, when there is only one frame whose similarity is less than the predetermined threshold, the selection unit 132 may select that frame whose similarity is less than the predetermined threshold as the (N + 1)-th frame. Also, when there are two or more frames whose similarity is less than the predetermined threshold, the selection unit 132 may select, as the (N + 1)-th frame, the frame that is displayed at a time closer to the time when the N-th frame is displayed from among the two or more frames whose similarity is less than the predetermined threshold. For example, the selection unit 132 may calculate a value obtained by subtracting the time when the N-th frame is displayed from the time when each of the two or more frames whose similarity is less than the predetermined threshold is displayed. Subsequently, the selection unit 132 may select, as the (N + 1)-th frame, the frame for which the calculated value is the smallest from among the two or more frames whose similarity is less than the predetermined threshold. Also, when the selection unit 132 selects the (N + 1)-th frame, the selection unit 132 may output the (N + 1)-th frame to the generation unit 133. Further, the generation unit 133 acquires the (N + 1)-th frame from the selection unit 132. Also, when the generation unit 133 acquires the (N + 1)-th frame, the generation unit 133 generates a description text for the (N + 1)-th frame that describes the content of the (N + 1)-th frame.

[0028] As described above, the selection unit 132 selects this frame, which is the frame to be the target for generating the image description text, from among the frames constituting the video to be processed. Specifically, the selection unit 132 calculates the frame similarity between this frame (the Nth frame) selected one frame before and a frame after the frame selected one frame before, determines whether there is a frame whose frame similarity is lower than a predetermined threshold, and if it is determined that there is a frame whose frame similarity is lower than the predetermined threshold, selects the frame whose frame similarity is lower than the predetermined threshold as this frame (the (N + 1)th frame). For example, the selection unit 132 calculates, as the frame similarity, the similarity between the feature information indicating the features of the frame selected one frame before and the feature information indicating the features of a frame after the frame selected one frame before. Here, a frame after the frame selected one frame before refers to a frame that is displayed after the frame selected one frame before in the video to be processed. In other words, a frame after the frame selected one frame before refers to a frame that is displayed at a time after the time when the frame selected one frame before is displayed in the video to be processed.

[0029] On the other hand, if the selection unit 132 determines that there is no frame whose frame similarity is lower than the predetermined threshold, it may end the process. For example, if the selection unit 132 determines that there is no frame among the frames after the Nth frame whose calculated similarity is lower than the predetermined threshold, it may determine that there is no such frame among the frames after the Nth frame. Also, if the selection unit 132 determines that there is no such frame among the frames after the Nth frame, it may end the process.

[0030] In addition, the similarity between a predetermined frame and other frames in a video being below a predetermined threshold means that the predetermined frame and the other frames are not similar. Also, the fact that the predetermined frame and the other frames are not similar means that the scene (also referred to as a scene) corresponding to the predetermined frame and the scene corresponding to the other frames are not similar. Further, the fact that the scene corresponding to the predetermined frame and the scene corresponding to the other frames are not similar means that the scene in the video has switched between the predetermined frame and the other frames. That is, the information processing apparatus 100 can determine the switching of the scene in the video to be processed based on the similarity between a past frame (for example, the Nth frame) that is a frame for which a past description text (for example, the Nth frame description text) explaining the content of a frame before this frame was generated and a frame after the past frame. Also, after determining the switching of the scene in the video to be processed, the information processing apparatus 100 can select the frames corresponding to each scene in the video to be processed as frames to be targets for generating image description texts.

[0031] (Generation unit 133) The generation unit 133 may acquire the machine learning model output from the acquisition unit 131. For example, the generation unit 133 may acquire a vision - language model. FIG. 2 is a diagram for explaining the vision - language model 1 according to the embodiment. The generation unit 133 may acquire the vision - language model 1 learned based on learning data including a pair of an image and an image description text corresponding to the image. In FIG. 2, a state where the vision - language model 1 is learned based on a pair of an image 200 and an image description text 300 (not shown) corresponding to the image 200 is shown. The vision - language model 1 processes a sentence in units called tokens. In FIG. 2, the vision - language model 1 divides the image description text 300 into k tokens, the first token A1, the second token A2, …, the kth (k is a natural number) token Ak, and reads the tokens from the first token to the kth token in order. For example, the vision - language model 1 is learned to output the first token A1 following the start token A0 when the image 200 and the start token A0 are input. For example, the start token A0 is "<s>It may be something like "」. Also, when an image 200, a start token A0, and a first token A1 are input to the vision-language model 1, the vision-language model 1 is trained to output a second token A2 following the first token A1. Similarly, when the start token A0 and the first token A1 to the (k - 1)th token A(k - 1) are input, the vision-language model 1 is trained to output a kth token Ak following the (k - 1)th token A(k - 1). The generation unit 133 may obtain the vision-language model 1 trained to predict the correct next token from the image and the given (intermediate) tokens. For example, the generation unit 133 may obtain the vision-language model 1 which is a machine learning model trained to estimate the next token from the sequence of tokens being generated. For example, the generation unit 133 may obtain the vision-language model 1 which is a machine learning model trained to estimate and output the next token from the input image and the token sequence. For example, the generation unit 133 may obtain the vision-language model 1 such as CoCa, BLIP, BLIP2, GIT, etc.

[0032] FIG. 3 is a diagram showing an example of generation processing by the information processing apparatus 100 according to the embodiment. In FIG. 3, the generation unit 133 may obtain the first frame 211 output from the selection unit 132. In FIG. 3, the first frame 211 may be the first frame. The first frame 211 may be, for example, an image showing a person washing dishes in the kitchen. Also, when the generation unit 133 obtains the first frame 211, the generation unit 133 may generate a first frame description text 311A for explaining the content of the first frame 211 by inputting the start token A0 and the first frame 211 to the vision-language model 1. For example, the generation unit 133 may generate the first frame description text 311A which is the sentence "A person is washing dishes".

[0033] Further, when generating the first frame description text 311A, the generation unit 133 may acquire the second frame 212 output from the selection unit 132. The second frame 212 may be, for example, an image showing a person putting away dishes in a cupboard or taking them out. Further, when acquiring the second frame 212, the generation unit 133 may generate the second frame description text 312 for explaining the content of the second frame 212 by inputting the first frame description text 311A and the second frame 212 into the vision-language model 1. In this way, the generation unit 133 inputs, into the machine learning model (for example, the vision-language model 1), the current frame (for example, the second frame 212) which is the frame to be generated with a description text among the frames constituting the video to be processed, and the past description text (for example, the first frame description text 311A) for explaining the content of the frame before the current frame, thereby generating the current frame description text (for example, the second frame description text 312) for explaining the content of the current frame. Specifically, the generation unit 133 may further input the pre-prepared conjunction 411 into the vision-language model 1 to generate the second frame description text 312. For example, the pre-prepared conjunction 411 may be "and then,". Note that the pre-prepared conjunction 411 is not limited to "and then," and may be, for example, "but," etc. For example, the generation unit 133 may generate the second frame description text 312 by inputting into the vision-language model 1 the sentence "A person is washing dishes and then," which is the sentence with the pre-prepared conjunction 411 connected after the first frame description text 311A and the second frame 212. For example, the generation unit 133 may generate the second frame description text 312 which is the sentence "putting away dishes in the cupboard".

[0034] Here, from only the second frame 212, it is unclear whether a person is putting the dish in the cupboard or taking it out. In contrast, the first-frame description text 311A is information indicating the content of the scene corresponding to the first frame 211 before the second frame 212. For example, the first-frame description text 311A is information indicating that a person is washing a dish. Also, in the video to be processed, information indicating the content of the scene corresponding to a frame before a predetermined frame is information indicating the (past) context in the video to be processed. Therefore, the generation unit 133 can generate the second-frame description text 312 based on the content of the scene corresponding to the first frame 211 before the second frame 212 by inputting the first-frame description text 311A together with the second frame 212 into the vision-language model 1. For example, the generation unit 133 can generate the second-frame description text 312 indicating that a person is putting the dish in the cupboard (after washing the dish) based on the context that a person is washing a dish. Also, the generation unit 133 can generate the second-frame description text 312 following the pre-prepared conjunction 411 by inputting the pre-prepared conjunction 411 into the vision-language model 1.

[0035] Further, the generation unit 133 generates a second video description 312A including the first main frame description 311A and the second main frame description 312. In this way, the generation unit 133 generates a video description (e.g., the second video description 312A) including a past description (e.g., the first main frame description 311A) and the current frame description (e.g., the second main frame description 312). Specifically, when the generation unit 133 generates the second main frame description 312, it connects the first main frame description 311A and the second main frame description 312 with a pre-prepared conjunction 411, thereby generating the second video description 312A including the first main frame description 311A and the second main frame description 312. Here, the second video description 312A can be regarded as a video description that explains the content of the second video composed of frames up to the second main frame 212 among the frames constituting the video to be processed. Here, the second video may be a video composed of frames from the first frame to the second main frame 212 among the frames constituting the video to be processed. For example, the generation unit 133 may generate the second video description 312A by arranging the pre-prepared conjunction 411 behind the first main frame description 311A and arranging the second main frame description 312 behind the pre-prepared conjunction 411 to form a sentence. The generation unit 133 may generate the second video description 312A by connecting the first main frame description 311A and the second main frame description 312 with the pre-prepared conjunction 411. For example, the generation unit 133 may generate the second video description 312A which is a sentence "A person is washing dishes and then, putting away dishes in the cupboard".

[0036] Further, when generating the second video description text 312A, the generation unit 133 may acquire the third key frame 213 output from the selection unit 132. The third key frame 213 may be, for example, an image showing a person cleaning the kitchen. Also, when the generation unit 133 acquires the third key frame 213, the generation unit 133 may input the second video description text 312A and the third key frame 213 into the vision-language model 1 to generate a third key frame description text 313 that describes the content of the third key frame 213. In this way, the generation unit 133 inputs, into the machine learning model (for example, the vision-language model 1), the key frame (for example, the third key frame 213) that is the target frame for generating the description text among the frames constituting the video to be processed, and the past description text (for example, the second video description text 312A) that describes the content of the frame before this key frame, thereby generating a key frame description text (for example, the third key frame description text 313) that describes the content of this key frame. Specifically, the generation unit 133 may further input a pre-prepared conjunction 412 into the vision-language model 1 to generate the third key frame description text 313. For example, the pre-prepared conjunction 412 may be "and then,". Note that the pre-prepared conjunction 412 is not limited to "and then," and may be, for example, "but," etc. For example, the generation unit 133 may input into the vision-language model 1 the sentence "A person is washing dishes and then, putting away dishes in the cupboard and then," which is a sentence with the pre-prepared conjunction 412 connected after the second video description text 312A, and the third key frame 213 to generate the third key frame description text 313. For example, the generation unit 133 may generate the third key frame description text 313 which is the sentence "cleaned the kitchen".

[0037] Here, the second video description text 312A is a text in which the first main frame description text 311A and the second main frame description text 312 are connected. In other words, the second video description text 312A is a text that includes information indicating the content of the scene corresponding to the first main frame 211 (a scene where a person is washing dishes) and information indicating the content of the subsequent scene corresponding to the second main frame 212 (a scene where (a person) is putting the dishes away in the cupboard). At this time, for example, if the second main frame description text 312 is connected after the first main frame description text 311A, the second video description text 312A will be a text arranged in the order of scene occurrence. Conversely, for example, if the first main frame description text 311A is connected after the second main frame description text 312, the second video description text 312A will be a text indicating the causal relationship between the scenes. In other words, the second video description text 312A is information indicating the context formed by the stacking of each scene up to the second main frame 212. Therefore, the generation unit 133 inputs the second video description text 312A together with the third main frame 213 into the visual language model 1, so as to generate the third main frame description text 313 based on the context formed by the stacking of each scene up to the second main frame 212 before the third main frame 213. For example, the generation unit 133 can generate the third main frame description text 313 indicating the process of a person cleaning the kitchen by washing the dishes and putting them away in the cupboard based on the context that a person is washing the dishes and then (a person) is putting the dishes away in the cupboard. Also, for example, the generation unit 133 can generate the third main frame description text 313 indicating the causal relationship that a person is putting the dishes away in the cupboard because the dishes have been washed. Further, the generation unit 133 can generate the third main frame description text 313 connected by the pre-prepared conjunction 412 by inputting the pre-prepared conjunction 412 into the visual language model 1.

[0038] The generation unit 133 generates a third video description text 313A including the second video description text 312A and the third frame description text 313. In this way, the generation unit 133 generates a video description text (for example, the third video description text 313A) including a past description text (for example, the second video description text 312A) and the current frame description text (for example, the third frame description text 313). Specifically, when the generation unit 133 generates the third frame description text 313, it connects the second video description text 312A and the third frame description text 313 with a pre-prepared conjunction 412, thereby generating the third video description text 313A including the second video description text 312A and the third frame description text 313. Here, the third video description text 313A can be regarded as a video description text that describes the content of the third video composed of the frames up to the third frame 213 among the frames constituting the video to be processed. Here, the third video may be a video composed of the frames from the first frame to the third frame 213 among the frames constituting the video to be processed. For example, the generation unit 133 may generate the third video description text 313A by arranging the pre-prepared conjunction 412 behind the second video description text 312A and arranging the third frame description text 313 behind the pre-prepared conjunction 412. The generation unit 133 may generate the third video description text 313A by connecting the second video description text 312A and the third frame description text 313 with the conjunction 412. For example, the generation unit 133 may generate the third video description text 313A which is the sentence "A person is washing dishes and then, putting away dishes in the cupboard and then, cleaned the kitchen".

[0039] Here, the third video description text 313A is a text in which the second video description text 312A and the third frame description text 313 are connected. In other words, the third video description text 313A is a text that includes information indicating the content of the scenes (a scene where a person is washing dishes and a scene where a person is putting the dishes away in the cupboard) included in the second video description text 312A, and information indicating the content of the subsequent scene (a scene of cleaning the kitchen) corresponding to the third frame 213. At this time, for example, if the third frame description text 313 is connected after the second video description text 312A, the third video description text 313A becomes a text arranged in the order of occurrence of the scenes. Conversely, for example, if the second video description text 312A is connected after the third frame description text 313, the third video description text 313A becomes a text indicating the causal relationship of the scenes. In other words, the third video description text 313A is information indicating the context formed by the stacking of each scene up to the third frame 213. Therefore, the generation unit 133 can generate a fourth frame description text (not shown) based on the context formed by the stacking of each scene up to the third frame 213, which is before the fourth frame (not shown), by inputting the third video description text 313A into the visual language model 1 together with the fourth frame (not shown).

[0040] Similarly, when the generation unit 133 generates the Nth video description text, it may acquire the (N + 1)th frame output from the selection unit 132. N may be a natural number of 2 or more. Here, the Nth video description text can be regarded as a video description text that describes the content of the Nth video composed of the frames up to the Nth frame among the frames constituting the video to be processed. Here, the Nth video may be a video composed of the frames from the first frame to the Nth frame among the frames constituting the video to be processed. Further, when the generation unit 133 acquires the (N + 1)th frame, it may generate the (N + 1)th frame description text that describes the content of the (N + 1)th frame by inputting the Nth video description text and the (N + 1)th frame into the vision language model 1. In this way, the generation unit 133 inputs, into the machine learning model (for example, the vision language model 1), the current frame (for example, the (N + 1)th frame) which is the frame to be the target of generating the description text among the frames constituting the video to be processed, and the past description text (for example, the Nth video description text) that describes the content of the frames before the current frame, and thereby generates the current frame description text (for example, the (N + 1)th frame description text) that describes the content of the current frame. Specifically, the generation unit 133 may further input a pre-prepared conjunction into the vision language model 1 to generate the (N + 1)th frame description text. The generation unit 133 may generate the (N + 1)th frame description text by inputting into the vision language model 1 the sentence with the pre-prepared conjunction arranged after the Nth video description text and the (N + 1)th frame. For example, the pre-prepared conjunction may be "and then,". Note that the pre-prepared conjunction is not limited to "and then," and may be, for example, "but," or the like.

[0041] Here, the N-th video description text is a text in which the (N - 1)-th video description text and the N-th frame description text are connected. In other words, the N-th video description text is a text that includes information indicating the content of the scene corresponding to the first frame, information indicating the content of the scene corresponding to the subsequent second frame,..., and information indicating the content of the scene corresponding to the subsequent N-th frame. At this time, for example, if the N-th frame description text is connected after the (N - 1)-th video description text, the N-th video description text becomes a text arranged according to the occurrence order of the scenes. Conversely, for example, if the (N - 1)-th video description text is connected after the N-th frame description text, the N-th video description text becomes a text indicating the causal relationship of the scenes. In other words, the N-th video description text is information indicating the context formed by the stacking of each scene from the first frame to the N-th frame. Therefore, the generation unit 133 can generate the (N + 1)-th frame description text based on the context formed by the stacking of each scene from the first frame to the N-th frame, which is before the (N + 1)-th frame, by inputting the N-th video description text together with the (N + 1)-th frame into the visual language model 1. In addition, the generation unit 133 can generate the (N + 1)-th frame description text following the pre-prepared conjunction by inputting the pre-prepared conjunction into the visual language model 1.

[0042] Further, the generation unit 133 generates a (N + 1)-th video description text including the N-th video description text and the (N + 1)-th frame description text. N may be a natural number of 2 or more. In this way, the generation unit 133 generates a video description text (e.g., the (N + 1)-th video description text) including a past description text (e.g., the N-th video description text) and a frame description text (e.g., the (N + 1)-th frame description text). Specifically, when the generation unit 133 generates the (N + 1)-th frame description text, the generation unit 133 connects the N-th video description text and the (N + 1)-th frame description text with a pre-prepared conjunction, thereby generating a (N + 1)-th video description text including the N-th video description text and the (N + 1)-th frame description text. Here, the (N + 1)-th video description text can be regarded as a video description text that describes the content of the (N + 1)-th video composed of the frames up to the (N + 1)-th frame among the frames constituting the video to be processed. Here, the (N + 1)-th video may be a video composed of the frames from the first frame to the (N + 1)-th frame among the frames constituting the video to be processed. For example, the generation unit 133 may generate the (N + 1)-th video description text by arranging a pre-prepared conjunction behind the N-th video description text and arranging the (N + 1)-th frame description text behind the pre-prepared conjunction. The generation unit 133 may generate the (N + 1)-th video description text by connecting the N-th video description text and the (N + 1)-th frame description text with a pre-prepared conjunction.

[0043] As described above, the generation unit 133 inputs, into the machine learning model (e.g., the vision language model 1), the current frame (e.g., the (N + 1)-th frame), which is the frame to be used for generating the description text among the frames constituting the video to be processed, and the past description text (e.g., the N-th video description text) that describes the content of the frames before the current frame. By doing so, the generation unit 133 generates the current frame description text (e.g., the (N + 1)-th frame description text) that describes the content of the current frame, and generates a video description text (e.g., the (N + 1)-th video description text) that includes the past description text and the current frame description text. As a result, the information processing apparatus 100 can generate the current frame description text based on the context formed by the accumulation of each scene up to the current frame, which is the frame to be used for generating the image description text, before the current frame. Also, the information processing apparatus 100 can generate a video description text based on the context formed by the accumulation of each scene up to the current frame in the video to be processed. Therefore, the information processing apparatus 100 can generate a video description text that reflects the context formed by the accumulation of each scene in the video to be processed.

[0044] Further, the generation unit 133 further inputs a pre-prepared conjunction (e.g., "and then,") into the machine learning model, generates the current frame description text (e.g., the (N + 1)-th frame description text), and connects the past description text (e.g., the N-th video description text) and the current frame description text with the pre-prepared conjunction, thereby generating a video description text (e.g., the (N + 1)-th video description text). As a result, the information processing apparatus 100 can generate a video description text in which the image description texts corresponding to each scene in the video to be processed are appropriately connected by the pre-prepared conjunction.

[0045] [3. Processing Procedure] FIG. 4 is a flowchart showing a procedure of information processing by the information processing apparatus 100 according to the embodiment. In FIG. 4, an acquisition unit 131 of the information processing apparatus 100 acquires a video to be processed (step S11). Further, a selection unit 132 of the information processing apparatus 100 determines whether or not there is a current frame, which is a frame to be a target for generating an image description text, among the frames constituting the video to be processed (step S12). When the selection unit 132 determines that the current frame does not exist (step S12; No), the processing ends. On the other hand, when the selection unit 132 determines that the current frame exists (step S12; Yes), the selection unit 132 selects the current frame from among the frames constituting the video to be processed (step S13). Further, a generation unit 133 of the information processing apparatus 100 generates a current frame description text for describing the content of the current frame by inputting a past description text for describing the content of a frame before the current frame selected by the selection unit 132 and the current frame into a machine learning model (step S14). Subsequently, the generation unit 133 generates a video description text including the past description text and the current frame description text (step S15).

[0046] [4. Modification Example] The processing according to the above-described embodiment may be implemented in various different forms other than the above embodiment.

[0047] In the above-described embodiment, the generation unit 133 further inputs a pre-prepared conjunction into the machine learning model to generate the present frame description text, and connects the past description text and the present frame description text with the pre-prepared conjunction to generate the video description text. However, the present invention is not limited to this. Specifically, the generation unit 133 may generate the present frame description text including a conjunction for connecting the past description text to the present frame description text using the machine learning model, and generate the text in which the past description text and the present frame description text are connected by the conjunction as the video description text. For example, when the generation unit 133 generates the past description text (for example, the Nth video description text), the generation unit 133 may acquire the present frame (for example, the (N + 1)th present frame) output from the selection unit 132. Further, when the generation unit 133 acquires the present frame, the generation unit 133 may input the past description text and the present frame into the vision language model to generate the present frame description text (for example, the (N + 1)th present frame description text) including a conjunction for connecting the past description text to the present frame description text. Further, when the generation unit 133 generates the present frame description text including a conjunction for connecting the past description text to the present frame description text, the generation unit 133 may generate the text in which the past description text and the present frame description text are connected by the conjunction as the video description text. For example, the generation unit 133 may generate, as the video description text, a text in which the present frame description text including a conjunction for connecting the past description text to the present frame description text is arranged after the past description text. Thereby, the information processing apparatus 100 can generate a video description text in which the image description texts corresponding to the respective scenes in the video to be processed are appropriately connected by the conjunction generated by the machine learning model.

[0048] In addition, when the length of the video description text exceeds a predetermined length, the generation unit 133 generates a summary text obtained by summarizing the video description text using a text summarization model. The generation unit 133 may generate a summary text obtained by summarizing the video description text of this video using a text summarization model when the length of the video description text of this video exceeds a predetermined length. For example, when generating the video description text of this video, the generation unit 133 may determine whether the length of the video description text of this video exceeds a predetermined length. When the generation unit 133 determines that the length of the video description text of this video does not exceed a predetermined length, the generation unit 133 may end the process. On the other hand, when the generation unit 133 determines that the length of the video description text of this video exceeds a predetermined length, the generation unit 133 may refer to the storage unit 120 to obtain a text summarization model. Subsequently, the generation unit 133 may input the video description text of this video into the text summarization model to generate a summary text obtained by summarizing the video description text of this video.

[0049] In addition, in the above-described embodiment, the case where the selection unit 132 calculates the similarity as the frame similarity between the feature information indicating the features of the past frame and the feature information indicating the features of the frame after the past frame has been described. However, the present invention is not limited to this. For example, the selection unit 132 may calculate the similarity as the frame similarity between the vector obtained by arranging the pixel values of the past frame and the vector obtained by arranging the pixel values of each frame after the past frame. Alternatively, the selection unit 132 may calculate the similarity as the frame similarity between the average of the pixel values of the past frame and the average of the pixel values of each frame after the past frame.

[0050] 〔5. Effects〕 As described above, the information processing apparatus 100 according to the embodiment includes a machine learning model that generates a video description text, which is a text describing the content included in the video to be processed, from the video to be processed. The information processing apparatus 100 includes an acquisition unit 131 and a generation unit 133. The acquisition unit 131 acquires the video to be processed. The generation unit 133 inputs, to the machine learning model, a current frame, which is a frame to be a target for generating a description text among the frames constituting the video to be processed, and a past description text that describes the content of a frame earlier than the current frame, thereby generating a current frame description text that describes the content of the current frame, and generates a video description text including the past description text and the current frame description text.

[0051] Thereby, the information processing apparatus 100 can generate a current frame description text based on the context formed by stacking each scene up to the current frame, which is a frame to be a target for generating an image description text, earlier than the current frame. Further, the information processing apparatus 100 can generate a video description text based on the context formed by stacking each scene up to the current frame in the video to be processed. Therefore, the information processing apparatus 100 can be enabled to generate a video description text reflecting the context formed by stacking each scene in the video to be processed. Also, since the information processing apparatus 100 can be enabled to generate a video description text reflecting the context formed by stacking each scene in the video to be processed, it can contribute to the achievement of Goal 9, "Build the infrastructure for industry and innovation," of the Sustainable Development Goals (SDGs).

[0052] Also, the information processing apparatus 100 further includes a selection unit 132. The selection unit 132 selects the current frame from among a plurality of frames constituting the video to be processed. The selection unit 132 calculates the frame similarity between a past frame, which is a frame for which the past description text has been generated, and a plurality of frames later than the past frame, and selects, as the current frame, a frame for which the frame similarity is less than a predetermined threshold value.

[0053] As described above, the similarity between a predetermined frame and other frames in a video being below a predetermined threshold means that the predetermined frame and the other frames are not similar. Also, the fact that the predetermined frame and the other frames are not similar means that the scene corresponding to the predetermined frame and the scenes corresponding to the other frames are not similar. Further, the fact that the scene corresponding to the predetermined frame and the scenes corresponding to the other frames are not similar means that the scene in the video has switched between the predetermined frame and the other frames. That is, the information processing apparatus 100 can determine the switching of scenes in the video to be processed based on the similarity between a past frame and a frame after the past frame. Also, the information processing apparatus 100 can select the current frame corresponding to each scene in the video to be processed after determining the switching of scenes in the video to be processed. Also, the information processing apparatus 100 can be made capable of generating a current frame description text corresponding to each scene in the video to be processed from the current frame corresponding to each scene in the video to be processed. Also, the information processing apparatus 100 can be made capable of generating a video description text corresponding to the video to be processed based on the current frame description text corresponding to each scene in the video to be processed. Therefore, the information processing apparatus 100 can be made capable of generating a video description text that appropriately reflects the context formed by the stacking of each scene in the video to be processed.

[0054] Also, the selection unit 132 calculates, as the frame similarity, the similarity between the feature information indicating the features of the past frame and the feature information indicating the features of a frame after the past frame.

[0055] Thereby, the information processing apparatus 100 can determine the switching of scenes in the video to be processed based on the similarity between the feature information indicating the features of the past frame and the feature information indicating the features of a frame after the past frame. Thereby, the information processing apparatus 100 can more accurately determine the switching of scenes in the video to be processed.

[0056] Further, the generation unit 133 further inputs a pre-prepared conjunction into the machine learning model, generates the present frame description text, and connects the past description text and the present frame description text with the pre-prepared conjunction to generate video description text.

[0057] Thereby, the information processing apparatus 100 can generate video description text in which the image description texts corresponding to the respective scenes in the video to be processed are appropriately connected by the pre-prepared conjunction.

[0058] Further, the generation unit 133 uses the machine learning model to generate the present frame description text including the conjunction for connecting the past description text to the present frame description text, and generates the text in which the past description text and the present frame description text are connected by the conjunction as the video description text.

[0059] Thereby, the information processing apparatus 100 can generate video description text in which the image description texts corresponding to the respective scenes in the video to be processed are appropriately connected by the conjunction generated by the machine learning model.

[0060] Further, when the length of the video description text exceeds a predetermined length, the generation unit 133 uses the text summarization model to generate a summary text obtained by summarizing the video description text.

[0061] Thereby, the information processing apparatus 100 can generate video description text with an appropriate text length.

[0062] [6. Hardware Configuration] Further, the information processing apparatus 100 according to the above-described embodiment is realized by a computer 1000 having a configuration as shown in FIG. 5, for example. FIG. 5 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing apparatus 100. The computer 1000 includes a CPU 1100, a RAM 1200, a ROM 1300, an HDD 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.

[0063] The CPU 1100 operates based on programs stored in the ROM 1300 or the HDD 1400 and controls each part. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 is started up, programs dependent on the hardware of the computer 1000, and the like.

[0064] The HDD 1400 stores programs executed by the CPU 1100, data used by such programs, and the like. The communication interface 1500 receives data from other devices via a predetermined communication network, sends it to the CPU 1100, and sends data generated by the CPU 1100 to other devices via the predetermined communication network.

[0065] The CPU 1100 controls output devices such as displays and printers, and input devices such as keyboards and mice, via the input / output interface 1600. The CPU 1100 acquires data from the input devices via the input / output interface 1600. Also, the CPU 1100 outputs the generated data to the output devices via the input / output interface 1600.

[0066] The media interface 1700 reads a program or data stored in the recording medium 1800 and provides it to the CPU 1100 via the RAM 1200. The CPU 1100 loads such a program from the recording medium 1800 onto the RAM 1200 via the media interface 1700 and executes the loaded program. The recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc), a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.

[0067] For example, when the computer 1000 functions as the information processing apparatus 100 according to the embodiment, the CPU 1100 of the computer 1000 realizes the functions of the control unit 130 by executing the program loaded on the RAM 1200. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800. As another example, these programs may be acquired from other devices via a predetermined communication network.

[0068] As described above, some of the embodiments of the present application have been described in detail with reference to the drawings. However, these are merely examples, and the present invention can be implemented in other forms with various modifications and improvements based on the knowledge of those skilled in the art, including the aspects described in the column of the disclosure of the invention.

[0069] [7. Others] In addition, among the respective processes described in the above embodiments and modification examples, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the illustrated information.

[0070] In addition, each component of each illustrated device is a functional concept, and does not necessarily have to be physically configured as illustrated. That is, the specific form of the distribution and integration of each device is not limited to that shown, and all or part of it can be functionally or physically distributed and integrated in arbitrary units according to various loads, usage situations, etc.

[0071] In addition, the above-described embodiments and modification examples can be appropriately combined as long as the processing contents do not conflict.

Description of Reference Numerals

[0072] 100 Information processing apparatus 110 Communication Unit 120 Memory Unit 130 Control Unit 131 Acquisition Unit 132 Selection Unit 133 Generation Unit< / s>

Claims

1. An information processing apparatus including a machine learning model that generates a video description text, which is a text describing the content included in the video to be processed, from the video to be processed, an acquisition unit that acquires the video to be processed, a generation unit that generates a current frame description text for describing the content of the current frame by inputting the current frame, which is a frame to be a target for generating a description text among the frames constituting the video to be processed, and a past description text for describing the content of a frame earlier than the current frame, into the machine learning model, and generates the video description text including the past description text and the current frame description text; An information processing apparatus comprising:

2. The information processing apparatus according to claim 1, further comprising a selection unit that selects the current frame from among a plurality of frames constituting the video to be processed, wherein the selection unit calculates a frame similarity between a past frame, which is a frame for generating the past description text, and a plurality of frames later than the past frame, and selects, as the current frame, a frame in which the frame similarity is less than a predetermined threshold value. The information processing apparatus according to claim 1.

3. The selection unit according to claim 2, calculates, as the frame similarity, a similarity between feature information indicating features of the past frame and feature information indicating features of a frame later than the past frame. The information processing apparatus according to claim 2.

4. The generation unit according to claim 1, further inputs a previously prepared conjunction into the machine learning model to generate the current frame description text, and generates the video description text by connecting the past description text and the current frame description text with the previously prepared conjunction. The information processing apparatus according to claim 1.

5. The generation unit according to claim 1, generates the current frame description text including a conjunction for connecting the past description text to the current frame description text using the machine learning model, and generates, as the video description text, a sentence in which the past description text and the current frame description text are connected by the conjunction. The information processing apparatus according to claim 1.

6. The generation unit according to claim 1, when the length of the video description text exceeds a predetermined length, generates a summary text obtained by summarizing the video description text using a text summarization model. The information processing apparatus according to claim 1.

7. An information processing program executed by an information processing apparatus including a machine learning model that generates a video description text, which is a text describing the content included in the video to be processed, from the video to be processed, An acquisition procedure for acquiring the video to be processed; A generation procedure for generating a description text for the current frame that describes the content of the current frame by inputting the current frame, which is a frame to be a target for generating a description text among the frames constituting the video to be processed, and a past description text that describes the content of a frame earlier than the current frame, into the machine learning model, and generating the video description text including the past description text and the description text for the current frame; An information processing program for causing the information processing apparatus to execute the above.

Citation Information

Patent Citations

  • Video retrieval method based on natural language description

    CN111651635A

  • Automated generation and use of building videos with accompanying narration from analysis of acquired images and other building information

    EP4328866A1

  • Abnormality monitoring system

    JP2018101317A

  • Correlating videos and sentences

    US20140369596A1

  • Unified referring video object segmentation network

    US20210383171A1