Caption selection method, caption selection program, and caption selection device
The caption selection method addresses the challenge of accurately selecting video captions by using an image caption generation model and selecting captions based on similarity with frame images, ensuring accurate representation of video content.
Patent Information
- Application Number
- PCT/JP2024/043465
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-09
- Publication Date
- 2025-06-26
AI Technical Summary
Existing methods for generating video captions, such as using video caption generation models or image caption generation models, face challenges in accurately selecting captions that match the content of videos, particularly for short scenes where frequently occurring words may not be included.
A caption selection method that extracts frame images from a video and inputs them into an image caption generation model to generate captions. The method then selects a specific caption based on the similarity between the generated captions and the frame images, using both appearance frequency and feature vector similarities.
This approach allows for the appropriate selection of captions that accurately represent the content of videos, even in cases where short scenes may lack frequently occurring words, thereby improving caption accuracy and relevance.
Smart Images

Figure JP2024043465_26062025_PF_FP_ABST
Abstract
Description
Caption selection method, caption selection program, and caption selection device
[0001] The present disclosure relates to a caption selection method, a caption selection program, and a caption selection device.
[0002] In recent years, video caption generation models that generate captions for videos have been disclosed. However, training these video caption generation models requires the preparation of dedicated training data, which can lead to high generation costs. For example, annotating videos with captions requires manual review of the video content and the creation of captions. This work also requires the annotators' labor costs, training time, and work time.
[0003] Here, instead of the video caption generation model, it is possible to generate video captions using an image caption generation model that generates captions for images. However, when video captions are generated using an image caption generation model, there is a risk that captions that do not match the content of the video will be assigned.
[0004] For example, Patent Document 1 discloses a sentence selection device that selects a sentence that represents the outline of content from a set of sentences that represent the substance of the content. This device assigns captions to each scene based on frequently occurring words.
[0005] Japanese Patent Application Laid-Open No. 2021-141364
[0006] However, because the sentence selection device of Patent Document 1 selects captions based on frequently occurring words, there is a risk that it may not be possible to appropriately select captions that correspond to the content of the video. For example, in the case of a short scene, there is a risk that frequently occurring words will not be included in the caption. In this case, if captions are selected based on frequently occurring words, captions that correspond to the content of the video will not be selected.
[0007] The present disclosure aims to provide a caption selection method, a caption selection program, and a caption selection device that appropriately select captions according to the content of a video.
[0008] The caption selection method of the present disclosure includes extracting a plurality of frame images from a video, inputting the plurality of frame images into an image caption generation model to generate a plurality of captions corresponding to the plurality of frame images, and selecting a specific caption from the plurality of captions based on a first similarity between the plurality of captions and a second similarity between the plurality of frame images and the plurality of captions.
[0009] The caption selection program of the present disclosure causes a computer to extract multiple frame images from a video, input the multiple frame images into an image caption generation model to generate multiple captions corresponding to the multiple frame images, and select a specific caption from the multiple captions based on a first similarity between the multiple captions and a second similarity between the multiple frame images and the multiple captions.
[0010] The caption selection device of the present disclosure includes an extraction unit that extracts multiple frame images from a video, a generation unit that inputs the multiple frame images into an image caption generation model to generate multiple captions corresponding to the multiple frame images, and a selection unit that selects a specific caption from the multiple captions based on a first similarity between the multiple captions and a second similarity between the multiple frame images and the multiple captions.
[0011] According to the present disclosure, it is possible to appropriately select captions according to the content of a video.
[0012] FIG. 1 is a diagram showing how a caption is generated by an image caption generation model. FIG. 2 is a diagram showing an overview of a first embodiment. FIG. 3 is a block diagram showing the functional configuration of a caption selection device. FIG. 4 is a block diagram showing the hardware configuration of a caption selection device. FIG. 5 is a flowchart showing caption selection processing. FIG. 6 is a diagram showing how a first similarity is calculated. FIG. 7 is a diagram showing how a second similarity is calculated. FIG. 8 is a diagram showing how a specific caption is selected. FIG. 9 is a diagram showing an overview of a second embodiment. FIG. 10 is a block diagram showing the functional configuration of a re-learning device. FIG. 11 is a flowchart showing re-learning processing.
[0013] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0014] First Embodiment [Overview] It has been proposed to generate captions corresponding to a plurality of frame images F constituting a moving image by inputting the frame images F into an image caption generation model, as shown in Fig. 1. In this case, the image caption generation model may generate a caption Ca that does not match the content of a portion of the frame images Fa among the plurality of frame images F.
[0015] 2 , the caption selection device of the present disclosure extracts a plurality of frame images F from a video, inputs the plurality of frame images F into an image caption generation model 9, and generates a plurality of captions C corresponding to the plurality of frame images F. Then, the caption selection device selects a specific caption C1 from the plurality of captions C based on a first similarity between the plurality of captions C and a second similarity between the plurality of frame images F and the plurality of captions C.
[0016] [Caption Selection Device] Next, a detailed description will be given of the functional configuration of the caption selection device 1. As shown in FIG.
[0017] The extraction unit 2 extracts a plurality of frame images F from the moving image. For example, the extraction unit 2 may extract frame images F included in a predetermined time period, that is, a predetermined number of frame images F, from the moving image.
[0018] The generation unit 3 inputs a plurality of frame images F into the image caption generation model 9 to generate a plurality of captions C corresponding to the plurality of frame images. Here, the image caption generation model 9 is a machine learning model trained using a dataset of images and captions for those images, and may be realized by, for example, a convolutional neural network. Furthermore, the captions C are sentences indicating the content of the video or frame images F, and may include, for example, explanatory text.
[0019] The selection unit 4 selects a specific caption C1 from the plurality of captions C based on a first similarity between the plurality of captions C and a second similarity between the plurality of frame images F and the plurality of captions C.
[0020] Next, we will explain in detail the hardware configuration of the caption selection device 1. For example, as shown in Figure 4, the caption selection device 1 has a storage device 5, a processor 6, a user interface (UI) device 7, and a communication device 8, which are interconnected via a bus B.
[0021] The programs or instructions for realizing the various functions and processes of the caption selection device 1 may be downloaded from any external device via a network, etc. Also, the programs or instructions for realizing the various functions and processes of the caption selection device 1 may be provided from a removable storage medium such as a CD-ROM (Compact Disk-Read Only Memory) or a flash memory.
[0022] The storage device 5 may be realized by one or more non-transitory storage media, such as random access memory, flash memory, or a hard disk drive, and may store installed programs or instructions as well as files or data used in the execution of the programs or instructions.
[0023] The processor 6 may be realized by one or more central processing units (CPUs), which may be configured with one or more processor cores, graphics processing units (GPUs), processing circuitry, etc. The processor 6 executes various functions and processes of the caption selection device 1 in accordance with data such as programs, instructions, programs, or parameters (e.g., parameters required to execute instructions) stored in the storage device 5.
[0024] The user interface device 7 may be composed of input devices such as a keyboard, mouse, camera, or microphone, output devices such as a display, speaker, headset, or printer, and input / output devices such as a smartphone, tablet, or touch panel.
[0025] The communication device 8 is realized by various communication circuits that execute wired and / or wireless communication processing with communication networks such as external devices, the Internet, a LAN (Local Area Network), and a cellular network.
[0026] It should be noted that the above-described hardware configuration is merely an example, and the caption selection device 1 according to the present disclosure may be realized by any other appropriate hardware configuration.
[0027] [Caption Selection Process] Next, the caption selection process performed by the caption selection device 1 will be described with reference to the flowchart shown in FIG.
[0028] First, in step S1, the extraction unit 2 shown in FIG. 3 extracts multiple frame images F from a video. For example, when a video is input by a user, the extraction unit 2 may extract a predetermined number of frame images F based on the playback time of the video. For example, if the video is composed of 30 frames per second (FPS), the extraction unit 2 will extract 30 frame images F if the extraction time is one second. In this case, the extraction unit 2 may sequentially extract multiple frame images F every predetermined time (e.g., one second). Alternatively, the extraction unit 2 may sequentially extract frame images F included in a predetermined time (e.g., one second) while shifting the start position of extraction (e.g., by sequentially shifting the start position of extraction by ten frame images F). The extraction unit 2 outputs the extracted multiple frame images F to the generation unit 3.
[0029] In this way, the extraction unit 2 extracts a predetermined number of frame images F based on the playback time of the video, which facilitates the processing described below (e.g., caption generation processing or selection processing). Note that the extraction unit 2 preferably sets the extraction time between 0.1 and 5 seconds, and more preferably sets it to 1 second, for example.
[0030] In step S2, the generation unit 3 generates a plurality of captions C corresponding to the plurality of frame images F from the plurality of frame images F extracted by the extraction unit 2. Specifically, as shown in FIG. 2 , the generation unit 3 inputs the plurality of frame images F into an image caption generation model 9, thereby generating a plurality of captions C corresponding to the plurality of frame images F. The generation unit 3 outputs the plurality of frame images F and the plurality of captions C to the selection unit 4.
[0031] When the selection unit 4 receives a plurality of frame images F and a plurality of captions C, it calculates a first similarity between the plurality of captions C and a second similarity between the plurality of frame images F and the plurality of captions C in step S3.
[0032] For example, as shown in FIG. 6 , the selection unit 4 may calculate the first similarity S1 based on the frequency of appearance of captions C that indicate the same content among a plurality of captions C. Specifically, the selection unit 4 categorizes the plurality of captions C based on the content of each caption C. For example, the selection unit 4 may categorize the plurality of captions C based on whether the sentences in the captions C completely match. This allows the selection unit 4 to accurately and quickly calculate the first similarity S1. Here, it is assumed that the selection unit 4 categorizes the plurality of captions C as "a man is cooking," "a woman is cooking," ... "a man is eating food."
[0033] Next, the selection unit 4 calculates, as a first similarity S1, the frequency of appearance of the classified caption C in the plurality of captions C. For example, the selection unit 4 calculates the frequency of appearance of "a man is cooking" to be 20 times, the frequency of appearance of "a woman is cooking" to be 3 times, ... and the frequency of appearance of "a man is eating food" to be 1 time.
[0034] In this way, the selection unit 4 calculates the first similarity S1 based on the appearance frequency of captions C that indicate the same content, and therefore can accurately calculate the first similarity S1. Note that the selection unit 4 classified the multiple captions C based on whether the sentences in the captions C completely match, but this is not limiting. For example, the selection unit 4 may also classify the multiple captions C based on the degree of similarity of the words used in the captions C.
[0035] 7, the selection unit 4 may calculate the second similarity S2 based on feature vectors T1 to T30 of a plurality of frame images F and feature vectors I1 to I30 of a plurality of captions C. For example, the selection unit 4 may calculate the second similarity S2 using a language-image model such as CLIP (Contrastive Language-Image Pre-training). Here, the language-image model is a machine learning model trained using a data set of language and image, and may be configured, for example, from a language and image embedding model.
[0036] Specifically, the selection unit 4 inputs a plurality of captions C into a text encoder, thereby converting the plurality of captions C into feature vectors T1 to T30 indicating their features. The selection unit 4 also inputs a plurality of frame images F into an image encoder, thereby converting the plurality of frame images F into feature vectors I1 to I30 indicating their features. The selection unit 4 then calculates the second similarity S2 based on the distance between the feature vectors T1 to T30 of the plurality of captions C and the feature vectors I1 to I30 of the plurality of frame images F. For example, the selection unit 4 may calculate, as the second similarity S2, a similarity score indicating the dot product of the feature vectors T1 to T30 and the feature vectors I1 to I30. In this case, the selection unit 4 calculates, as the second similarity S2, the similarity scores (I1·T1, I2·T2, I3·T3, ..., I30·T30) between corresponding frame images F and captions C.
[0037] In this way, the selection unit 4 calculates the second similarity S2 based on the feature vectors T1 to T30 of multiple frame images F and the feature vectors I1 to I30 of multiple captions C, and therefore can accurately calculate the second similarity S2.
[0038] Next, the selection unit 4 selects a specific caption C1 from the multiple captions C based on the calculated first similarity S1 and second similarity S2. For example, the selection unit 4 may compare the magnitude of the first similarity S1 and the second similarity S2 for each classified caption C and select a specific caption C1 based on the magnitude. Specifically, as shown in FIG. 8 , the selection unit 4 multiplies the appearance frequency indicating the first similarity S1 by the similarity score indicating the second similarity S2 for each classified caption C. Then, the selection unit 4 selects a specific caption C1, "A man is cooking," based on the magnitude of the multiplied value. At this time, if the similarity scores differ, the selection unit 4 may multiply the average similarity score by the appearance frequency. For example, if the 20 similarity scores (I1·T1, I2·T2, etc.) classified into the caption C "A man is cooking" have different values, the selection unit 4 may multiply the average of the 20 similarity scores by the appearance frequency.
[0039] In this way, the selection unit 4 selects a specific caption C1 from among the multiple captions C based on the first similarity S1 between the multiple captions C and the second similarity S2 between the multiple frame images F and the multiple captions C. This allows the selection unit 4 to appropriately select a caption C1 that matches the content of the video. At this time, the selection unit 4 includes in the selection conditions not only the first similarity S1 between the captions C but also the second similarity S2 between the frame image F and the caption C. This second similarity S2 indicates not just the similarity between the frame images F, but the similarity in content between the frame image F and the caption C. This allows the selection unit 4 to appropriately select a specific caption C1.
[0040] Furthermore, the selection unit 4 selects a specific caption C1 from among a plurality of captions C in accordance with the frequency of appearance of captions C that indicate the same content. This allows the selection unit 4 to more appropriately select a caption C1 that corresponds to the content of the video.
[0041] Furthermore, the selection unit 4 selects a specific caption C1 according to the magnitude of the second similarity S2 between the frame image F and the caption C. This allows the selection unit 4 to more appropriately select a caption C1 that corresponds to the content of the video.
[0042] When the selection unit 4 selects a specific caption C1, it assigns the specific caption C1 to a frame image F of the video. For example, when a video is recorded, the selection unit 4 may assign the specific caption C1 to multiple frame images F extracted by the extraction unit 2. Furthermore, when a video is played in real time, the selection unit 4 may assign the specific caption C1 to a frame image F that follows multiple frame images F extracted by the extraction unit 2. The selection unit 4 then repeats the selection of the specific caption C1 each time multiple frame images F are extracted by the extraction unit 2.
[0043] According to this embodiment, the selection unit 4 selects a specific caption C1 from among the multiple captions C based on the first similarity S1 between the multiple captions C and the second similarity S2 between the multiple frame images F and the multiple captions C. This allows the selection unit 4 to appropriately select a caption C1 that corresponds to the content of the video.
[0044] 9 , a re-learning device 21 re-trains an image caption generation model 9 using a data set D1 of frame images F and captions C that satisfy a positive example condition, among a plurality of frame images F and a plurality of captions C, as a positive example. The re-learning device 21 may also re-train an image caption generation model 9 using a data set D2 of frame images F and captions C that are not positive examples, among a plurality of frame images F and a plurality of captions C, as a negative example.
[0045] After retraining the image caption generation model 9, the retraining device 21 provides the generated retrained image caption generation model 22 to the caption selection device 1. As a result, the caption selection device 1 inputs a plurality of frame images F extracted from a video to the retrained image caption generation model 22. Then, the caption selection device 1 selects a specific caption C1 from the plurality of captions C output from the retrained image caption generation model 22.
[0046] [Re-learning Device] Fig. 10 shows the functional configuration of the relearning device 21. The relearning device 21 includes an acquisition unit 23 and a relearning unit 24. The acquisition unit 23 acquires, from the caption selection device 1, frame images F and captions C that satisfy a positive example condition, as a positive example dataset D1, from among a plurality of frame images F and a plurality of captions C. Here, the positive example condition is that the first similarity S1 is the highest and the second similarity S2 is equal to or greater than a predetermined threshold. The acquisition unit 23 may also acquire, from the caption selection device 1, frame images F and captions C other than the positive examples, from among the plurality of frame images F and a plurality of captions C, as a negative example dataset D2.
[0047] The re-learning unit 24 re-trains the image caption generation model 9 using the positive example dataset D1. The re-learning unit 24 may also re-train the image caption generation model 9 using the negative example dataset D2. Then, the re-learning unit 24 provides the generated re-trained image caption generation model 22 to the caption selection device 1. Note that, like the caption selection device 1, the re-learning device 21 may be configured with hardware such as a storage device, a processor, a user interface device, and a communication device.
[0048] [Relearning Process] The relearning process of the image caption generation model 9 performed by the relearning device 21 will be described with reference to the flowchart shown in FIG.
[0049] First, in step S21, the acquisition unit 23 acquires positive example training data from the caption selection device 1. Specifically, the caption selection device 1 provides the re-learning device 21 with frame images F and captions C that satisfy the positive example conditions, out of the multiple frame images F and the multiple corresponding captions C extracted by the extraction unit 2. That is, the caption selection device 1 provides the re-learning device 21 with the frame image F and caption C that have the highest first similarity S1 and whose second similarity S2 is equal to or greater than a predetermined threshold, as the positive example data set D1.
[0050] For example, as shown in Fig. 6, the caption selection device 1 may classify multiple captions C based on the content of the captions C. Then, the caption selection device 1 may select, as a positive example candidate, the caption C that appears most frequently among the multiple captions C. Then, as shown in Fig. 7, the caption selection device 1 may select, from the positive example candidates, frame images F and captions C whose corresponding similarity scores (I1·T1, I2·T2, I3·T3, ..., I30·T30) are equal to or greater than a predetermined threshold, as the positive example dataset D1.
[0051] In this case, there may be multiple data sets that satisfy the positive example condition. For example, there may be multiple captions C with the highest appearance frequency. In this case, the caption selection device 1 may provide the re-learning device 21 with one data set D1 randomly selected from the multiple data sets that satisfy the positive example condition as a positive example.
[0052] Furthermore, the caption selection device 1 may provide the re-learning device 21 with a data set D1 including a specific caption C1 as a positive example.
[0053] The acquisition unit 23 may also acquire negative example learning data from the caption selection device 1. Specifically, the caption selection device 1 provides the re-learning device 21 with frame images F and captions C other than the positive examples, out of the multiple frame images F and the multiple corresponding captions C extracted by the extraction unit 2, as a negative example dataset D2.
[0054] At this time, the caption selection device 1 may select a negative example data set D2 from among the data sets other than the positive examples in accordance with the magnitude of the first similarity S1 or the second similarity S2. For example, the caption selection device 1 may preferentially select a data set including a caption C with a high appearance frequency from among the data sets other than the positive examples.
[0055] As a result, when the acquisition unit 23 acquires the training data of positive examples and negative examples, it outputs the training data to the re-learning unit 24. Then, in step S22, the re-learning unit 24 re-learns the image caption generation model 9 using the training data.
[0056] Specifically, the re-learning unit 24 re-learns the image caption generation model 9 using positive example training data. That is, the re-learning unit 24 re-learns the image caption generation model 9 using a frame image F and a caption C that satisfy the positive example condition as a positive example data set D1. This allows the re-learning unit 24 to re-learn the image caption generation model 9 effectively.
[0057] Here, if there are multiple datasets that satisfy the positive example condition, the positive example dataset D1 is randomly selected from the multiple datasets, which allows the re-training unit 24 to more effectively re-train the image caption generation model 9.
[0058] Furthermore, the re-learning unit 24 may use the data set D2 other than the positive examples as negative examples to re-learn the image caption generation model 9. This allows the re-learning unit 24 to re-learn the image caption generation model 9 more effectively.
[0059] Here, the negative example dataset D2 is selected depending on the magnitude of the first similarity or the second similarity, which allows the retraining unit 24 to retrain the image caption generation model 9 more effectively.
[0060] At this time, the re-learning unit 24 may adjust the parameters of the image caption generation model 9 by using, for example, cross-entropy as a loss function. This allows the re-learning unit 24 to more effectively re-learn the image caption generation model 9 by using the training data of positive examples and negative examples.
[0061] After retraining the image caption generation model 9, the retraining unit 24 provides the generated retrained image caption generation model 22 to the caption selection device 1. This allows the caption selection device 1 to select a specific caption C1, similar to the first embodiment.
[0062] That is, the extraction unit 2 of the caption selection device 1 extracts a plurality of frame images F from a video. Next, the generation unit 3 inputs the plurality of frame images F into the re-trained image caption generation model 22, thereby generating a plurality of captions C corresponding to the plurality of frame images F. At this time, the re-trained image caption generation model 22 has been re-trained using training data of positive examples and negative examples, and therefore can generate captions C corresponding to the frame images F with high accuracy. Then, the selection unit 4 selects a specific caption C1 from the plurality of captions C based on the first similarity S1 and the second similarity S2.
[0063] According to this embodiment, the re-learning unit 24 re-learns the image caption generation model 9 using, as a positive example, a data set D1 of frame images and captions that satisfy the positive example conditions, that is, the first similarity S1 is highest and the second similarity S2 is equal to or greater than a predetermined value, among a plurality of frame images F and a plurality of captions C. This allows the re-learning unit 24 to re-learn the image caption generation model 9 effectively.
[0064] In the first embodiment, when a plurality of captions C are extracted at predetermined time intervals throughout the entire video, the selection unit 4 selects a specific caption C1 for each of the plurality of captions C. However, the selection unit 4 may select a specific caption C1 by limiting the selection to a plurality of captions C included in a specific position in the video.
[0065] For example, the selection unit 4 may select a specific caption C1 by limiting the captions C extracted from each section at predetermined intervals to those captions C in sections where the average value of the second similarity S2 is equal to or less than a predetermined threshold. Specifically, the selection unit 4 limits the captions C to those captions C in sections where the average value of the similarity scores (I1·T1, I2·T2, I3·T3, ..., I30·T30) shown in FIG. 7 is equal to or less than a predetermined threshold. The selection unit 4 then selects a specific caption C1 from among the captions C in the limited section based on the first similarity S1 and the second similarity S2. As a result, the selection unit 4 assigns the specific caption C1 to the frame image F in the limited section. In this way, the selection unit 4 selects the specific caption C1 only from sections in which appropriate captions C have not been generated in the video, thereby enabling the selection unit 4 to quickly select an appropriate caption C1 for that section.
[0066] Furthermore, the selection unit 4 may select a specific caption C1 by limiting the captions C extracted for each section every predetermined time to a plurality of captions C for which the variation rate of the variance of the second similarity S2 is equal to or greater than a predetermined threshold. Specifically, the selection unit 4 limits the captions C to a plurality of captions C for sections for which the variation rate of the variance of the similarity scores (I1·T1, I2·T2, I3·T3, ..., I30·T30) shown in FIG. 7 is equal to or greater than a predetermined threshold. In other words, the selection unit 4 limits the captions C to sections with a lot of noise in the video. Then, the selection unit 4 selects a specific caption C1 from the plurality of captions C for the limited section based on the first similarity S1 and the second similarity S2. In this way, the selection unit 4 selects a specific caption C1 for a section with a lot of noise, thereby enabling the selection unit 4 to quickly select a caption C1 appropriate for that section.
[0067] In the first embodiment, the extractor 2 sequentially extracts a predetermined number of frame images F. However, the present invention is not limited to this. For example, the extractor 2 may extract a different number of frame images F for each position in the video.
[0068] Although specific examples of the present disclosure have been described in detail above, these are merely examples and do not limit the scope of the claims. The technology described in the claims includes various modifications and alterations of the specific examples exemplified above.
[0069] The disclosures of the specification, drawings and abstract contained in Japanese Patent Application No. 2023-216243, filed on December 21, 2023, are incorporated herein by reference in their entirety.
[0070] The caption selection method according to the present disclosure can be used as a method for selecting captions to be added to videos.
[0071] REFERENCE SIGNS LIST 1 Caption selection device 2 Extraction unit 3 Generation unit 4 Selection unit 5 Storage device 6 Processor 7 User interface device 8 Communication device 9 Image caption generation model 21 Re-learning device 22 Re-learned image caption generation model 23 Acquisition unit 24 Re-learning unit C, C1, Ca Caption D1, D2 Data set F, Fa Frame image S1 First similarity S2 Second similarity
Claims
1. A computer-implemented caption selection method comprising: extracting a plurality of frame images from a video; inputting the plurality of frame images into an image caption generation model to generate a plurality of captions corresponding to the plurality of frame images; and selecting a particular caption from the plurality of captions based on a first similarity between the plurality of captions and a second similarity between the plurality of frame images and the plurality of captions.
2. The caption selection method according to claim 1, wherein the first similarity is calculated based on the frequency of appearance of captions indicating the same content among the plurality of captions.
3. The caption selection method according to claim 2, wherein the specific caption is selected according to the frequency of appearance.
4. A caption selection method according to claim 1, wherein the second similarity is calculated based on feature vectors of the plurality of frame images and feature vectors of the plurality of captions.
5. The caption selection method according to claim 1, wherein the plurality of frame images are extracted based on the playback time of the video.
6. A caption selection method as described in claim 1, which extracts the multiple captions from the video for each specified time period, and selects the specific caption from among the multiple captions extracted for each specified time period, limited to those captions whose average value of the second similarity is below a specified threshold.
7. A caption selection method as described in claim 1, which extracts the multiple captions from the video for each specified time period, and selects the specific caption from among the multiple captions extracted for each specified time period, limited to those captions for which the rate of variation of the variance of the second similarity is equal to or greater than a specified threshold.
8. A caption selection method as described in claim 1, in which the image caption generation model is retrained using a data set of frame images and captions among the plurality of frame images and the plurality of captions that satisfy the condition that the first similarity is the highest and the second similarity is equal to or greater than a predetermined value as a positive example.
9. The caption selection method according to claim 8, wherein, when there are multiple datasets that satisfy the condition, the positive example dataset is randomly selected from the multiple datasets.
10. A caption selection method as described in claim 8, in which the image caption generation model is retrained using a data set of frame images and captions other than the positive examples from among the plurality of frame images and the plurality of captions as negative examples.
11. The method of claim 10, wherein the negative example dataset is selected depending on the magnitude of the first similarity or the second similarity.
12. A caption selection program that causes a computer to execute the following steps: extracting a plurality of frame images from a video; inputting the plurality of frame images into an image caption generation model to generate a plurality of captions corresponding to the plurality of frame images; and selecting a specific caption from the plurality of captions based on a first similarity between the plurality of captions and a second similarity between the plurality of frame images and the plurality of captions.
13. A caption selection device comprising: an extraction unit that extracts a plurality of frame images from a video; a generation unit that inputs the plurality of frame images into an image caption generation model to generate a plurality of captions corresponding to the plurality of frame images; and a selection unit that selects a specific caption from the plurality of captions based on a first similarity between the plurality of captions and a second similarity between the plurality of frame images and the plurality of captions.
Citation Information
Patent Citations
Discovery of semantic similarities between images and text
US20170061250A1
Automated video and audio annotation techniques
WO2023149898A1