Response generation device, response generation method, and response generation program

The response generation device addresses the challenge of extracting refined video features from videos of arbitrary length by using a frame complementation unit and moving image feature extraction unit, ensuring consistent feature extraction and improved task performance.

JP7694794B2Active Publication Date: 2025-06-18NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024500885
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-18
Publication Date
2025-06-18
Estimated Expiration
2042-02-18

AI Technical Summary

Technical Problem

Existing video feature extraction models struggle to obtain refined feature quantities from videos of arbitrary length, as they assume a fixed number of frames, leading to inconsistent time information density and suboptimal performance in tasks requiring precise video content capture.

Method used

A response generation device that includes a frame complementation unit to ensure a fixed number of frames is maintained, and a moving image feature extraction unit that extracts features from complemented frames, allowing for refined feature extraction from videos of varying lengths.

Benefits of technology

Enables the extraction of refined video feature quantities from videos of arbitrary length, improving the performance of subsequent tasks by maintaining consistent feature extraction across varying video lengths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694794000001
    Figure 0007694794000001
  • Figure 0007694794000002
    Figure 0007694794000002
  • Figure 0007694794000003
    Figure 0007694794000003
Patent Text Reader

Abstract

If the number of frames sampled at equal intervals from divided video images obtained by dividing a video image to be processed at equal intervals from the beginning does not satisfy a prescribed condition, a response generation device (20) supplements the frames sampled from the divided video images. The response generation device (20) then extracts a feature quantity for each section by inputting a plurality of frames that are contained in the section and consist of frames sampled from the divided video images and supplementary frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a response generation device, a response generation method, and a response generation program.

Background Art

[0002] In recent years, a multimodal question-and-answer technology that outputs an appropriate answer using moving image data and a question about the moving image as inputs has been proposed.

[0003] For example, in Non-Patent Document 1, in order to perform multimodal question-and-answer with multiple turns, a method is disclosed in which moving image data, a dialogue history, and the text of the question content are used as inputs, and an answer sentence is output for each token based on a neural autoregressive model.

[0004] For example, in Non-Patent Document 2, a method is disclosed in which moving images are used as inputs, and moving image feature amounts are output in an intermediate layer of a network based on machine learning such as a Transformer.

Prior Art Documents

Non-Patent Documents

[0005]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, in the prior art, it has not been possible to obtain refined video feature quantities from videos having an arbitrary time length using a video feature extraction model that assumes input and output of a fixed number of frames.

[0007] The pre-trained model of Non-Patent Document 2 extracts features from video frames for a fixed number of frames regardless of the length of the video. Therefore, feature quantities of a fixed number of frames are obtained for videos of any time length, and the density of the time information contained in the feature quantities changes according to the samples, which may have an adverse effect on the subsequent tasks. For this reason, for tasks that require precise capture of video content, a method that enables feature extraction of an arbitrary number of frames even from a video feature extraction model with fixed extracted time frames is desirable.

[0008] The present invention has been made in view of the above, and an object thereof is to obtain refined video feature quantities from videos having an arbitrary time length using a video feature extraction model that assumes input and output of a fixed number of frames.

Means for Solving the Problems

[0009] In order to solve the above-described problems and achieve the object, the response generation device of the present invention has a frame complementation unit that complements frames to be sampled from the divided moving image when the number of frames obtained by equally sampling the divided moving images obtained by equally dividing the moving image to be processed at equal intervals from the beginning does not satisfy a predetermined condition, and a moving image feature extraction unit that extracts feature amounts for each interval by inputting a plurality of frames included in one interval of the frames sampled from the divided moving image and the frames complemented by the frame complementation unit.

Effect of the Invention

[0010] According to the present invention, it becomes possible to obtain refined moving image feature amounts from a moving image having an arbitrary time length using a moving image feature extraction model assuming input / output of a fixed number of frames.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Embodiments for Carrying Out the Invention

[0012] Hereinafter, embodiments of a response generation device, a response generation method, and a response generation program according to the present application will be described in detail with reference to the drawings. Note that the present invention is not limited to the embodiments described below.

[0013] [Configuration of Feature Extraction Device] FIG. 1 is a block diagram illustrating the configuration of the feature extraction device of the present embodiment. As illustrated in FIG. 1, the feature extraction device 10 of the present embodiment includes a communication processing unit 11, an input unit 12, an output unit 13, a control unit 14, and a storage unit 15. The feature extraction device 10 and the response generation device 20 are communicably connected by wire or wirelessly via a predetermined communication network (network N).

[0014] The feature extraction device 10 is an information processing device for obtaining a refined moving image feature amount from a moving image having an arbitrary time length. For example, the feature extraction device 10 is realized by a server device, a cloud system, or the like. For example, the feature extraction device 10 extracts a moving image feature amount and transmits it to the response generation device 20.

[0015] The response generation device 20 is an information processing device aimed at generating multimodal question responses from moving image features. For example, the feature extraction device 10 is realized by a server device, a cloud system, or the like. For example, the response generation device 20 generates question responses using moving image features. Note that the question responses are based on, for example, the QA method.

[0016] The communication processing unit 11 is realized by a NIC (Network Interface Card) or the like, and controls the communication between the response generation device 20 (see FIG. 7), which generates multimodal question responses from moving image features, and the control unit 14 via a telecommunications line such as a LAN (Local Area Network) or the Internet. For example, the communication processing unit 11 transmits the moving image features output by the control unit 14 to the response generation device 20.

[0017] The input unit 12 is realized using an input device such as a keyboard or a mouse, and inputs various instruction information such as a processing start to the control unit 14 in response to an input operation by an operator. The output unit 13 is realized by a display device such as a liquid crystal display. For example, the output unit 13 outputs the result of the extraction process of the moving image features.

[0018] The storage unit 15 stores data and programs necessary for various processes by the control unit 14, and has a moving image storage unit 15a and a feature extraction model storage unit 15b. For example, the storage unit 15 is a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk.

[0019] The moving image storage unit 15a stores moving images that are the targets of the extraction process of moving image features. In the example shown in FIG. 2, an example is shown in which conceptual information such as "moving image #1" and "moving image #2" is stored in the "moving image", but actually, moving image data and the like are stored. Also, in the "moving image", for example, a URL where the moving image data is located, a file path name indicating the storage location, or the like may be stored. FIG. 2 is a diagram showing an example of the moving image storage unit 15a.

[0020] The feature extraction model storage unit 15b stores a pre-trained moving image feature extraction model for extracting moving image feature amounts. Note that the moving image feature extraction model is a model obtained by, for example, pre-training the parameters of the neural network that constitutes the moving image feature extraction unit 14d described later using a large-scale dataset. In the example shown in FIG. 3, an example is shown in which conceptual information such as "feature extraction model #1" and "feature extraction model #2" is stored in the "moving image feature extraction model", but actually, the internal parameters of the moving image feature extraction model and the like are stored. FIG. 3 is a diagram showing an example of the feature extraction model storage unit 15b.

[0021] The control unit 14 has an internal memory for storing programs that define various processing procedures and the like and required data, and executes various processes based on these. For example, the control unit 14 includes an interval division unit 14a, a moving image sampling unit 14b, a frame complementation unit 14c, a moving image feature extraction unit 14d, a complement frame deletion unit 14e, and an interval combination unit 14f. Here, the control unit 14 is an electronic circuit such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an MPU (Micro Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0022] Here, prior to the description of the interval division unit 14a, the moving image sampling unit 14b, the frame complementation unit 14c, the moving image feature extraction unit 14d, the complement frame deletion unit 14e, and the interval combination unit 14f included in the control unit 14, an overview of the processing executed by the feature extraction device 10 together with the response generation device 20 by the processing executed by the control unit 14 will be described.

[0023] For example, when moving image data and a question about the moving image are input, a process of outputting a response according to the content of the moving image is performed. In an AVSD (Audio Visual Scene-Aware Dialog) that has the task of correctly answering questions about moving images, global feature extraction in the spatio-temporal direction is required. For example, when the number of people shown in a moving image is questioned, it is necessary to identify people from each frame of the moving image and track / identify each person throughout the entire moving image.

[0024] However, in conventional multimodal question-and-answer technologies, a moving image feature extraction model based on a convolutional neural network (CNN) is often used. In such a CNN, local feature extraction tends to be performed in the spatio-temporal direction, and it may be difficult to extract global features. As a result, it is not possible to accurately identify people from each frame of the moving image, track / identify each person throughout the entire moving image, and there is a risk that the number of people shown in the moving image cannot be appropriately answered.

[0025] Therefore, a technique of generating a response to a question about a moving image using the moving image feature amount obtained by a moving image feature extraction model based on a technique called a transformer disclosed in Non-Patent Document 2 can be considered. Such a moving image feature extraction model based on a transformer can extract more globally spatio-temporal features compared to a moving image feature extraction model based on a CNN, and thus is suitable for an action recognition task.

[0026] Therefore, it is considered that by using the feature amount of a moving image feature extraction model based on a transformer in the model of AVSD, better moving image understanding and response can be achieved. The moving image feature extraction model of this embodiment is based on a transformer. That is, in this embodiment, response generation is performed using the feature amount of a moving image feature extraction model based on a transformer.

[0027] Hereinafter, various processing procedures of the control unit 14 will be described with reference to FIG. 4. FIG. 4 shows a specific example of the processing of moving image feature extraction in the present invention. For comparison, FIG. 5 shows a specific example of the processing of conventional moving image feature extraction. Note that G1 and G2 included in the moving image D1 and the moving image D2 are images of predetermined frames in the moving image, and V1 and V2 are feature amounts (feature vectors) of predetermined frames of the moving image feature amounts. In FIGS. 4 and 5, for convenience of explanation, symbols are attached only to predetermined frames such as G1, G2, V1, and V2, but symbols may be attached to images and feature vectors of all frames in the moving image.

[0028] When receiving a moving image, the section division unit 14a outputs, in order from the head, a moving image of a certain section having a time length (assumed time length) assumed as an input by the pre-trained moving image feature extraction model as a moving image series (moving image series) (step S101). Note that the last element of the moving image series may be less than the assumed time length. In FIG. 4, the section division unit 14a outputs DK1 to DK4 as a moving image series from the moving image D1.

[0029] The moving image sampling unit 14b outputs a moving image series (hereinafter, appropriately referred to as "sampled moving image series") obtained by sampling the moving image series output in step S101 at equal intervals for each section (step S102). Note that the number of frames for each series is the number of frames (assumed number of frames) assumed as an input by the pre-trained moving image feature extraction model. Also, the last element of the sampled moving image series may be less than the assumed number of frames. In FIG. 4, the moving image sampling unit 14b outputs HDK1 to HDK4 as a sampled moving image series from DK1 to DK4.

[0030] Here, there may be a case where a specific element less than the assumed number of frames is included in the moving image. And when it is less than the assumed number of frames, the moving image feature extraction model may not be able to handle it, and there may be a case where the feature amount cannot be appropriately extracted. Therefore, in the present embodiment, the number of frames is complemented so that the number of frames of the specific element matches the assumed number of frames.

[0031] The frame completion unit 14c outputs a moving image series (hereinafter, appropriately referred to as the "post-frame-completed moving image series") obtained by complementing the image frames of the post-sampling moving image series output in step S102 so that the number of frames of the last element of the post-sampling moving image series matches the assumed number of frames (step S103). In FIG. 4, the frame completion unit 14c outputs FDK1 to FDK4 as the post-frame-completed moving image series from HDK1 to HDK4. Note that the completion is performed by duplicating the image of the last frame among the last elements of the post-sampling moving image series and adding it to the end. In FIG. 4, the frame completion unit 14c outputs FDK4 by duplicating the image of the last frame of HDK4 and adding it to the end of HDK4. Note that this addition method is an example, and for example, a predictor may be separately used to predict future frames and add them to the end. Note that G11 and G12 are images of the frames of the complemented part of the moving image.

[0032] The moving image feature extraction unit 14d inputs the post-frame-completed moving image series output in step S103 to the pre-trained moving image feature extraction model, and outputs the corresponding moving image feature amount series (hereinafter, appropriately referred to as the "post-frame-completed moving image feature amount series") (step S104). In FIG. 4, the moving image feature extraction unit 14d outputs FDTK1 to FDTK4 as the post-frame-completed moving image feature amount series from FDK1 to FDK4. Note that the extraction is based on the internal parameters of the pre-trained moving image feature extraction model for each element of the post-frame-completed moving image series.

[0033] The complementary frame deletion unit 14e outputs a moving image feature amount series excluding the frame complementary part by excluding the feature amounts corresponding to the image frames complemented in step S103 from the moving image feature amount series after frame complementation output in step S104 (step S105). In FIG. 4, the complementary frame deletion unit 14e outputs DTK1 to DTK4 as a moving image feature amount series from FDTK1 to FDTK4. Specifically, the complementary frame deletion unit 14e outputs DTK1 to DTK4 by excluding the feature amounts (V11 and V12) corresponding to the replication of the last frame of HDK4 from FDTK4. Note that V11 and V12 are the feature vectors of the frames of the complementary part of the moving image feature amount.

[0034] The interval combination unit 14f outputs one moving image feature amount by combining the feature amounts of the moving image feature amount series output in step S104 in the time direction over the entire interval (step S106). In FIG. 4, the interval combination unit 14f outputs DTK11 as one moving image feature amount from DTK1 to DTK4. That is, DTK11 is the moving image feature amount corresponding to the moving image D1. Then, the interval combination unit 14f transmits the output one moving image feature amount to the response generation device 20.

[0035] On the other hand, in FIG. 5, by sampling the image frames of the moving image D2 at equal intervals, HD21 is output as the sampled moving image (sampled moving image). Further, by inputting the sampled moving image into the pre-trained moving image feature extraction model, the corresponding moving image feature amount is output. Specifically, DTK21 is output as the moving image feature amount from HD21. That is, DTK21 is the moving image feature amount corresponding to the moving image D2.

[0036] [Processing Procedure of Feature Extraction Device] Next, an example of the processing procedure of the processing executed by the feature extraction device 10 will be described with reference to FIG. 6. FIG. 6 is a flowchart showing an example of the processing procedure of the feature extraction process.

[0037] As illustrated in FIG. 6, the feature extraction device 10 takes a moving image as input, performs interval division processing, and outputs a moving image sequence (step S201).

[0038] Then, the feature extraction device 10 takes the moving image sequence as input, performs moving image sampling processing, and outputs the sampled moving image sequence (step S202).

[0039] Then, the feature extraction device 10 takes the sampled moving image sequence as input and determines whether there are specific elements less than the assumed number of frames (step S203).

[0040] Also, when the feature extraction device 10 determines that there are specific elements less than the assumed number of frames (step S203; YES), it takes the sampled moving image sequence as input, performs frame completion processing, and outputs the moving image sequence after frame completion (step S204).

[0041] Then, the feature extraction device 10 takes the moving image sequence after frame completion as input, performs moving image feature extraction processing, and outputs the moving image feature amount sequence after frame completion (step S205).

[0042] Then, the feature extraction device 10 takes the moving image feature amount sequence after frame completion as input, performs complementary frame deletion processing, and outputs the moving image feature amount sequence (step S206).

[0043] Then, the feature extraction device 10 takes the moving image feature amount sequence as input, performs interval combination processing, and outputs the moving image feature amount (step S207).

[0044] On the other hand, when the feature extraction device 10 determines that there are no specific elements less than the assumed number of frames (step S203; NO), it omits step S204. In step S205, it takes the sampled moving image sequence as input, performs moving image feature extraction processing, and outputs the moving image feature amount sequence. Then, it omits step S206 and performs the processing of step S207.

[0045] [Configuration of Response Generation Device] FIG. 7 is a block diagram illustrating the configuration of the response generation device according to the present embodiment. As illustrated in FIG. 7, the response generation device 20 according to the present embodiment includes a communication processing unit 21, an input unit 22, an output unit 23, a control unit 24, and a storage unit 25.

[0046] The communication processing unit 21 is implemented by a NIC or the like, and controls communication between the feature extraction device 10 and the control unit 24 via a telecommunication line such as a LAN or the Internet. For example, the communication processing unit 21 receives moving image feature amounts from the feature extraction device 10.

[0047] The input unit 22 is implemented using an input device such as a keyboard or a mouse, and inputs various instruction information such as a processing start to the control unit 24 in response to an input operation by an operator. The output unit 23 is implemented by a display device such as a liquid crystal display. For example, the output unit 23 outputs the result of the generation process of the multimodal question response.

[0048] The storage unit 25 stores data and programs necessary for various processes by the control unit 24, and includes a dialogue history / question text storage unit 25a and a multimodal question response model storage unit 25b. For example, the storage unit 25 is a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk.

[0049] The dialogue history / question text storage unit 25a stores the dialogue history and questions that are the targets of the generation process of the multimodal question response. In the example shown in FIG. 8, an example is shown in which conceptual information such as "dialogue history / question text #1" and "dialogue history / question text #2" is stored in the "dialogue history / question text", but actually, text data indicating the dialogue history and questions is stored. FIG. 8 is a diagram showing an example of the dialogue history / question text storage unit 25a.

[0050] The multimodal question-and-answer model storage unit 25b stores a pre-trained multimodal question-and-answer model for multimodal question-and-answer. In the example shown in FIG. 9, conceptual information such as "multimodal question-and-answer model #1" and "multimodal question-and-answer model #2" is stored in the "multimodal question-and-answer model", but actually, internal parameters of the multimodal question-and-answer model and the like are stored. FIG. 9 is a diagram showing an example of the multimodal question-and-answer model storage unit 25b.

[0051] The control unit 24 has an internal memory for storing a program that defines various processing procedures and the like and required data, and executes various processes based on these. For example, the control unit 24 includes a response generation unit 24a. Here, the control unit 24 is an electronic circuit such as a CPU or MPU, or an integrated circuit such as an ASIC or FPGA.

[0052] Here, in AVSD, it is required to correctly answer questions regarding moving image data, acoustic data, and the like. For example, in the example shown in FIG. 10, it is required to derive an answer such as A10 from the moving image data VD1, acoustic data SD1, conversation history (Q1, A1, Q2, A2, ···, Q9, A9), and question Q10.

[0053] The response generation unit 24a inputs the moving image feature amount, conversation history, and question into the multimodal question-and-answer model, and outputs a response to the question. Note that the response generation unit 24a may include the acoustic feature amount as an input. In the learning of the multimodal question-and-answer model, for example, the moving image feature amount, conversation history, question, and response to the question are used as a data set. At this time, the moving image feature amount uses the moving image feature amount received from the feature extraction device 10. Further, when it is assumed that the acoustic feature amount is included in the input of the response generation unit 24a, the acoustic feature amount may be further included in the data set.

[0054] Note that the feature extraction device 10 and the response generation device 20 may be integrated. For example, the feature extraction device 10 may be included in the configuration of the response generation device 20. For example, the response generation device 20 may have, in the control unit 24, an interval division unit 14a, a moving image sampling unit 14b, a frame complementation unit 14c, a moving image feature extraction unit 14d, a complemented frame deletion unit 14e, and an interval combination unit 14f, and the response generation device 20 may execute each process of the feature extraction device 10.

[0055] [Processing Procedure of Response Generation Device] Next, with reference to FIG. 11, an example of the processing procedure of the process executed by the response generation device 20 will be described. FIG. 11 is a flowchart showing an example of the processing procedure of the response generation process.

[0056] As illustrated in FIG. 11, the response generation device 20 acquires the moving image feature amount output by the feature extraction process executed by the feature extraction device 10 (step S301).

[0057] Then, the response generation device 20 inputs the moving image feature amount, the dialogue history, and the question into the multimodal question answering model, performs a response generation process, and outputs a response (step S302).

[0058] [Effects of Embodiment] As described above, when the number of frames obtained by sampling at equal intervals the divided moving images obtained by dividing the moving image to be processed at equal intervals from the beginning does not satisfy a predetermined condition, the response generation device 20 according to the embodiment complements the frames sampled from the divided moving images, and extracts feature amounts for each interval by inputting the frames sampled from the divided moving images and the complemented frames, which are a plurality of frames included in one interval.

[0059] According to the present invention, since feature extraction can be performed with the number of frames corresponding to the time length of the moving image, and a neural network for solving a desired task can be trained using the obtained moving image feature amount, it is possible to improve the performance of the task.

[0060] In a moving image feature extraction model that assumes input and output of a fixed number of frames in the past, it was not possible to obtain refined moving image feature amounts from a moving image having an arbitrary time length. According to the present invention, by dividing a moving image into intervals at regular time intervals, performing moving image feature extraction for each interval, and combining the obtained feature amounts in the time direction, it becomes possible to extract moving image feature amounts from a moving image having an arbitrary number of frames.

[0061] Further, according to the present invention, by using the obtained moving image feature amounts as input feature amounts of a model in a multimodal question-and-answer task, it becomes possible to improve the degree of coincidence between the generated sentence and the correct sentence of the model.

[0062] FIG. 12 is a diagram showing the results of a verification experiment for a question-and-answer task regarding a moving image capturing a daily scene of a person. In the verification experiment of the results shown in FIG. 12, the structure and learning method of the neural network were those shown in Non-Patent Document 1, and the moving image feature amounts were the feature amounts obtained by the method of Non-Patent Document 2. Further, as a learning / evaluation dataset, the moving images shown in Non-Patent Document 3 and the dataset shown in Non-Patent Document 4 as question-and-answer pairs for each moving image were used. From the dataset, 7,659 moving image / question-and-answer pairs were used for learning, and 1,710 moving image / question-and-answer pairs were used for evaluation.

[0063]

Non-Patent Document 3

Non-Patent Document 4

[0064] A method for extracting feature quantities for a fixed number of frames from the moving images shown in FIG. 5 and a multi-modal question-and-answer model were learned using the moving image feature quantities extracted by the method according to the present invention shown in FIG. 4, and the degree of coincidence between the generated sentence and the correct sentence for the evaluation data was evaluated. Also, BLEU-1 was used as the evaluation index for the degree of coincidence. When the feature quantities extracted by the method for extracting feature quantities for a fixed number of frames as shown in FIG. 5 were used (corresponding to the first case), BLEU-1 was 0.692, and when the feature quantities extracted by the method according to the present invention as shown in FIG. 4 were used (corresponding to the second case), BLEU-1 was 0.695. From this result, it was shown that the degree of coincidence with the correct sentence was improved by using the moving image feature quantities according to the present invention.

[0065] 〔System configuration, etc.〕 Each component of each of the illustrated devices according to the above embodiment is a functional concept, and does not necessarily have to be physically configured as shown in the drawings. That is, the specific form of the distribution and integration of each device is not limited to that shown in the drawings, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads, usage situations, etc. Furthermore, each processing function performed by each device can be realized in whole or in any part by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware by wired logic.

[0066] Also, among the processes described in the above embodiments, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, regarding the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings, they can be arbitrarily changed unless otherwise specified.

[0067] 〔Program〕 Also, it is possible to create a program that describes the processes executed by the feature extraction device 10 and the response generation device 20 described in the above embodiments in a language executable by a computer. In this case, by having the computer execute the program, the same effects as in the above embodiments can be obtained. Furthermore, such a program may be recorded on a computer-readable recording medium, and the same processes as in the above embodiments may be realized by having the computer read and execute the program recorded on this recording medium.

[0068] FIG. 13 is a diagram showing a computer that executes a program. As illustrated in FIG. 13, the computer 1000 has, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070, and these components are connected by a bus 1080.

[0069] As illustrated in FIG. 13, the memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System), for example. The hard disk drive interface 1030 is connected to a hard disk drive 1090 as illustrated in FIG. 13. The disk drive interface 1040 is connected to a disk drive 1100 as illustrated in FIG. 13. For example, a removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120 as illustrated in FIG. 13. The video adapter 1060 is connected to, for example, a display 1130 as illustrated in FIG. 13.

[0070] Here, as illustrated in FIG. 13, the hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the above programs are stored, for example, in the hard disk drive 1090 as program modules in which instructions to be executed by the computer 1000 are described.

[0071] Also, the various data described in the above embodiments are stored, for example, in the memory 1010 or the hard disk drive 1090 as program data. Then, the CPU 1020 reads out the program module 1093 and the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as needed and executes various processing procedures.

[0072] Note that the program module 1093 and program data 1094 related to the program are not limited to being stored in the hard disk drive 1090. For example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive or the like. Alternatively, the program module 1093 and program data 1094 related to the program may be stored in another computer connected via a network (LAN, WAN (Wide Area Network), etc.) and read by the CPU 1020 via the network interface 1070.

[0073] As described above, the embodiments to which the invention made by the present inventor is applied have been described. However, the present invention is not limited by the description and drawings that form a part of the disclosure of the present invention according to this embodiment. That is, all other embodiments, examples, operation techniques, etc. made by those skilled in the art based on this embodiment are included in the scope of the present invention.

[0074] Regarding the above embodiments, the following additional remarks are disclosed.

[0075] (Supplementary Note 1) A memory, At least one processor connected to the memory, Including, The processor, When the number of frames obtained by equally sampling the frames of the divided moving image obtained by equally dividing the moving image to be processed at equal intervals from the beginning does not satisfy a predetermined condition, the frames to be sampled from the divided moving image are complemented, By inputting a plurality of frames included in one section of the frames sampled from the divided moving image and the complemented frames, feature amounts are extracted for each section Response generation device.

[0076] (Supplementary Note 2) The processor, Using a pre-trained moving image feature extraction model, which is a transformer-based moving image feature extraction model trained to extract a predetermined number of feature amounts from a predetermined number of frames, to extract the feature amounts The response generation device according to the above-mentioned appended claim 1

[0077] (Appended claim 3) The processor Taking the extracted feature amounts, the dialogue history, and the question as inputs, and outputting a response to the question based on a multimodal question-answering model The response generation device according to the above-mentioned appended claim 2

[0078] (Appended claim 4) The processor Deleting the feature amounts for the complemented frames from the extracted feature amounts The response generation device according to the above-mentioned appended claim 1

[0079] (Appended claim 5) A non-transitory storage medium storing a program executable by a computer to execute a response generation process, The response generation process When the number of frames obtained by equally sampling the divided moving images obtained by equally dividing the moving image to be processed at equal intervals from the beginning does not satisfy a predetermined condition, complementing the frames to be sampled from the divided moving image, Extracting feature amounts for each interval by inputting a plurality of frames included in one interval of the frames sampled from the divided moving image and the complemented frames Non-transitory storage medium

Description of symbols

[0080] 10 Feature extraction device 11 Communication processing unit 12 Input unit 13 Output unit 14 Control unit 14a Interval division unit 14b Moving image sampling unit 14c Frame Completion Unit 14d Moving Image Feature Extraction Unit 14e Complemented Frame Deletion Unit 14f Interval Combining Unit 15 Memory Unit 15a Moving Image Memory Unit 15b Feature Extraction Model Memory Unit 20 Response Generation Device 21 Communication Processing Unit 22 Input Unit 23 Output Unit 24 Control Unit 24a Response Generation Unit 25 Memory Unit 25a Dialogue History / Question Text Memory Unit 25b Multimodal Question-Response Model Memory Unit N Network

Claims

1. a frame complementing unit that complements frames sampled from a divided video obtained by dividing a video to be processed at equal intervals from a beginning of the video when the number of frames sampled at equal intervals from the divided video does not satisfy a predetermined condition; a video feature extraction unit that extracts a feature for each section by inputting a plurality of frames included in one section, the frames sampled from the divided video images and the frames complemented by the frame complement unit; A response generating device comprising:

2. The video feature extraction unit The response generation device described in claim 1, characterized in that the features are extracted using a pre-trained video feature extraction model, which is a Transformer-based video feature extraction model trained to extract a predetermined number of features from a predetermined number of frames.

3. The video feature extraction unit includes:

3. The response generation device according to claim 2, wherein the extracted features, a dialogue history, and a question are input, and a response to the question is output based on a multimodal question answering model.

4. The video feature extraction unit 4. The response generating device according to claim 1, further comprising: a feature amount for the interpolated frame being deleted from the extracted feature amounts.

5. A response generation method executed by a response generation device, comprising: a frame interpolation step of interpolating frames sampled from a divided video sequence obtained by dividing the video sequence to be processed at equal intervals from the beginning of the video sequence when the number of frames sampled at equal intervals from the divided video sequence does not satisfy a predetermined condition; a video feature extraction step of extracting a feature for each section by inputting a plurality of frames included in one section, the frames sampled from the divided video images and the frames interpolated by the frame interpolation step; A response generation method characterized by including

6. When the number of frames obtained by equally sampling the divided moving images obtained by dividing the moving image to be processed at equal intervals from the beginning does not satisfy a predetermined condition, a frame complementation step of complementing the frames to be sampled from the divided moving images, and A moving image feature extraction step of extracting feature amounts for each interval by inputting a plurality of frames included in one interval of the frames sampled from the divided moving images and the frames complemented by the frame complementation step, and A response generation program characterized by causing a computer to execute

Citation Information

Patent Citations

  • System and method for attention-based configurable convolutional neural network (abc-CNN) for visual question answering

    JP2017091525A

  • Method, apparatus, electronic device, computer-readable storage medium, and computer program for image-based data processing

    JP2020135852A