Information-of-interest detecting device, information-of-interest detecting method, and program
By generating and comparing object feature sequences with text features across multiple images, the interest information detection device improves the accuracy of identifying relevant video segments and objects, addressing inefficiencies in existing video-text information extraction methods.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2025-11-06
- Publication Date
- 2026-05-15
AI Technical Summary
Existing techniques for obtaining information from video data using text data, such as those described in Patent Document 1 and Shen Yan et al.'s UnLoc framework, individually compare feature amounts of each image with text features, leading to inefficiencies and reduced accuracy in detecting relevant information.
The interest information detection device and method utilize an acquisition unit to acquire query text and video data, calculate text features, generate object feature sequences, and detect interest information by comparing these sequences, thereby considering the features of objects across multiple images and changes in object states.
This approach enhances the accuracy of detecting interest information by leveraging time-series object features, allowing for precise identification of matching intervals and objects within video data based on text queries.
Smart Images

Figure JP2025038951_15052026_PF_FP_ABST
Abstract
Description
Interest Information Detection Device, Interest Information Detection Method, and Program
[0001] The present disclosure relates to an interest information detection device, an interest information detection method, and a program.
[0002] Techniques for obtaining information from video data using text data have been developed. For example, Patent Document 1 discloses a technique that receives input of text and video, and detects an image that matches the text from each individual image constituting the input video.
[0003] Japanese Patent Application Laid-Open No. 2024-126309
[0004] Shen Yan, et al. 7, "UnLoc: A Unified Framework for Video Localization Tasks", [online], arXiv, August 21, 2023, [retrieval date April 10, 2024], Internet <URL:https: / / arxiv.org / pdf / 2308.11062.pdf>
[0005] In Patent Document 1, the feature amount of each image constituting the video is individually compared with the feature amount of the text. The present disclosure has been made in view of this problem, and one of its purposes is to provide a new technique for obtaining information from video data using text data.
[0006] The interest information detection device according to the present disclosure includes an acquisition unit that acquires query text data and video data, a calculation unit that calculates a text feature amount from the query text data, a first generation unit that generates, for each object, an object feature amount sequence that is a sequence of object feature amounts from the video data, and a detection unit that detects interest information from the video data using the text feature amount and the object feature amount sequences of each of the objects. The interest information includes at least an interest section that is a section that matches the content of the query text data, or an interest object that is an object that matches the content of the query text data.
[0007] The interest information detection method relating to this disclosure is performed by a computer. The interest information detection method includes an acquisition step of acquiring query text data and video data; a calculation step of calculating text features from the query text data; a first generation step of generating an object feature sequence, which is a sequence of object features, for each object from the video data; and a detection step of detecting interest information from the video data using the text features and the object feature sequence for each object. The interest information includes at least an interest interval, which is an interval that matches the content of the query text data, or an interest object, which is an object that matches the content of the query text data.
[0008] The program relating to this disclosure causes a computer to perform the following steps: an acquisition step of acquiring query text data and video data; a calculation step of calculating text features from the query text data; a first generation step of generating an object feature sequence, which is a sequence of object features, for each object from the video data; and a detection step of detecting interest information from the video data using the text features and the object feature sequence for each object. The interest information includes at least an interest interval, which is an interval that matches the content of the query text data, or an interest object, which is an object that matches the content of the query text data.
[0009] This disclosure provides a new technology for obtaining information from video data using text data.
[0010] This is a diagram illustrating the overview of the operation of the interest information detection device. This is a block diagram illustrating the functional configuration of the interest information detection device. This is a block diagram illustrating the hardware configuration of the computer that implements the interest information detection device. This is a flowchart illustrating the flow of processing performed by the interest information detection device. This is a diagram illustrating the configuration of object tracking information. This is a diagram illustrating the first example of the configuration of the first linked data. This is a diagram illustrating the second example of the configuration of the first linked data. This is a diagram illustrating the flow of generating the first related feature sequence using the first related feature sequence generation model. This is a diagram illustrating the training method of the first related feature sequence generation model and the interest interval detection model. This is a diagram illustrating the training method of the first related feature sequence generation model and the interest object determination model. This is a diagram illustrating the flow of interest information detection. This is a diagram illustrating the training method of the interest interval detection model. This is a diagram illustrating the training method of the interest object determination model. This is a diagram illustrating the overview of the operation of the interest information detection device. This is a diagram illustrating the configuration of the interest information detection device. This is a flowchart illustrating the flow of processing performed by the interest information detection device. This is a diagram illustrating the method of generating the third related feature sequence. This is a diagram illustrating the training method of the interest interval detection model. This diagram illustrates a training method for an interest object determination model. This diagram illustrates a training method for an association score calculation model. This diagram illustrates a method for detecting interest information.
[0011] Embodiments of the present disclosure will be described in detail below with reference to the drawings. In each drawing, the same or corresponding elements are denoted by the same reference numerals, and redundant explanations are omitted as necessary for clarity. Unless otherwise specified, predetermined values such as specified values and thresholds are stored in advance in a storage device accessible from the device that uses those values. Furthermore, unless otherwise specified, the storage unit is composed of one or any number of storage devices.
[0012] [Embodiment 1] <Overview> Figure 1 is a diagram illustrating the overview of the operation of the interest information detection device 2000. Here, Figure 1 is a diagram intended to facilitate understanding of the overview of the interest information detection device 2000, and the operation of the interest information detection device 2000 is not limited to the operation shown in Figure 1.
[0013] The interest information detection device 2000 detects information that matches the query from the video data 20. The search query represents the conditions (i.e., search conditions) for the information to be obtained from the video data 20. The information detected from the video data 20 is, for example, an interval that matches the query (hereinafter referred to as the interest interval 52) or an object that matches the query (hereinafter referred to as the interest object 54).
[0014] The interest interval 52 is a section of the video data 20 that represents a scene that matches the conditions specified in the query. The interest object 54 is an object captured in the video data 20 that matches the conditions specified in the query. Hereafter, the information detected by the interest information detection device 2000, such as the interest interval 52 and the interest object 54, will be collectively referred to as interest information 50.
[0015] The interest information detection device 2000 acquires query text data 10, which is text data representing the content of the query. The query text data 10 represents conditions related to objects. More specifically, the query text data 10 represents conditions related to relationships between objects. Hereinafter, "relationships between objects" refers to any relationship between a pair of objects, namely subject and object. Hereafter, relationships between objects will also be referred to as "inter-object relationships."
[0016] The object relationship represented by the query text data 10 is, for example, "a person is sitting on a chair." In this object relationship, the object pair is a subject, "a person," and an object, "a chair." The relationship between the subject and the object is "sitting." The interest information detection device 2000, having acquired this query text data 10, detects an interest interval from the video data 20 that represents the scene "a person is sitting on a chair," and detects the person and chair in that scene as objects of interest.
[0017] Object-object relationships can also be represented by a triplet of the form (subject, object, relationship). Here, a triplet refers to a combination of three elements arranged in a predetermined order (in other words, a tuple with three elements). For example, the object-object relationship "a person is sitting in a chair" can be represented by the triplet (person, chair, sitting).
[0018] Note that the conditions expressed by the query text data 10 are not limited to conditions relating to relationships between objects. For example, the query text data 10 may indicate a condition relating to a single object. A condition relating to a single object might be, for example, "a person is sitting."
[0019] For example, the interest information detection device 2000 operates as follows: The interest information detection device 2000 acquires query text data 10 and video data 20. The interest information detection device 2000 calculates text features 30 from the query text data 10. Text features 30 are text features calculated from the query text data 10.
[0020] The interest information detection device 2000 generates an object feature sequence 40 for each object included in the video data 20. Here, the objects included in the video data 20 refer to the objects captured in the video data 20. Specifically, the interest information detection device 2000 calculates object feature quantities 42 for each object included in each of the multiple video frames 22 included in the video data 20. Using the multiple object feature quantities 42 calculated for each object, the interest information detection device 2000 generates an object feature sequence 40 for each object, which is time-series data of the object feature quantities 42.
[0021] The object feature quantities 42 represent features of the object captured in the image. The object features obtained from the video frame 22 include at least the external appearance of the object. In addition, the object features obtained from the video frame 22 may also include the object's position, type of object, and so on.
[0022] The interest information detection device 2000 detects interest information 50 from the video data 20 using text features 30 and object feature sequences 40 for each object.
[0023] <Example of effect> According to the interest information detection device 2000, interest information 50 is detected from the video data 20 using text features 30 calculated from the query text data 10 and an object feature sequence 40, which is a sequence of feature quantities for each object detected from the video data 20. In this way, the interest information detection device 2000 provides a new technology for obtaining information from video data using text data.
[0024] Furthermore, in the interest information detection device 2000, instead of individually comparing the feature quantities of each video frame 22 with the text feature quantities 30, the object feature quantity sequence 40 is compared with the text feature quantities 30. Therefore, by considering the features of objects across multiple images, interest information 50 can be detected. Thus, compared to the case where the feature quantities of each video frame 22 are individually compared with the text feature quantities 30, interest information 50 can be detected with higher accuracy.
[0025] Furthermore, the interest information detection device 2000 utilizes time-series data of object features (object feature sequence 40). Therefore, it can detect interest information 50 by considering the characteristics of changes in the state of the object. Thus, even when the state of the object changes in the video data 20, interest information 50 can be detected with high accuracy from the video data 20.
[0026] The interest information detection device 2000 of this embodiment will be described in more detail below.
[0027] <Example of Functional Configuration> Figure 2 is a block diagram illustrating the functional configuration of the interest information detection device 2000. For example, the interest information detection device 2000 has an acquisition unit 2020, a calculation unit 2040, a first generation unit 2060, and a detection unit 2080. The acquisition unit 2020 acquires query text data 10 and video data 20. The calculation unit 2040 calculates text features 30 from the query text data 10. The first generation unit 2060 generates an object feature sequence 40 for each object from the video data 20. The detection unit 2080 detects interest information 50 from the video data 20 using the video data 20 and the object feature sequence 40 for each object.
[0028] <Example of Hardware Configuration> Each functional component of the interest information detection device 2000 is implemented by hardware that realizes each functional component. Here, the hardware that realizes each functional component is, for example, a hardwired electronic circuit. In addition, each functional component of the interest information detection device 2000 is implemented by a combination of hardware and software. Here, a specific example of a combination of hardware and software is a combination of an electronic circuit and a program that controls it.
[0029] Figure 3 is a block diagram illustrating the hardware configuration of the computer 1000 that implements the interest information detection device 2000. The computer 1000 is any computer. For example, the computer 1000 is a stationary computer such as a PC (Personal Computer) or a server machine. Alternatively, the computer 1000 is a portable computer such as a smartphone or a tablet terminal. Alternatively, the computer 1000 is an integrated circuit such as a SoC (System on Chip). The computer 1000 may be a dedicated computer designed to implement the interest information detection device 2000, or it may be a general-purpose computer.
[0030] For example, by installing a predetermined application on the computer 1000, the various functions of the interest information detection device 2000 are realized on the computer 1000. The above application consists of a program for realizing each functional component of the interest information detection device 2000.
[0031] The method of obtaining the above program is arbitrary. For example, the program can be obtained from the storage medium on which it is stored. The storage medium on which the program is stored can be any storage medium, such as a DVD (Digital Versatile Disk) or a USB (Universal Serial Bus) memory. Alternatively, the program can be obtained by downloading it from a server device that manages the storage device on which the program is stored.
[0032] Computer 1000 includes a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. The bus 1020 is a data transmission path for the processor 1040, memory 1060, storage device 1080, input / output interface 1100, and network interface 1120 to send and receive data to and from each other. However, the method of connecting the processor 1040 and the other components is not limited to bus connection.
[0033] The processor 1040 is a variety of processor such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an FPGA (Field-Programmable Gate Array). The memory 1060 is a main memory device implemented using RAM (Random Access Memory), etc. The storage device 1080 is an auxiliary storage device implemented using a hard disk, SSD (Solid State Drive), memory card, or ROM (Read Only Memory), etc.
[0034] The input / output interface 1100 is an interface for connecting the computer 1000 with input / output devices. For example, input devices such as keyboards and output devices such as display devices are connected to the input / output interface 1100.
[0035] The network interface 1120 is an interface for connecting the computer 1000 to a network. This network may be a LAN (Local Area Network) or a WAN (Wide Area Network).
[0036] The storage device 1080 stores programs that implement each functional component of the interest information detection device 2000 (programs that implement the aforementioned applications). The processor 1040 reads these programs into the memory 1060 and executes them to implement each functional component of the interest information detection device 2000.
[0037] The interest information detection device 2000 may be implemented using one computer 1000 or multiple computers 1000. In the latter case, the configuration of each computer 1000 does not need to be the same and can be different.
[0038] <Processing Flow> Figure 4 is a flowchart illustrating the processing flow performed by the interest information detection device 2000. The acquisition unit 2020 acquires query text data 10 (S102). The acquisition unit 2020 acquires video data 20 (S104). The calculation unit 2040 calculates text features 30 from the query text data 10 (S106). The first generation unit 2060 generates object feature sequence 40 for each object from the video data 20 (S108). The detection unit 2080 detects interest information 50 from the video data 20 using the text features 30 and the object feature sequence 40 for each object (S110).
[0039] The processing flow performed by the interest information detection device 2000 is not limited to the flow shown in Figure 4. For example, the acquisition of query text data 10 (S102) and the acquisition of video data 20 may be performed in the reverse order of the order shown in Figure 4, or they may be performed in parallel. However, as will be described later, if the video data 20 to be acquired is determined based on the content of the query text data 10, the query text data 10 is acquired before the video data 20.
[0040] Also, the calculation of the text feature quantity 30 (S106) and the generation of the object feature quantity column 40 (S108) may be executed in an order reverse to the order shown in FIG. 4, or may be executed in parallel.
[0041] <Acquisition of query text data 10: S102> The acquisition unit 2020 acquires the query text data 10 (S102). There are various methods for the acquisition unit 2020 to acquire the query text data 10. For example, the acquisition unit 2020 provides a screen for inputting the query text data 10 to the user of the interest information detection device 2000. The acquisition unit 2020 acquires the text input on this screen as the query text data 10. Additionally, for example, the acquisition unit 2020 acquires the query text data 10 by receiving the query text data 10 transmitted from another device such as a terminal used by the user. Here, the terminal is, for example, a PC or a smartphone.
[0042] Additionally, for example, the query text data 10 may be stored in advance in the storage unit in a manner that can be acquired from the interest information detection device 2000. In this case, the acquisition unit 2020 acquires the query text data 10 by reading out the query text data 10 from the storage unit.
[0043] As described above, for example, the query text data 10 indicates one or more object relationships. There are various ways to represent the object relationships in the query text data 10. For example, the query text data 10 represents the object relationship in a sentence such as "a person is sitting on a chair". Additionally, for example, the query text data 10 represents the object relationship in a triplet such as (person, chair, sitting).
[0044] The query text data 10 can indicate a plurality of conditions. Here, the plurality of conditions are, for example, a plurality of object relationships. When a plurality of conditions are indicated by the query text data 10, the conditions represented by the query text data 10 are represented by the logical sum or logical product of the plurality of conditions.
[0045] Whether to adopt the logical sum or the logical product of a plurality of conditions as the search condition may be determined in advance and fixed in the interest information detection device 2000, or may be specified in the query text data 10. In the latter case, the logical sum of a plurality of conditions can be expressed as, for example, "a person is sitting on a chair, or a bag is placed on the floor". Also, the logical product of a plurality of conditions can be expressed as, for example, "a person is sitting on a chair and a bag is placed on the floor".
[0046] In addition, when three or more conditions are shown in the query text data 10, the logical sum and the logical product can be arbitrarily combined.
[0047] Here, the acquisition unit 2020 may acquire data (hereinafter referred to as query data) that includes both the query text data 10 and other search conditions. In this case, the query text data 10 is acquired as part of the query data. The method for acquiring the query data is the same as the method for acquiring the query text data 10 described above.
[0048] The search conditions that can be specified by the query data are various. For example, the search condition is a condition for specifying the video data 20 to be searched. The video data 20 can be specified, for example, by the file name of the video data.
[0049] In addition, for example, the search condition indicates the time range to be searched. When the time range is shown as the search condition, the acquisition unit 2020 acquires each video data within the time range as the video data 20. For example, the acquisition unit 2020 acquires the video data 20 by calculating only the video frames generated within the time range from the video data stored in the storage unit or the video data input by the user.
[0050] In addition, for example, the search condition indicates the space range to be searched. Here, the space range is, for example, a geographical range. When the space range is shown as the search condition, for example, the acquisition unit 2020 acquires the video data generated by the camera located within the space range as the video data 20.
[0051] The spatial range is specified by location information, such as an address or GPS (Global Positioning System) coordinates. In this case, information indicating the camera's installation location (hereinafter referred to as installation location information) is prepared in advance. The installation location information is stored in an arbitrary memory unit in a manner accessible from the interest information detection device 2000, for example. The acquisition unit 2020 uses the installation location information to identify cameras installed at the locations indicated in the search conditions.
[0052] The spatial range may be specified by a camera identifier. In other words, the search condition may be a condition that specifies a camera. In this case, the acquisition unit 2020 acquires the video data generated by the camera specified by the search condition as video data 20.
[0053] <Acquisition of video data 20: S104> The acquisition unit 2020 acquires the video data 20 (S104). There are various ways in which the acquisition unit 2020 acquires the video data 20. For example, the video data 20 is stored in a storage unit in advance in a manner that can be acquired from the interest information detection device 2000. In this case, the acquisition unit 2020 acquires the video data 20 from the storage unit.
[0054] The acquisition unit 2020 may acquire all of the video data 20 stored in the storage unit, or it may acquire only some of the video data 20. In the latter case, for example, the acquisition unit 2020 acquires video data that matches the search criteria as video data 20, as described above.
[0055] The acquisition unit 2020 may receive video data transmitted from another device, such as a terminal used by the user, and use the received video data as video data 20. Here, the terminal may be, for example, a PC or a smartphone.
[0056] Here, the acquisition unit 2020 may acquire multiple video data 20. In this case, the acquisition unit 2020 detects interest information 50 for each of the multiple video data 20. Specifically, the processes S108 and S110 in the flowchart of Figure 4 are performed for each video data 20.
[0057] <Calculation of text features 30: S106> The calculation unit 2040 calculates text features 30, which are the features of the text represented by the query text data 10 (S106). Text features 30 are calculated using a machine learning model, such as a neural network. Hereinafter, the model that calculates text features 30 from the query text data 10 will be called a text feature calculation model. A text feature calculation model may also be called a text encoder.
[0058] For example, a text feature extraction model is configured to output the features of a given text token in response to that token being input. Text tokens can be words, subwords, and so on.
[0059] In this case, the calculation unit 2040 divides the query text data 10 into multiple text tokens and inputs each text token into the text feature calculation model. As a result, the calculation unit 2040 obtains features for each of the multiple text tokens that make up the query text data 10. In this way, a column of features is generated from the column of text tokens. For example, the calculation unit 2040 uses this column of features as text features 30.
[0060] The text feature extraction model may be configured to extract features from each text token, taking into account the characteristics of its relationships with other text tokens. Here, the relationships between text tokens in the query text data 10 can also be called the context in the query text data 10. As a model that takes context into account, machine learning models that handle time-series data, such as RNN (Recurrent Neural Network) or Transformer encoders, can be employed.
[0061] Even when context is considered, for example, the calculation unit 2040 inputs a column of text tokens obtained from the query text data 10 into the text feature calculation model. As a result, the feature of each text token is output from the text feature calculation model. The calculation unit 2040 uses the column of feature outputs from the text feature calculation model as the text features 30.
[0062] In addition, for example, a text feature extraction model may be configured to output a single text feature composed of a sequence of text tokens, in response to a sequence of text tokens being input. Here, the text is, for example, a sentence. Such a text feature extraction model can also utilize machine learning models that handle time series data.
[0063] In this case, for example, the calculation unit 2040 inputs a column of text tokens obtained from the query text data 10 into the text feature calculation model. The calculation unit 2040 then uses the features output from the text feature calculation model as text features 30.
[0064] <Generation of object feature sequence 40: S108> The first generation unit 2060 generates an object feature sequence 40 for each object from the video data 20 (S108). To do this, the first generation unit 2060 detects objects from each video frame 22 included in the video data 20. Furthermore, for each object detected from the video frame 22, the first generation unit 2060 calculates the object feature quantities 42 of that object. Then, for each object, it generates an object feature sequence 40, which is time-series data of multiple object feature quantities 42 calculated for that object.
[0065] Here, it is preferable that the object feature sequence 40 is generated as time-series data of the same length as the video data 20. That is, it is preferable that the number of object features 42 constituting the object feature sequence 40 is the same as the number of video frames 22 constituting the video data 20.
[0066] Thus, assume that the length of the object feature sequence 40 is the same as the length of the video data 20. In this case, in the object feature sequence 40 of a certain object X, the object feature 42 corresponding to the video frame 22 containing object X represents the feature of object X detected from that video frame 22. On the other hand, in the object feature sequence 40, the object feature 42 corresponding to the video frame 22 that does not contain object X represents a predetermined feature. This predetermined feature is, for example, a feature where the value of all cells is 0.
[0067] There are various specific methods for detecting objects from video frames 22. For example, the first generation unit 2060 detects objects from video frames 22 using a machine learning model that has been pre-trained to detect objects from images. Here, the machine learning model used to implement this model is, for example, a neural network. Unless otherwise specified, the examples of machine learning models include neural networks. Hereafter, this model will be referred to as the object detection model.
[0068] An object detection model is configured, for example, to take an image as input and output information about the image region representing each of the one or more objects contained in that image. Hereafter, the image region representing an object will be called the object region, and the information about the object region will be called the object region information.
[0069] An object region is, for example, an image region representing the interior of the bounding rectangle of an object. Object region information includes, for example, image data of the object region, the position of a specific point within the object region, and the size of the object region. Here, the specific point within the object region is, for example, the upper left corner. The size of the object region is, for example, the width and height.
[0070] The first generation unit 2060 inputs each video frame 22 to the object detection model. This allows the first generation unit 2060 to obtain object region information for each video frame 22. If a video frame 22 contains multiple objects, object region information is obtained for each object.
[0071] The first generation unit 2060 identifies objects using object region information obtained from each of the multiple video frames 22. That is, it determines whether objects detected in different video frames 22 are the same object or not. Various methods, such as tracking, can be used to identify objects across multiple images in a time series.
[0072] As a result of object identification, the first generation unit 2060 generates object tracking information that includes a history of object region information (time-series data of object region information) for each object. Figure 5 is a diagram illustrating the configuration of object tracking information. In Figure 5, the object tracking information 60 shows an identifier 62 and a history 64.
[0073] Identifier 62 is a unique identifier assigned to each object.
[0074] The history 64 shows one or more pairs of time point 66 and object region information 68 for an object to which the identifier shown in the corresponding identifier 62 has been assigned. The time point 66 indicates the generation time of the video frame 22 to which the corresponding object region information 68 was generated. The object region information 68 shows the image data of the object region, the position of the object region, and the size of the object region.
[0075] In Figure 5, img1 and img2 represent image data of an object region extracted from video frame 22. p1 and p2 represent the position of the object region. w1 and w2 represent the width of the object region. h1 and h2 represent the height of the object region.
[0076] The first generation unit 2060 generates object feature quantities 42 for each object from the object region information 68 shown in the object's history 64. In this way, time-series data of object feature quantities 42 is obtained for each object.
[0077] Object features 42 are calculated using machine learning models, such as neural networks. Hereafter, models that calculate features of object regions will be referred to as object feature calculation models. Any model capable of calculating features from images can be used as an object feature calculation model.
[0078] For example, a text feature model and an object feature model are configured to share a feature space. With this configuration, the closer the distance between the features obtained from text using the text feature model and the features obtained from images using the object feature model, the more similar the situations represented by the text and the image are.
[0079] Text and object feature acquisition models configured to share a feature space can be, for example, the text encoder and image encoder in a vision and language model, respectively. A method for training the text and object feature acquisition models so that they share a feature space can be the same as the method used to train the text and image encoders in a vision and language model.
[0080] The first generation unit 2060 calculates the feature quantities of an object region by inputting the object region of an object detected from a video frame 22 into an object feature calculation model. For example, the first generation unit 2060 uses the feature quantities of the object region calculated by the object feature calculation model as object feature quantities 42.
[0081] The first generation unit 2060 may convert the position and size of the object region into features and further include these features in the object features 42. For example, the first generation unit 2060 generates the object features 42 by concatenating the features of the object region calculated using an object feature calculation model, the features of the position of the object region, and the features of the size of the object region using a method such as concatenation. Various methods can be used to convert the position and size of the image region into features.
[0082] Similarly, the first generation unit 2060 may convert the type of object into a feature and further include that feature in the object feature 42.
[0083] The first generation unit 2060 may be configured to detect only some objects from the video frame 22. For example, the types of objects to be acquired from the video frame 22 are predetermined. Here, the types of objects are, for example, people or vehicles. The first generation unit 2060 detects only the objects of the predetermined types from the video frame 22.
[0084] Alternatively, for example, the first generation unit 2060 may use the query text data 10 to identify objects to be detected from the video frame 22. Specifically, the first generation unit 2060 detects words representing the name or type of an object from the query text data 10. Then, the first generation unit 2060 detects only the objects identified by the detected words from the video frame 22.
[0085] The first generation unit 2060 may also perform the process of generating an object feature sequence 40 from the video data 20 in advance. That is, the first generation unit 2060 may acquire video data 20 that can be compared with the query text data 10 in advance, and generate an object feature sequence 40 from the acquired video data 20. By generating the object feature sequence 40 in advance, the time and computing resources required to detect the interest information 50 can be reduced.
[0086] If the object feature sequence 40 is generated in advance, the first generation unit 2060 stores one or more object feature sequences 40 generated from the video data 20 in the storage unit, associating them with the identifier of the video data 20. When the detection of interest information 50 is performed, the interest information detection device 2000 receives a designation of the video data 20 and retrieves the object feature sequence 40 corresponding to the designated video data 20 from the storage unit.
[0087] <Detection of interest information 50: S110> The detection unit 2080 detects interest information 50 using the text features 30 and the object feature sequence 40 for each object (S110). To do this, for example, the detection unit 2080 further generates a feature sequence (hereinafter referred to as the first related feature sequence) for each object using the object feature sequence 40 for each object and the text features 30. Then, the detection unit 2080 detects interest information 50 using the first related feature sequence for each object.
[0088] The following sections will specifically explain how to generate the first related feature sequence and how to detect interest information 50 from the first related feature sequence.
[0089] <<Method for generating the first related feature sequence>> For example, the detection unit 2080 generates the first concatenated data by concatenating the object feature sequence 40 and text feature sequence 30 for each object. Then, the detection unit 2080 generates the first related feature sequence for each object using the first concatenated data of that object.
[0090] There are various methods for generating the first concatenated data from text features 30 and object feature sequences 40. For example, the detection unit 2080 generates the first concatenated data having the configuration shown in Figure 6. Figure 6 is a diagram illustrating the configuration of the first concatenated data. In this example, the detection unit 2080 generates the first concatenated data 70 by concatenating the text features 30 with the entire object feature sequence 40.
[0091] In the example in Figure 6, the text feature 30 is concatenated at the beginning of the object feature sequence 40. However, the text feature 30 may also be concatenated at the end of the object feature sequence 40.
[0092] In addition, for example, the detection unit 2080 generates first concatenated data 70 having the configuration shown in Figure 7. Figure 7 is a diagram illustrating the configuration of the first concatenated data 70. In this example, the detection unit 2080 generates the first concatenated data 70 by concatenating text features 30 with each object feature 42 included in the object feature sequence 40.
[0093] In the example in Figure 7, the text features 30 are concatenated before each object feature 42. However, the text features 30 may also be concatenated after each object feature 42.
[0094] The detection unit 2080 generates a first related feature sequence 80 from the first linked data 70. The first related feature sequence 80 is generated using a machine learning model, such as a neural network. Hereinafter, the model used to generate the first related feature sequence 80 will also be called the first related feature sequence generation model.
[0095] The first related feature sequence generation model is configured to output a different feature sequence in response to the feature sequence it receives as input. For example, the first related feature sequence generation model is configured to extract the features of the relationships between features (i.e., context) from the input feature sequence and output a feature sequence in which the features of that context are embedded.
[0096] By utilizing a first relation feature sequence generation model configured to extract features of relationships between features, a first relation feature sequence 80 is generated considering the relationships between multiple object features 42 that constitute the object feature sequence 40. Here, the relationships between the multiple object features 42 that constitute the object feature sequence 40 can also be said to be the relationships between the features of objects obtained from multiple video frames 22. By utilizing such a first relation feature sequence 80, interest information 50 is detected considering the relationships between the features of objects obtained from multiple video frames 22. Therefore, compared to the case where interest information 50 is detected by analyzing the video frames 22 individually, interest information 50 can be detected with higher accuracy.
[0097] Specific configurations for the first related feature generation model include, for example, the encoder configuration of transformer or the Video-text Fusion configuration disclosed in Non-Patent Document 1. However, the first related feature generation model can be constructed using any machine learning model and is not limited to models with the same configuration as the transformer encoder or Video-text Fusion.
[0098] Figure 8 illustrates the process of generating the first related feature sequence 80 using the first related feature sequence generation model. The detection unit 2080 concatenates the text features 30 and the object feature sequence 40 to generate the first concatenated data 70. The detection unit 2080 then inputs the first concatenated data 70 into the first related feature sequence generation model 90 to generate the first related feature sequence 80.
[0099] However, some machine learning models cannot take into account the order of data in a time series simply by inputting that data. An example of such a model is the transformer encoder.
[0100] Therefore, it is preferable for the detection unit 2080 to add information representing the time-series position of each data to the text features 30 and object feature sequences 40 included in the first concatenated data 70. For example, a position code can be used to represent the time-series position of the data.
[0101] For example, the detection unit 2080 adds information representing the rank of each object feature 42 included in the first concatenated data 70. The rank assigned to each object feature 42 is the rank of the video frame 22 used to generate that object feature 42 in the video data 20.
[0102] For example, suppose video data 20 contains four video frames f1, f2, f3, and f4 in that order. Also, suppose object X is detected from each of these video frames. In this case, in the object feature sequence 40 for object X, the object feature 42 of object X calculated from f1 is assigned rank 1. The object feature 42 calculated from f2 is assigned rank 2.
[0103] Similarly, if the text feature 30 is a sequence of features, the detection unit 2080 adds information representing the rank of each text token to the text feature 30 when generating the first concatenated data 70.
[0104] <<Method for detecting interest information 50>> The detection unit 2080 detects interest information 50 from the video data 20 using the first related feature sequence 80. The detection methods for the interest interval 52 and the interest object 54 will be described below.
[0105] <<<Method for detecting the interest interval 52>>> The detection unit 2080 detects the interest interval 52 using the first related feature sequence 80 for each object. Specifically, the detection unit 2080 uses the first related feature sequence 80 to detect intervals from the video data 20 that have a high degree of association with the query text data 10. The detected intervals are then treated as the interest interval 52.
[0106] The detection of the interest interval 52 is performed using a machine learning model, such as a neural network. This machine learning model is also called an interest interval detection model. For the configuration of the interest interval detection model, for example, the Head configuration disclosed in Non-Patent Document 1 can be used. However, the interest interval detection model can be composed of any machine learning model, and its configuration is not limited to the Head configuration described above.
[0107] The interest interval detection model outputs information that can identify an interest interval in response to the input of the first related feature sequence 80. This information that can identify an interest interval is, for example, a combination of the start and end positions of the interest interval. The start position of the interest interval is represented, for example, by the identifier of the video frame 22 located at the beginning of the interest interval or by the generation time of the said video frame 22. Similarly, the end position of the interest interval is represented, for example, by the identifier of the video frame 22 located at the end of the interest interval or by the generation time of the said video frame 22. Here, the identifier of the video frame 22 is, for example, the frame number.
[0108] Here, even if the first related feature sequence 80 is used, it is not guaranteed that the interest interval 52 will be detected. For example, if the first related feature sequence 80 generated for objects unrelated to the query is used, the interest interval 52 will not be detected. Furthermore, it is possible that the interest interval 52 does not exist in the video data 20 at all. In other words, it is possible that the video data 20 does not contain any scenes that are highly related to the query text data 10. To put it another way, it is possible that the video data 20 does not contain any scenes that match the query text data 10.
[0109] Therefore, it is preferable that the interest interval detection model be configured to output data indicating that the interest interval 52 is not detected. For example, the interest interval detection model is configured to output predetermined values as the start and end positions when the interest interval 52 is not detected. Here, these predetermined values are, for example, 0 or a negative value.
[0110] For example, the detection unit 2080 inputs each of the multiple first related feature sequences 80 into the interest interval detection model. Suppose all outputs obtained from the interest interval detection model indicate that the interest interval 52 is not detected. In this case, the detection unit 2080 determines that the video data 20 does not contain the interest interval 52. On the other hand, suppose that at least one output obtained from the interest interval detection model is information that can identify the interest interval 52. In this case, the interest interval 52 identified by that information is detected.
[0111] Alternatively, for example, the detection unit 2080 may consolidate multiple first related feature sequences 80 into a single first related feature sequence 80 and input this single first related feature sequence 80 into the interest interval detection model. Various pooling processes, such as mean pooling and maximum pooling, can be used to consolidate multiple first related feature sequences 80 into a single first related feature sequence 80. More specifically, the detection unit 2080 performs pooling on multiple first related feature sequences 80 for each set of features that are at the same position in the time series. The pooling process performed on multiple features involves pooling cells that are at the same position in each of these multiple features.
[0112] Note that the interest interval 52 detected by the interest interval detection model is not limited to one. The interest interval detection model may be configured to output multiple interest intervals 52.
[0113] Here, we will explain the training method for the interest interval detection model. Hereafter, the device used to train the model will be referred to as the training device. The training device may be the interest information detection device 2000, or it may be a device other than the interest information detection device 2000. In the latter case, the hardware configuration of the training device may be similar to that of the interest information detection device 2000 (see Figure 3).
[0114] The interest interval detection model is trained, for example, together with the first related feature sequence generation model 90 described above. Figure 9 illustrates the training method for the first related feature sequence generation model 90 and the interest interval detection model.
[0115] The first related feature generation model 90 and the interest interval detection model 100 are trained using multiple training samples 110. Each training sample 110 includes a combination of a training query 112, a training image sequence 114, and a ground truth interval 116.
[0116] The training query 112 is data corresponding to the query text data 10. The training image sequence 114 is data corresponding to the video data 20. The ground truth interval 116 represents the ground truth interest interval that should be output by the interest interval detection model 100 for the combination of the training query 112 and the training image sequence 114. That is, the ground truth interval 116 indicates the interval of scenes in the training image sequence 114 that match the training query 112.
[0117] The training device obtains the first concatenated data 70 using the training query 112 and the training image sequence 114. To do this, the training device calculates text features 30 from the training query 112. The training device also generates an object feature sequence 40 from the training image sequence 114. Then, the training device concatenates the text features 30 and the object feature sequence 40 to generate the first concatenated data 70.
[0118] The training device inputs the first linked data 70 into the first related feature sequence generation model 90 to obtain the first related feature sequence 80. The training device then inputs the first related feature sequence 80 into the interest interval detection model 100 to obtain the interest interval 52.
[0119] The training device calculates the loss using the interest interval 52 and the ground truth interval 116. Based on the calculated loss, the training device updates the trainable parameters of the first related feature sequence generation model 90 and the interest interval detection model 100, respectively. Here, trainable parameters include, for example, the weights and biases of the neural network.
[0120] The training device repeatedly updates the parameters of the first related feature sequence generation model 90 and the interest interval detection model 100 using multiple training samples 110. In this way, the training device trains the first related feature sequence generation model 90 and the interest interval detection model 100.
[0121] <<<Method for detecting the object of interest 54>>> The detection unit 2080 detects the object of interest 54 using the first related feature sequence 80 for each object. Specifically, for each object, the detection unit 2080 determines whether or not that object is the object of interest 54 using the first related feature sequence 80 for that object.
[0122] The determination of whether an object corresponding to the first related feature sequence 80 is an object of interest 54 is performed, for example, using a machine learning model such as a neural network. This machine learning model is called an object of interest determination model. For the configuration of the object of interest determination model, for example, the configuration of Head disclosed in Non-Patent Document 1 can be adopted. However, the object of interest determination model can be constructed with any machine learning model and is not limited to a model with the same configuration as Head.
[0123] The interest object determination model is configured to output a determination result indicating whether or not the object corresponding to the input sequence of first related features 80 is an interest object 54, in response to the input sequence of first related features 80. More specifically, the interest object determination model calculates the probability that the object corresponding to the input sequence of first related features 80 is an interest object. If the calculated probability is greater than or equal to a predetermined threshold, the interest object determination model outputs a value indicating that the object corresponding to the input sequence of first related features 80 is an interest object. Here, the output value is, for example, 1. On the other hand, if the calculated probability is less than the predetermined threshold, the interest object determination model outputs a value indicating that the object corresponding to the input sequence of first related features 80 is not an interest object. Here, the output value is, for example, 0.
[0124] The detection unit 2080 determines whether each object is an object of interest 54 by inputting the first related feature sequence 80 of each object into the object of interest determination model. For example, suppose that the first related feature sequence 80 has been obtained for each of objects X1, X2, and X3. In this case, the detection unit 2080 inputs the first related feature sequence 80 of object X1, the first related feature sequence 80 of object X2, and the first related feature sequence 80 of object X3 into the object of interest determination model, respectively.
[0125] Here, suppose the object of interest determination model outputs 1 in response to receiving the first related feature sequence 80 corresponding to object X1. Also, suppose the object of interest determination model outputs 0 in response to receiving the first related feature sequence 80 corresponding to object X2. Furthermore, suppose the object of interest determination model outputs 0 in response to receiving the first related feature sequence 80 corresponding to object X3. In this case, the detection unit 2080 detects object X1 as the object of interest 54.
[0126] The training method for the interest object determination model will be explained. The interest object determination model can be trained together with the first related feature sequence generation model 90 in the same way as the interest interval detection model 100. Figure 10 is an example of the training method for the first related feature sequence generation model 90 and the interest object determination model.
[0127] The training of the object of interest determination model 120 is performed using training samples 110, similar to the training of the interval detection model 100. However, the training samples 110 used to train the object of interest determination model 120 include a ground truth flag 118 instead of a ground truth interval 116. The training samples 110 include a ground truth flag 118 for each object included in the training image sequence 114.
[0128] The ground truth flag 118 indicates whether the corresponding object is an object of interest. If the corresponding object is an object of interest, the ground truth flag 118 shows 1. On the other hand, if the corresponding object is not an object of interest, the ground truth flag 118 shows 0.
[0129] The training device generates a first related feature sequence 80 for each object from the training query 112 and the training image sequence 114, similar to how it trains the interest interval detection model 100. The training device then inputs the first related feature sequence 80 for each object into the interest object determination model 120.
[0130] The interest object determination model 120 calculates the probability that an object corresponding to the first related feature sequence 80 is an interest object, based on the input of the first related feature sequence 80. The training device calculates a loss using the calculated probability of each object being an interest object and the ground truth flag 118 included in the training sample 110 for each object. The training device then updates the parameters of the first related feature sequence generation model 90 and the interest object determination model 120 based on the calculated loss.
[0131] The training device repeatedly updates the parameters of the first related feature sequence generation model 90 and the interest object determination model 120 using multiple training samples 110. In this way, the training device trains the first related feature sequence generation model 90 and the interest object determination model 120.
[0132] <Output by the Interest Information Detection Device 2000> The interest information detection device 2000 outputs information representing the processing result (hereinafter referred to as output information). The output information includes the detected interest information 50. Specifically, the output information includes information about the detected interest interval 52 and information about the detected interest object 54.
[0133] The output information may include any information that identifies the interest interval 52, as information relating to the interest interval 52. For example, the output information may indicate the start and end positions of the interest interval 52. Alternatively, the output information may include video data consisting of video frames 22 included in the interest interval 52. This video data can also be described as a short clip consisting of video frames 22 from the start position of the interest interval 52 to the end position of the interest interval 52.
[0134] The output information may include any information that identifies the object of interest 54, as information relating to the object of interest 54. For example, the output information may include a sequence of object features 40 of the object of interest 54. In addition, the output information may include, for example, the history 64 of the object of interest 54.
[0135] When the detection of interest information 50 is performed on multiple video data 20, it is preferable that the output information further indicate the identifier of the video data 20 in which the interest information 50 was detected. Here, the identifier of the video data 20 is, for example, the file name. In addition, the output information may indicate information about the camera that generated the video data 20, either together with the identifier of the video data 20 or in place of the identifier of the video data 20. Here, the information about the camera is, for example, the camera identifier or installation location.
[0136] By obtaining information about the camera that generated the video data 20 in which the interest interval 52 was detected, the user of the interest information detection device 2000 can easily understand when and where scenes matching the query text data 10 were captured. Furthermore, by obtaining information about the camera that generated the video data 20 in which the object of interest 54 was detected, the user of the interest information detection device 2000 can easily understand where objects matching the query text data 10 were captured. In other words, the user can easily understand the location where objects of interest appeared, or the location and time when events of interest occurred.
[0137] <Modification> In the above description, the detection unit 2080 uses the first related feature sequence 80 to detect the interest information 50. However, the method for detecting the interest information 50 is not limited to the method using the first related feature sequence 80. Below, other examples of how the detection unit 2080 detects the interest information 50 will be described.
[0138] Figure 11 illustrates the process by which interest information 50 is detected. For each object, the detection unit 2080 generates a second related feature sequence 130 from the object feature sequence 40 of that object. The second related feature sequence 130 is a feature sequence in which the contextual features of the object feature sequence 40 are embedded.
[0139] The second related feature sequence 130 is generated, for example, using the second related feature sequence generation model 140. The detection unit 2080 obtains the second related feature sequence 130 for each object by inputting the object feature sequence 40 for each object into the second related feature sequence generation model 140.
[0140] The second related feature generation model 140 is a machine learning model such as a neural network. The configuration of the second related feature generation model 140 can be the same as that of the first related feature generation model 90.
[0141] Furthermore, the detection unit 2080 generates second concatenated data 150 for each object by concatenating the second related feature sequence 130 and the text feature sequence 30 of that object. The method for generating the second concatenated data 150 by concatenating the text feature sequence 30 and the second related feature sequence 130 can be the same as the method for generating the first concatenated data 70 by concatenating the text feature sequence 30 and the object feature sequence 40.
[0142] The detection unit 2080 detects interest information 50 using second linked data 150 generated for each object. When an interest interval 52 is detected using the second linked data 150, the interest interval detection model 100 is configured to output the interest interval 52 in response to the input of the second linked data 150. Furthermore, when an object of interest 54 is detected using the second linked data 150, the object of interest determination model 120 is configured to output a determination result indicating whether or not the object corresponding to the second linked data 150 is an object of interest in response to the input of the second linked data 150.
[0143] The process of generating the second related feature sequence 130 from the video data 20 may be performed in advance. That is, the interest information detection device 2000 may acquire video data 20 that can be compared with the query text data 10 in advance, and generate the second related feature sequence 130 from the acquired video data 20. By generating the second related feature sequence 130 in advance, the time and computing resources required to detect the interest information 50 can be reduced.
[0144] If the second related feature sequence 130 is generated in advance, the interest information detection device 2000 stores the second related feature sequence 130 generated from the video data 20 in the storage unit, associating it with the identifier of the video data 20. When the detection of interest information 50 is performed, the interest information detection device 2000 receives a designation of the video data 20 and retrieves the second related feature sequence 130 corresponding to the designated video data 20 from the storage unit.
[0145] The training method for the interest interval detection model 100 using the second concatenated data 150 is the same as the training method for the interest interval detection model 100 using the first related feature sequence 80. Figure 12 is an example of the training method for the interest interval detection model 100. The training device generates an object feature sequence 40 for each object from the training image sequence 114. Furthermore, the training device inputs the object feature sequence 40 of each object into the second related feature sequence generation model 140 to obtain a second related feature sequence 130 for each object. The training device generates second concatenated data 150 for each object by combining the text features 30 obtained from the training query 112 with the second related feature sequence 130 of each object. The training device inputs the second concatenated data 150 into the interest interval detection model 100 to detect an interest interval 52. The training device calculates a loss from the detected interest interval 52 and ground truth interval 116, and updates the interest interval detection model 100 and the second related feature sequence generation model 140 with the calculated loss.
[0146] The training method for the interest object determination model 120 using the second linked data 150 is the same as the training method for the interest object determination model 120 using the first related feature sequence 80. Figure 13 is an example of the training method for the interest object determination model 120. In the example in Figure 13, the process for generating the second linked data 150 is the same as the process for generating the second linked data 150 in the example in Figure 12.
[0147] The training device inputs the second linked data 150 for each object into the interest interval detection model 100 to calculate the probability that each object is an object of interest. The training device calculates a loss from the probability that each object is an object of interest and the ground truth flag 118 for each object, and updates the object of interest determination model 120 and the second related feature sequence generation model 140 with the calculated loss.
[0148] [Embodiment 2] <Overview> Figure 14 is a diagram illustrating an overview of the operation of the interest information detection device 2000. Here, Figure 14 is a diagram intended to facilitate understanding of the overview of the interest information detection device 2000, and the operation of the interest information detection device 2000 is not limited to the operation shown in Figure 14.
[0149] The interest information detection device 2000 of Embodiment 2 further generates an image feature sequence 160 from the video data 20. The image feature sequence 160 is time-series data of feature quantities (image feature quantities 162) calculated from each of the multiple video frames 22 included in the video data 20. The image feature quantities 162 represent the outward appearance of the entire scene captured in the video frame 22. In other words, the image feature quantities 162 represent the outward appearance of the entire image captured in the video frame 22.
[0150] The interest information detection device 2000 detects interest information 50 from video data 20 using text features 30, object feature sequences 40 generated for each object, and image feature sequences 160.
[0151] <Example of effect> According to the interest information detection device 2000, in addition to the object feature sequence 40, which is a sequence of feature sequences for each object, an image feature sequence 160, which is a sequence of feature sequences for the entire video frame 22, is also generated from the video frame 22. Therefore, in addition to the features of each object contained in the video frame 22, the features of the entire scene captured in the video frame 22 are further considered when detecting interest information 50. Thus, according to the interest information detection device 2000, information that matches the content of the query text data 10 can be detected from the video data 20 with higher accuracy.
[0152] The interest information detection device 2000 of this embodiment will be described in more detail below.
[0153] <Example of Functional Configuration> Figure 15 is a diagram illustrating the configuration of the interest information detection device 2000. In the example in Figure 14, the interest information detection device 2000 has a second generation unit 2100 in addition to the acquisition unit 2020, calculation unit 2040, first generation unit 2060, and detection unit 2080. In the interest information detection device 2000 of Embodiment 2, the operation of the acquisition unit 2020, calculation unit 2040, and first generation unit 2060 is the same as the operation in the interest information detection device 2000 of Embodiment 1.
[0154] The second generation unit 2100 generates an image feature sequence 160 from the video data 20. The detection unit 2080 of Embodiment 2 detects interest information 50 from the video data 20 using the text feature sequence 30, the object feature sequence 40 for each object, and the image feature sequence 160.
[0155] <Example of Hardware Configuration> The hardware configuration of the interest information detection device 2000 of Embodiment 2 can be represented in Figure 3, similar to the hardware configuration of the interest information detection device 2000 of Embodiment 1. However, the storage device 1080 of Embodiment 2 stores programs for realizing each function of the interest information detection device 2000 of Embodiment 2.
[0156] <Processing Flow> Figure 16 is a flowchart illustrating the processing flow performed by the interest information detection device 2000. Note that the contents of S102 to S108 shown in Figure 16 are the same as the contents of S102 to S108 shown in Figure 4.
[0157] The second generation unit 2100 generates an image feature sequence 160 from the video data 20 (S202). The detection unit 2080 uses the text features 30, the object feature sequence 40 for each object, and the image feature sequence 160 to detect information of interest from the video data 20 (S204).
[0158] <Generation of Image Feature Sequence 160: S202> The second generation unit 2100 generates an image feature sequence 160 by calculating image features 162 for each of the multiple video frames 22 contained in the video data 20 (S202). The process of calculating image features 162 from video frames 22 is performed using, for example, a machine learning model such as a neural network. This machine learning model is also called an image feature calculation model.
[0159] The second generation unit 2100 inputs each of the multiple video frames 22 into an image feature calculation model to obtain image features 162 for each video frame 22. Then, the second generation unit 2100 generates an image feature sequence 160 by arranging the multiple image features 162 obtained in this way in chronological order.
[0160] Image feature interpolation models, like object feature interpolation models, are preferably trained to share a feature space with text feature interpolation models. This means that the closer the distance between the features obtained from text using the text feature interpolation model and the features obtained from images using the image feature interpolation model, the more similar the situations represented by the text and the image are.
[0161] For example, an image feature extraction model is configured to share a feature space with a text feature extraction model. With this configuration, the closer the distance between the features obtained from text using the text feature extraction model and the features obtained from images using the object feature extraction model, the more similar the situations represented by the text and the image are.
[0162] As an image feature model that shares a feature space with a text feature model, for example, an image encoder in a vision and language model can be used. A method for training the text feature model and the image feature model so that they share a feature space can be used, similar to how text encoders and image encoders are trained in a vision and language model.
[0163] The image feature calculation model, like the object feature calculation model, is a model that calculates features from an image. Therefore, the image feature calculation model can be configured in the same way as the object feature calculation model. Furthermore, the interest information detection device 2000 may use the object feature calculation model as the image feature calculation model. In other words, a single model common to both the calculation of object features 42 and the calculation of image features 162 may be used.
[0164] The second generation unit 2100 may have previously performed the process of generating an image feature sequence 160 from the video data 20. That is, the second generation unit 2100 may have previously acquired video data 20 that can be compared with the query text data 10, and generated an image feature sequence 160 from the acquired video data 20.
[0165] If the image feature sequence 160 is generated in advance, the second generation unit 2100 stores the image feature sequence 160 generated from the video data 20 in the storage unit, associating it with the identifier of the video data 20. When the detection of interest information 50 is performed, the interest information detection device 2000 receives the designation of the video data 20 and retrieves the image feature sequence 160 corresponding to the designated video data 20 from the storage unit.
[0166] <Detection of interest information 50: S204> The detection unit 2080 detects interest information 50 from the video data 20 using text features 30, object feature sequences 40 for each object, and image feature sequences 160 (S204). As mentioned above, for example, the detection unit 2080 generates a first related feature sequence 80 for each object from the text features 30 and each object feature sequence 40. Furthermore, the detection unit 2080 generates a third related feature sequence using the text features 30 and image feature sequences 160. Then, the detection unit 2080 detects interest information 50 from the video data 20 using the first related feature sequence 80 and the third related feature sequence.
[0167] As explained below, the method for generating the third related feature sequence using the text feature sequence 30 and the image feature sequence 160 is the same as the method for generating the first related feature sequence 80 using the text feature sequence 30 and the object feature sequence 40.
[0168] <<Method for generating the third related feature sequence>> The detection unit 2080 concatenates the text feature 30 and the image feature sequence 160 to generate the third concatenated data. Then, the detection unit 2080 generates the third related feature sequence using the third concatenated data.
[0169] The method for generating third-concatenation data from text features 30 and image feature sequence 160 can be the same as the method for generating first-concatenation data 70 from text features 30 and object feature sequence 40. For example, the detection unit 2080 generates third-concatenation data by concatenating text features 30 with image feature sequence 160. Alternatively, for example, the detection unit 2080 generates third-concatenation data by concatenating text features 30 with each image feature 162 included in image feature sequence 160.
[0170] The detection unit 2080 generates a sequence of third-related features from the third-connected data. The sequence of third-related features is generated using a machine learning model, such as a neural network. Hereinafter, the model used to generate the sequence of third-related features will also be called the third-related feature calculation model.
[0171] Figure 17 illustrates a method for generating a third related feature sequence. The detection unit 2080 concatenates the text features 30 and the image feature sequence 160 to generate the third concatenated data 190. The detection unit 2080 then inputs the third concatenated data 190 into the third related feature sequence generation model 210. The third related feature sequence generation model 210 outputs a third related feature sequence 200 in response to the input of the third concatenated data 190.
[0172] The configuration of the third related feature sequence generation model 210 can be the same as that of the first related feature sequence generation model 90.
[0173] As mentioned above, it is preferable that the first concatenated data 70 be accompanied by information indicating the position of each data point. Similarly, in the third concatenated data 190, it is preferable that the text features 30 and image features 162 be accompanied by information indicating the position of each data point in the time series. For example, a position code can be used to indicate the position of each data point in the time series.
[0174] Specifically, in the third concatenated data 190, the detection unit 2080 adds information to each image feature 162 that represents the rank of that image feature 162 in the image feature sequence 160. The rank of the image feature 162 in the image feature sequence 160 is represented by the rank of the video frame 22 used to generate that image feature 162 in the video data 20.
[0175] Similarly, if the text feature 30 is a sequence of features, the detection unit 2080 adds information representing the rank of each text token to the text feature 30 in the third concatenated data 190.
[0176] <<Method for detecting interest information 50>> The detection unit 2080 detects interest information 50 from the video data 20 using the first related feature sequence 80 and the third related feature sequence 200. For example, the detection unit 2080 generates fourth concatenated data for each object by concatenating the third related feature sequence 200 with the first related feature sequence 80 of that object. The fourth concatenated data is generated, for example, by concatenating the first related feature sequence 80 with the third related feature sequence 200. Alternatively, for example, the fourth concatenated data may be generated for each object by performing feature-wise addition between the third related feature sequence 200 and the first related feature sequence 80 of that object.
[0177] The detection unit 2080 uses the fourth linked data to detect the interest interval 52, the interest object 54, or both. The interest interval 52 is detected by inputting each of the fourth linked data to the interest interval detection model 100. For this purpose, the interest interval detection model 100 is pre-trained to output information that can identify the interest interval 52 in response to the input of the fourth linked data. The method for training the interest interval detection model 100 in this way will be described later.
[0178] Similarly, an object of interest 54 is detected by inputting each fourth linked data into the object of interest determination model 120. To this end, the object of interest determination model 120 is pre-trained to output a determination result indicating whether or not the object corresponding to the fourth linked data is an object of interest, in response to the input of the fourth linked data. The method for training the object of interest determination model 120 in this way will be described later.
[0179] <<Training of the Interest Interval Detection Model 100>> The interest interval detection model 100 of Embodiment 2 is trained using training samples 110, similar to the interest interval detection model 100 of Embodiment 1. Figure 18 is a diagram illustrating the training method for the interest interval detection model 100.
[0180] First, the training device generates a first related feature sequence 80 from the training query 112 and the training image sequence 114, following the same flow as shown in Figure 9.
[0181] Furthermore, the training device generates a third-related feature sequence 200 from the training query 112 and the training image sequence 114 in the following sequence: The training device generates an image feature sequence 160 from the training image sequence 114. The training device generates a third-connected data 190 by concatenating the text features 30 and the image feature sequence 160. The training device generates a third-related feature sequence 200 by inputting the third-connected data 190 into the third-related feature sequence generation model 210.
[0182] The training device generates a fourth concatenated data 220 for each object by concatenating the first related feature sequence 80 and the third related feature sequence 200 for that object. The training device detects an interest interval 52 by inputting the fourth concatenated data 220 into the interest interval detection model 100.
[0183] The training device calculates a loss from the detected interest interval 52 and ground truth interval 116, and uses this loss to update the parameters of the first related feature sequence generation model 90, the interest interval detection model 100, and the third concatenated data 190.
[0184] <<Training of the Interest Object Determination Model 120>> The interest object determination model 120 of Embodiment 2 is trained in the same manner as the interest interval detection model 100 of Embodiment 2. Figure 19 is a diagram illustrating the training method for the interest object determination model 120.
[0185] In the example shown in Figure 9, the training device generates fourth-connected data 220 for each object from the training sample 110 in a flow similar to that shown in Figure 18. The training device inputs the fourth-connected data 220 for each object into the object of interest determination model 120 to calculate the probability that each object is an object of interest 54. The training device calculates a loss using the probability of each object being an object of interest 54 and the ground truth flag 118 for each object. The training device then updates the parameters of the first related feature sequence generation model 90, the object of interest determination model 120, and the third related feature sequence generation model 210 using the calculated loss.
[0186] <Output by the interest information detection device 2000> The detection unit 2080 of Embodiment 2 outputs output information in the same manner as the detection unit 2080 of Embodiment 1. However, if the interest interval detection model 100 detects multiple interest intervals 52, the detection unit 2080 may narrow down the interest intervals 52 to be included in the output information to a portion of the multiple detected interest intervals 52.
[0187] For example, the detection unit 2080 calculates a score (hereinafter referred to as the interval score) that represents the degree to which each interest interval 52 is related to the query text data 10. Then, the detection unit 2080 uses the interval scores of each interest interval 52 to determine which interest intervals 52 should be included in the output information. For example, the detection unit 2080 includes the top predetermined number of interest intervals 52 in the output information in descending order of interval scores. In addition, for example, the detection unit 2080 also includes interest intervals 52 in the output information whose interval scores are equal to or greater than a threshold.
[0188] To calculate the interval score, for example, the detection unit 2080 uses the fourth linked data 220 to calculate a score (relevance score) for each video frame 22 that represents the degree to which that video frame 22 is related to the query text data 10. Then, the detection unit 2080 calculates a statistical value of the relevance scores calculated for multiple video frames 22 included in the interest interval 52 as the interval score for that interest interval 52. Here, the statistical value is, for example, the mean, the maximum, or the minimum.
[0189] The association score is calculated using a machine learning model, such as a neural network. This machine learning model is also called an association score calculation model. The association score calculation model outputs an association score for each video frame 22 in response to the input of the fourth linked data 220. The detection unit 2080 then inputs the fourth linked data 220 into the association score calculation model to obtain the association score for each video frame 22.
[0190] For example, the configuration of the Head disclosed in Non-Patent Document 1 can be used for the configuration of the association score calculation model. However, the association score calculation model can be composed of any machine learning model, and its configuration is not limited to the Head configuration described above.
[0191] The association score calculation model is trained in the same way as the interest interval detection model 100 and the interest object determination model 120. Figure 20 illustrates the training method for the association score calculation model. In the example in Figure 20, the training sample 110 shows the ground truth score 119 for each image (hereinafter referred to as training image) that makes up the training image sequence 114. The ground truth score 119 represents the degree to which the corresponding image is related to the training query 112.
[0192] The training device generates fourth-connected data 220 for each object from the training query 112 and the training image sequence 114, similar to how it trains the interest interval detection model 100. The training device then inputs the fourth-connected data 220 for each object into the association score calculation model 270 to obtain the association score for each training image.
[0193] The training device calculates a loss using the association score obtained for each training image and the ground truth score 119 for each training image. Then, the training device updates the parameters of the first association feature sequence generation model 90, the third association feature sequence generation model 210, and the association score calculation model 270 using the calculated loss.
[0194] <Modification> As described above, the detection unit 2080 may detect the interest information 50 without using the first related feature sequence 80. Similarly, the detection unit 2080 may detect the interest information 50 without using the third related feature sequence 200.
[0195] Figure 21 illustrates a method for detecting interest information 50. The detection unit 2080 generates second concatenated data 150 for each object from the text features 30 and the object feature sequence 40 of each object, using the same method as described with reference to Figure 11. The detection unit 2080 generates fifth concatenated data 250 from the text features 30 and the image feature sequence 160 in a similar manner.
[0196] To generate the fifth linked data 250, the detection unit 2080 first generates the fourth related feature sequence 230 from the image feature sequence 160. The fourth related feature sequence 230 is a feature sequence in which the contextual features of the image feature sequence 160 are embedded.
[0197] The fourth related feature sequence 230 is generated, for example, using the fourth related feature generation model 240. The detection unit 2080 generates the fourth related feature sequence 230 by inputting the image feature sequence 160 into the fourth related feature generation model 240.
[0198] The fourth related feature generation model 240 is a machine learning model such as a neural network. The configuration of the fourth related feature generation model 240 can be the same as that of the first related feature sequence generation model 90.
[0199] Furthermore, the detection unit 2080 generates fifth concatenated data 250 by concatenating the text features 30 and the fourth related features sequence 230. The method for generating fifth concatenated data 250 by concatenating the text features 30 and the fourth related features sequence 230 can be the same as the method for generating first concatenated data 70 by concatenating the text features 30 and the object features sequence 40.
[0200] The detection unit 2080 generates a sixth linked data 260 for each object by linking the second linked data 150 and the fifth linked data 250 for that object. The method for generating the sixth linked data 260 by linking the second linked data 150 and the fifth linked data 250 can be the same as the method for generating the first linked data 70 by linking the text feature quantity 30 and the object feature quantity sequence 40.
[0201] The detection unit 2080 detects interest information 50 using sixth linked data 260 generated for each object. When an interest interval 52 is detected using the sixth linked data 260, the interest interval detection model 100 is configured to output the interest interval 52 in response to the input of the sixth linked data 260. Furthermore, when an object of interest 54 is detected using the sixth linked data 260, the object of interest determination model 120 is configured to output a determination result indicating whether or not the object corresponding to the sixth linked data 260 is an object of interest in response to the input of the sixth linked data 260.
[0202] The process of generating the second related feature sequence 130 and the fourth related feature sequence 230 from the video data 20 may be performed in advance. That is, the interest information detection device 2000 may acquire video data 20 that can be compared with the query text data 10 in advance, and generate the second related feature sequence 130 and the fourth related feature sequence 230 from the acquired video data 20. By generating the second related feature sequence 130 and the fourth related feature sequence 230 in advance, the time and computing resources required to detect the interest information 50 can be reduced.
[0203] If the second related feature sequence 130 and the fourth related feature sequence 230 are generated in advance, the interest information detection device 2000 stores the second related feature sequence 130 and the fourth related feature sequence 230 generated from the video data 20 in the storage unit, associating them with the identifier of the video data 20. When the detection of interest information 50 is performed, the interest information detection device 2000 receives a designation of the video data 20 and retrieves the second related feature sequence 130 and the fourth related feature sequence 230 corresponding to the designated video data 20 from the storage unit.
[0204] The method for training the interest interval detection model 100 to output an interest interval 52 in response to the input of the sixth linked data 260 can be the same as the training method for the interest interval detection model 100 described above.
[0205] Specifically, the training device generates sixth-connected data 260 for each object using the training query 112 and the training image sequence 114. The training device detects interest intervals 52 by inputting the sixth-connected data 260 for each object into the interest interval detection model 100. The training device calculates a loss using the interest intervals 52 and the ground truth interval 116, and uses the calculated loss to update the parameters of the interest interval detection model 100, the second-related feature sequence generation model 140, and the fourth-related feature generation model 240.
[0206] The method for training the interest object determination model 120 to output an interest object 54 in response to the input of the sixth linked data 260 can be the same as the training method for the interest object determination model 120 described above.
[0207] Specifically, the training device generates sixth-connected data 260 for each object using the training query 112 and the training image sequence 114. The training device inputs the sixth-connected data 260 for each object into the object of interest determination model 120 to calculate the probability that each object is an object of interest 54. The training device calculates a loss using the probability that each object is an object of interest 54 and the ground truth flag 118 for each object, and uses the calculated loss to update the parameters of the object of interest determination model 120, the second related feature sequence generation model 140, and the fourth related feature generation model 240.
[0208] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure can be made as can be understood by those skilled in the art within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0209] For example, the concatenation of data is not limited to that shown in the embodiments described above. For example, the interest information detection device 2000 may generate a feature sequence from data obtained by concatenating the object feature sequence 40 and the image feature sequence 160. That is, the interest information detection device 2000 may generate a feature sequence from data obtained by concatenating the object feature sequence 40 and the image feature sequence 160, and use the data obtained by concatenating this feature sequence with the text feature 30 to detect interest information 50.
[0210] Each drawing is merely illustrative to illustrate one or more embodiments. Each drawing may be associated with one or more other embodiments, rather than being associated with only one specific embodiment. As those skilled in the art will understand, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings, for example, to create embodiments not explicitly shown or described. Not all features or steps shown in any one drawing to illustrate an exemplary embodiment are necessarily required, and some features or steps may be omitted. The order of steps described in any of the drawings may be changed as appropriate.
[0211] Some or all of the above embodiments may also be described as follows, but are not limited to the following: (Note 1) An interest information detection device comprising: acquisition means for acquiring query text data and video data; calculation means for calculating text features from the query text data; first generation means for generating an object feature sequence, which is a sequence of object features, from the video data for each object; and detection means for detecting interest information from the video data using the text features and the object feature sequence for each of the objects, wherein the interest information includes at least an interest interval, which is an interval that matches the content of the query text data, or an interest object, which is an object that matches the content of the query text data. (Note 2) The interest information detection device according to Note 1, wherein the detection means generates a first related feature sequence, which is a sequence of features calculated from data obtained by concatenating the object feature sequence and the text features of each of the objects, and detects the interest information using the first related feature sequence corresponding to each of the objects. (Note 3) The interest information detection device according to Note 2, wherein the first related feature sequence is composed of features in which the relationship features of multiple data that constitute data formed by concatenating the object feature sequence of the object and the text features of the object are embedded. (Note 4) The interest information detection device according to Note 1, wherein the detection means generates a second related feature sequence, which is a sequence of features calculated from the object feature sequence of the object for each object, and detects the interest information using data formed by concatenating the second related feature sequence and the text features. (Note 5) The interest information detection device according to Note 1, further comprising a second generation means for generating an image feature sequence, which is a sequence of features for the entire image of each of the multiple video frames included in the video data, wherein the detection means detects the interest information from the video data using the text features, the object feature sequence of each object, and the image feature sequence.(Note 6) The interest information detection device according to Note 5, wherein the detection means generates a first related feature sequence, which is a sequence of features calculated from data obtained by concatenating the object feature sequence and the text features of each object, generates a third related feature sequence, which is a sequence of features calculated from data obtained by concatenating the image feature sequence and the text features, and detects the interest information from the video data using data obtained by concatenating the first related feature sequence and the third related feature sequence. (Note 7) The interest information detection device according to Note 6, wherein the third related feature sequence is composed of features in which the relationship features of a plurality of features constituting the image feature sequence are embedded. (Note 8) The interest information detection device according to Note 5, wherein the detection means generates a second related feature sequence, which is a sequence of features calculated from the object feature sequence of the object, for each object, generates a fourth related feature sequence, which is a sequence of features calculated from the image feature sequence, and detects the interest information from the video data using data obtained by concatenating the text features and the second related feature sequence, and data obtained by concatenating the text features and the fourth related feature sequence. (Note 9) A computer-based method for detecting interest information, comprising: an acquisition step of acquiring query text data and video data; a calculation step of calculating text features from the query text data; a first generation step of generating an object feature sequence, which is a sequence of object features, for each object from the video data; and a detection step of detecting interest information from the video data using the text features and the object feature sequence for each object, wherein the interest information includes at least an interest interval, which is an interval that matches the content of the query text data, or an interest object, which is an object that matches the content of the query text data.(Note 10) A program that causes a computer to perform the following steps: an acquisition step of acquiring query text data and video data; a calculation step of calculating text features from the query text data; a first generation step of generating an object feature sequence, which is a sequence of object features, for each object from the video data; and a detection step of detecting interest information from the video data using the text features and the object feature sequence for each object, wherein the interest information includes at least an interest interval, which is an interval that matches the content of the query text data, or an interest object, which is an object that matches the content of the query text data.
[0212] Some or all of the elements (e.g., configurations and functions) described in Appendices 2 through 8 that are dependent on Appendice 1 may also be dependent on Appendices 9 and 10, respectively, in the same way as the dependencies between Appendices 2 through 8. Some or all of the elements described in any appendice may be applied to various hardware, software, recording means, systems, and methods for recording software.
[0213] This application claims priority based on Japanese Patent Application No. 2024-196833, filed on 11 November 2024, and incorporates all of its disclosures herein.
[0214] 10 Query text data 20 Video data 22 Video frame 30 Text features 40 Object feature sequence 42 Object features 50 Interest information 52 Interest interval 54 Interest object 60 Object tracking information 62 Identifier 64 History 66 Time point 68 Object region information 70 First concatenated data 80 First related feature sequence 90 First related feature sequence generation model 100 Interest interval detection model 110 Training sample 112 Training query 114 Training image sequence 116 Ground truth interval 118 Ground truth flag 119 Ground truth score 120 Interest object determination model 130 Second related feature sequence 140 Second related feature sequence generation model 150 Second concatenated data 160 Image feature sequence 162 Image features 190 Third concatenated data 200 Third related feature sequence 210 Third related feature sequence generation model 220 Fourth connected data 230 Fourth related feature sequence 240 Fourth related feature generation model 250 Fifth connected data 260 Sixth connected data 270 Related score calculation model 1000 Computer 1020 Bus 1040 Processor 1060 Memory 1080 Storage device 1100 Input / output interface 1120 Network interface 2000 Interest information detection device 2020 Acquisition unit 2040 Calculation unit 2060 First generation unit 2080 Detection unit 2100 Second generation unit
Claims
1. An interest information detection device comprising: acquisition means for acquiring query text data and video data; calculation means for calculating text features from the query text data; first generation means for generating an object feature sequence, which is a sequence of object features, for each object from the video data; and detection means for detecting interest information from the video data using the text features and the object feature sequence for each object, wherein the interest information includes at least an interest interval, which is an interval that matches the content of the query text data, or an interest object, which is an object that matches the content of the query text data.
2. The interest information detection device according to claim 1, wherein the detection means generates a first related feature sequence for each object, which is a sequence of features calculated from data obtained by concatenating the object feature sequence and the text feature of the object, and detects the interest information using the first related feature sequence corresponding to each object.
3. The interest information detection device according to claim 2, wherein the first related feature sequence is composed of feature quantities in which the relationship features of a plurality of data constituting data obtained by linking the object feature sequence of the object and the text feature quantities are embedded.
4. The interest information detection device according to claim 1, wherein the detection means generates a second related feature sequence, which is a sequence of features calculated from the object feature sequence of the object, for each object, and detects the interest information using data obtained by concatenating the second related feature sequence and the text features.
5. The interest information detection device according to claim 1, further comprising a second generation means for generating an image feature sequence which is a sequence of feature quantities for the entire image of each of a plurality of video frames contained in the video data, wherein the detection means detects the interest information from the video data using the text feature quantities, the object feature sequence for each of the objects, and the image feature sequence.
6. The interest information detection device according to claim 5, wherein the detection means generates a first related feature sequence for each object, which is a sequence of features calculated from data obtained by concatenating the object feature sequence and the text features of the object; generates a third related feature sequence, which is a sequence of features calculated from data obtained by concatenating the image feature sequence and the text features; and detects the interest information from the video data using data obtained by concatenating the first related feature sequence and the third related feature sequence.
7. The interest information detection device according to claim 6, wherein the third related feature sequence is composed of feature quantities in which the relationship features of a plurality of feature quantities constituting the image feature sequence are embedded.
8. The interest information detection device according to claim 5, wherein the detection means generates a second related feature sequence, which is a sequence of features calculated from the object feature sequence of the object, for each object, generates a fourth related feature sequence, which is a sequence of features calculated from the image feature sequence, and detects the interest information from the video data using data obtained by concatenating the text features and the second related feature sequence, and data obtained by concatenating the text features and the fourth related feature sequence.
9. A computer-based method for detecting interest information, comprising: an acquisition step of acquiring query text data and video data; a calculation step of calculating text features from the query text data; a first generation step of generating an object feature sequence, which is a sequence of object features, for each object from the video data; and a detection step of detecting interest information from the video data using the text features and the object feature sequence for each object, wherein the interest information includes at least an interest interval, which is an interval that matches the content of the query text data, or an interest object, which is an object that matches the content of the query text data.
10. A program that causes a computer to perform the following steps: an acquisition step of acquiring query text data and video data; a calculation step of calculating text features from the query text data; a first generation step of generating an object feature sequence, which is a sequence of object features, for each object from the video data; and a detection step of detecting interest information from the video data using the text features and the object feature sequence for each object, wherein the interest information includes at least an interest interval, which is an interval that matches the content of the query text data, or an interest object, which is an object that matches the content of the query text data.