Information processing unit, method and program

The information processing device enhances video-based estimation by using temporal and spatial encoders with cross-attention to generate precise text information, addressing the limitations of conventional methods in representing global and local video elements.

JP2025161193AActive Publication Date: 2025-10-24TOYOTA JIDOSHA KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024064176
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-10-24
Estimated Expiration
2044-04-11

AI Technical Summary

Technical Problem

Conventional video-based estimation techniques struggle with insufficient accuracy in verbalizing both global and local information, failing to accurately represent positions, actions, and gazes of individuals, and struggle with time-series motion changes and geometric information.

Method used

An information processing device employs a time direction encoder and a spatial direction encoder to extract temporal and spatial features from video, followed by a cross-attention operation and input into a language processing model to generate text information, considering both global and local aspects simultaneously.

Benefits of technology

The approach enables accurate and fine-grained verbalization of video content, including positions, behaviors, and gazes, ensuring chronological continuity and improving the reliability of video-based estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025161193000001_ABST
    Figure 2025161193000001_ABST
Patent Text Reader

Abstract

To improve estimation technique based upon a video.SOLUTION: There is provided an information processing unit that comprises a control part, which is configured to: input a video in a first predetermined period to a time-direction encoder to acquire a first time feature quantity; extract instance information from the video; input the instance information to a space direction encoder to acquire a space feature quantity; execute cross-attention operation, based upon language information in a second predetermined period, on the first time feature quantity to acquire a second time feature quantity; and input the second time feature quantity and the space feature quantity to a language processing model and output text information corresponding to the video in the first predetermined period.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, method, and program for generating text information based on video. [Background technology]

[0002] Conventionally, there is known an estimation technique for estimating the content of a video based on the video and generating text information. For example, Non-Patent Document 1 discloses a technique for estimating the content of a video and verbalizing it into text information. [Prior art documents] [Patent documents]

[0003] [Non-Patent Document 1] Muhammad Maaz et al., “Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models”, arXiv (2023) Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional techniques allow for either global (global) or local (local) verbalization of an entire scene. However, the accuracy of the output information is insufficient. For example, conventional techniques cannot simultaneously verbalize global and local information. As a result, conventional techniques cannot accurately verbalize the positions, actions, and gazes of people, objects, etc. contained in video, or verbalize these at a fine-grained level. Furthermore, conventional techniques have difficulty accurately verbalizing time-series motion changes and geometric information. As such, there is room for improvement in video-based estimation techniques.

[0005] In view of the above circumstances, an object of the present disclosure is to improve video-based estimation techniques. [Means for solving the problem]

[0006] An information processing device according to an embodiment of the present disclosure includes: An information processing device including a control unit, The control unit inputting video of a first predetermined period into a time direction encoder to obtain a first temporal feature; extracting instance information from the video, inputting the instance information to a spatial direction encoder to obtain spatial features; performing a cross-attention operation based on linguistic information for a second predetermined period on the first temporal feature to obtain a second temporal feature; The second temporal feature amount and the spatial feature amount are input to a language processing model, and text information corresponding to the video of the first predetermined period is output.

[0007] According to one embodiment of the present disclosure, a method comprises: A method executed by an information processing device, inputting video of a first predetermined period into a time direction encoder to acquire a first temporal feature; extracting instance information from the video, and inputting the instance information into a spatial direction encoder to obtain spatial features; performing a cross-attention operation based on linguistic information for a second predetermined period on the first temporal feature to obtain a second temporal feature; inputting the second temporal feature amount and the spatial feature amount into a language processing model and outputting text information corresponding to the video of the first predetermined period; Includes.

[0008] A program according to an embodiment of the present disclosure includes: On the computer, inputting video of a first predetermined period into a time direction encoder to acquire a first temporal feature; extracting instance information from the video, and inputting the instance information into a spatial direction encoder to obtain spatial features; performing a cross-attention operation based on linguistic information for a second predetermined period on the first temporal feature to obtain a second temporal feature; inputting the second temporal feature amount and the spatial feature amount into a language processing model and outputting text information corresponding to the video of the first predetermined period; Execute the following. [Effects of the Invention]

[0009] According to one embodiment of the present disclosure, video-based estimation techniques are improved. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a diagram illustrating a schematic configuration of an information processing device according to an embodiment of the present disclosure. [Figure 2] 10 is a flowchart showing the operation of the information processing device in a learning process. [Figure 3] FIG. 1 is a schematic diagram of functional blocks in a learning process. [Figure 4] FIG. 10 is a diagram illustrating an outline of a process for extracting geometry information from text information. [Figure 5] 10 is a flowchart showing the operation of the information processing device in an estimation process. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present disclosure will be described.

[0012] (Outline of the embodiment) First, an overview of this embodiment will be described with reference to Fig. 1. An information processing device 10 inputs video of a first predetermined period into a time-direction encoder to acquire first temporal features. The information processing device also extracts instance information from the video and inputs the instance information into a spatial-direction encoder to acquire spatial features. The information processing device 10 performs a cross-attention operation on the first temporal features based on linguistic information of a second predetermined period to acquire second temporal features. The information processing device 10 then inputs the second temporal features and spatial features into a language processing model and outputs text information corresponding to the video of the first predetermined period.

[0013] As described above, according to this embodiment, a time direction encoder and a spatial direction encoder are used to acquire features based on video and instance information, respectively, and these features are used to output text information corresponding to the video using a language processing model. This improves video-based estimation technology in that it is possible to output text information that takes video and instance information into consideration, in other words, to perform estimation that takes into consideration the entire video (global) and instance information (local) simultaneously.

[0014] (Configuration of information processing device) Next, a detailed description will be given of each component of the information processing device 10. As shown in Fig. 1, the information processing device 10 includes a control unit 11, a storage unit 12, an input unit 13, an output unit 14, and a communication unit 15.

[0015] The control unit 11 includes at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a central processing unit (CPU) or a graphics processing unit (GPU), or a dedicated processor specialized for a specific process. The dedicated circuit is, for example, a field-programmable gate array (FPGA) or an application specific integrated circuit (ASIC). The control unit 11 executes processes related to the operation of the information processing device 10 while controlling each unit of the information processing device 10.

[0016] The storage unit 12 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these. The semiconductor memory is, for example, a random access memory (RAM) or a read only memory (ROM). The RAM is, for example, a static random access memory (SRAM) or a dynamic random access memory (DRAM). The ROM is, for example, an electrically erasable programmable read only memory (EEPROM). The storage unit 12 functions as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 12 stores data used in the operation of the information processing device 10 and data obtained by the operation of the information processing device 10.

[0017] The input unit 13 includes at least one input interface. The input interface is, for example, a physical key, a capacitance key, a pointing device, or a touch screen integrated with a display. The input interface may also be, for example, a sound sensor that accepts voice input, or a camera that accepts gesture input. The input unit 13 accepts an operation to input data used for the operation of the information processing device 10. The input unit 13 may be connected to the information processing device 10 as an external input device instead of being provided in the information processing device 10. Any connection method may be used, for example, a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI) (registered trademark), or Bluetooth (registered trademark).

[0018] The output unit 14 includes at least one output interface. The output interface is, for example, a display that outputs information as a video, or a speaker that outputs information as a sound. The display is, for example, an LCD (liquid crystal display) or an organic EL (electro luminescence) display. The output unit 14 displays and outputs data obtained by the operation of the information processing device 10. The output unit 14 may be connected to the information processing device 10 as an external output device instead of being provided in the information processing device 10. Any connection method can be used, for example, USB, HDMI (registered trademark), or Bluetooth (registered trademark).

[0019] The communication unit 15 includes at least one external communication interface. The communication interface may be either a wired communication interface or a wireless communication interface. In the case of wired communication, the communication interface is, for example, a LAN (Local Area Network) interface or a USB (Universal Serial Bus). In the case of wireless communication, the communication interface is, for example, an interface compatible with mobile communication standards such as LTE (Long Term Evolution), 4G (4th generation), or 5G (5th generation), or an interface compatible with short-range wireless communication such as Bluetooth (registered trademark). The communication unit 15 receives data used in the operation of the information processing device 10 and transmits data obtained by the operation of the information processing device 10.

[0020] The functions of the information processing device 10 are realized by executing a program according to this embodiment on a processor corresponding to the information processing device 10. That is, the functions of the information processing device 10 are realized by software. The program causes a computer to execute the operations of the information processing device 10, thereby causing the computer to function as the information processing device 10. That is, the computer functions as the information processing device 10 by executing the operations of the information processing device 10 in accordance with the program.

[0021] In this embodiment, the program can be recorded on a computer-readable recording medium. The computer-readable recording medium includes non-transitory computer-readable media, such as a magnetic recording device, an optical disc, a magneto-optical recording medium, or a semiconductor memory. The program can be distributed, for example, by selling, transferring, or lending a portable recording medium, such as a DVD (digital versatile disc) or a CD-ROM (compact disc read only memory), on which the program is recorded. The program can also be distributed by storing the program in the storage of an external server and transmitting the program from the external server to another computer. The program can also be provided as a program product.

[0022] Some or all of the functions of the information processing device 10 may be implemented by a dedicated circuit equivalent to the control unit 11. In other words, some or all of the functions of the information processing device 10 may be implemented by hardware.

[0023] (Operation of information processing device) The operation of the information processing device 10 according to this embodiment in the learning process will be described with reference to FIGS. 2, 3, and 4. FIG.

[0024] Step S11: The control unit 11 of the information processing device 10 inputs a learning video of a first predetermined period into a time direction encoder to acquire a first temporal feature. In this embodiment, the first predetermined period is assumed to be a period from t-Δt to t. Δt may be, for example, several seconds. For example, Δt may be 3 seconds.

[0025] 3 shows a schematic diagram of the functional blocks of the learning process. The functional blocks of the learning process in this embodiment include a temporal visual encoder 101, a spatial visual encoder 103, a cross-attention processor 104 (Q-Former / cross-attention), a language processing model 105 (LLM: Large Language Model), an extraction unit 106 (Spatial Semantic Extractor), a text encoder 107, a text encoder 108, and a memory bank 110. In this embodiment, programs, data, etc. required for these functional blocks are stored in, for example, a storage unit 12, and the control unit 11 accesses the storage unit 12 to execute processing related to these functions.

[0026] The control unit 11 inputs a learning video 100 from a period t-Δt to t to a time direction encoder 101. The time direction encoder 101 outputs a first temporal feature corresponding to the input. Specifically, the time direction encoder 101 analyzes temporal changes between consecutive image frames of the video, encodes data based on the changes, and outputs a first temporal feature. The control unit 11 acquires the first temporal feature.

[0027] Step S12: The control unit 11 extracts instance information from the learning video for a first predetermined period and inputs the instance information to the spatial direction encoder to acquire spatial features. The instance information is information obtained by cropping an object to be verbalized from the video. The object includes any object. For example, the object includes a person, a vehicle, etc. As shown in FIG. 3, the control unit 11 extracts instance information 102 from the learning video 100 for the period from t-Δt to t. Any method can be used to extract instance information from video. In the example of FIG. 3, the objects to be verbalized are a woman and a black car. Note that in the example of FIG. 3, the number of objects to be verbalized is two, but this is not limited thereto. The number of objects to be verbalized may be one, three, or more. The control unit 11 inputs the instance information and necessary information to the spatial direction encoder 103. The necessary information here is mask information, etc. The mask information is information that represents the positional relationship, etc., of the instance information. To obtain the mask information, the control unit 11 may input, for example, the video 100 to the spatial direction encoder 103. The spatial direction encoder 103 outputs a spatial feature corresponding to the input. Specifically, the spatial direction encoder 103 analyzes spatial information in an image frame and outputs the spatial feature. More specifically, the spatial direction encoder 103 outputs the spatial feature by utilizing similarities, patterns, etc. of pixel values ​​of the input information. The control unit 11 acquires the spatial feature.

[0028] Step S13: The control unit 11 performs a cross-attention operation on the first temporal feature based on linguistic information for a second predetermined period to acquire a second temporal feature. The second predetermined period is a period prior to the first predetermined period. In this embodiment, the second predetermined period is assumed to be the period from t-Δ2t to t-Δt. As shown in FIG. 3, the control unit 11 inputs the first temporal feature output by the time direction encoder 101 to the cross-attention processing unit 104. The cross-attention processing unit 104 performs a cross-attention operation on the first temporal feature to output a second temporal feature. Specifically, the cross-attention processing unit 104 receives two different data sets (i.e., the first temporal feature and linguistic information for the second predetermined period) as input, calculates the degree of influence of elements of one data set on elements of the other data set, performs a cross-attention operation on the first temporal feature, and outputs a second temporal feature. The control unit 11 acquires the second temporal feature.

[0029] Here, the cross-attention processing unit 104 performs a cross-attention operation based on the linguistic information for the second predetermined period, as described above. The linguistic information for the second predetermined period is information generated by the language processing model 105, the extraction unit 106, and the text encoder 107. The language processing model 105 outputs text information for the second predetermined period based on an input corresponding to the second predetermined period. The extraction unit 106 extracts geometry information from the text information for the second predetermined period. Geometry information is information about the position, orientation, etc. of an object in the text information. Figure 4 shows an overview of the process of extracting geometry information from text information. The target video in Figure 4 is video 120. Text information 121 is text information corresponding to a predetermined period of this video 120. Specifically, the text information 121 is information corresponding to the video for the period from the start time (35.221 seconds) to the end time (37.223 seconds) of the entire duration of the video (1 minute 17 seconds). The text information 121 includes geometry information, attention information, behavior information, context information, etc. The extraction unit 106 extracts the geometry information from among these.Specifically, for example, the text information reads, ``The pedestrian, a male in his 20s, stood perpendicular to the vehicle and to the left. He was positioned diagonally to the right, in front of the vehicle, at a close distance. Slowly looking around, the pedestrian's line of sight was fixed on the vehicle. As for the environment, the weather was cloudy, and the brightness of the surroundings was dim. The road surface conditions were dry on the level asphalt road, which was classified as a residential road with two-way traffic. Notably, there were no sidewalks or roadside strips on both sides of the road, but there were street lights illuminating the area." (The pedestrian was a man in his 20s, standing perpendicular to the vehicle on the left side. He was positioned diagonally forward to the right, close to the vehicle. While slowly looking around, the pedestrian's gaze was fixed on the vehicle. He noticed the vehicle and appeared to be aware of its presence. Despite the traffic ahead of him walking in the lane, he intended to continue straight ahead.His speed was slow, which was commensurate with his careful actions. The weather was cloudy and the surroundings were dimly lit. The road surface was a flat, dry asphalt road and was classified as a residential road with two-way traffic. There were no sidewalks or shoulders on either side of the road, but streetlights lit the area.), the extraction unit 106 extracts the part "The pedestrian, a male in his 20s, stood perpendicular to the vehicle and to the left. He was positioned diagonally to the right, in front of the vehicle, at a close distance" as geometry information.

[0030] The text encoder 107 encodes the extracted geometry information to generate linguistic information for a second predetermined period. The cross-attention processor 104 adjusts the weighting based on the linguistic information for the second predetermined period. In other words, the cross-attention processor 104 adjusts parameters related to the cross-attention operation for the first predetermined period while taking into account the output for the previous second predetermined period.

[0031] Step S14: The control unit 11 inputs the second temporal feature and spatial feature into the language processing model to acquire text information corresponding to the video of the first predetermined period. At this time, the control unit 11 also inputs text related to the prompt question related to the language processing model into the language processing model as appropriate.

[0032] As shown in FIG. 3 , the prompt question is, for example, "Please describe the scene with the following conditions: XXX." The conditions in the prompt question can be set arbitrarily. The prompt question is input to a text encoder 108 and encoded into text in a format that can be input to the language processing model 105. The language processing model 105 outputs text information (Caption) 109 corresponding to the video 100 of the first predetermined time period based on the second temporal feature and spatial feature corresponding to the first predetermined time period and the text of the prompt question. The content of the text information 109 is, for example, "A woman is seen walking along the sidewalk and starts to cross the crossroad while a black car is turning left through the traffic lights...." The control unit 11 stores the text information corresponding to the video of each time period output by the language processing model 105 in a memory bank 110.

[0033] Step S15: The control unit 11 calculates a loss (Loss1) based on the text information corresponding to the video of the first predetermined period and the training information. Loss1 is also called a differential loss. Any method may be used in this calculation process. For example, the control unit 11 may calculate the differential loss between the text information corresponding to the video of the first predetermined period and the training information based on a predetermined loss function.

[0034] Step S16: The control unit 11 extracts geometry information from the text information corresponding to the video of the first predetermined period and projects it onto the feature space. As shown in FIG. 3, specifically, the extraction unit 106 extracts geometry information corresponding to the first predetermined period from the text information corresponding to the first predetermined period. The text encoder 107 encodes the geometry information corresponding to the first predetermined period and projects it onto the feature space. In other words, the text encoder 107 encodes the geometry information corresponding to the first predetermined period into features corresponding to that information.

[0035] Step S17: The control unit 11 calculates the distance (Loss2) between the spatial feature and the geometry information corresponding to the first predetermined period projected onto the feature space. Loss2 is also called the consistency loss. That is, the control unit 11 calculates the distance between the geometry information corresponding to the first predetermined period projected onto the feature space and the spatial feature acquired in step S12.

[0036] Step S18: The control unit 11 trains the time direction encoder and the space direction encoder based on the loss (Loss1) calculated in step S15 and the distance (Loss2) calculated in step S17. That is, the control unit 11 trains the time direction encoder and the space direction encoder so that the loss and the distance are optimized. Any method can be used for the optimization process of the loss and the distance. For example, the time direction encoder and the space direction encoder may be trained so that both the loss and the distance are minimized.

[0037] Once the time direction encoder and the spatial direction encoder are trained, the information processing device 10 can estimate text information corresponding to video based on the trained time direction encoder and the spatial direction encoder. Hereinafter, with reference to Fig. 5, the operation of the information processing device 10 according to this embodiment in the estimation process will be described.

[0038] Step S21: The control unit 11 of the information processing device 10 inputs video of a first predetermined period to a time direction encoder to acquire a first temporal feature. Specifically, the control unit 11 inputs video to be estimated from a period from t-Δt to t to the trained time direction encoder 101. The trained time direction encoder 101 outputs a first temporal feature corresponding to the input. The control unit 11 acquires the first temporal feature.

[0039] Step S22: The control unit 11 extracts instance information from the video of the first predetermined period and inputs the instance information to the trained spatial direction encoder to acquire spatial features. Specifically, the control unit 11 extracts instance information from the video of the estimation target during the period from t-Δt to t. The control unit 11 inputs the instance information and necessary information to the trained spatial direction encoder 103. The necessary information here is mask information, etc. To acquire the mask information, the control unit 11 may input, for example, the video of the estimation target to the trained spatial direction encoder 103. The trained spatial direction encoder 103 outputs spatial features corresponding to the input. The control unit 11 acquires the spatial features.

[0040] Step S23: The control unit 11 performs a cross-attention operation on the first temporal feature based on linguistic information for a second predetermined period to acquire a second temporal feature. Specifically, the control unit 11 inputs the first temporal feature output by the trained time direction encoder 101 to the cross-attention processing unit 104. The cross-attention processing unit 104 performs a cross-attention operation on the first temporal feature to output a second temporal feature. The control unit 11 acquires the second temporal feature.

[0041] Step S24: The control unit 11 inputs the second temporal feature and spatial feature into the language processing model to acquire text information corresponding to the video of the first predetermined period. At this time, the control unit 11 also inputs text related to a prompt question related to the language processing model into the language processing model as appropriate. The control unit 11 outputs the acquired text information. Any method can be used to output the information. For example, the control unit 11 may present the information through a user interface that is displayed and output by the output unit 14.

[0042] As described above, the information processing device 10 according to this embodiment uses a time direction encoder and a spatial direction encoder to acquire features based on video and instance information, respectively, and uses these features to output text information corresponding to the video using a language processing model.

[0043] This configuration allows for estimation that simultaneously considers the entire video (global) and instance information (local). Furthermore, estimation that simultaneously considers information related to time-series motion changes output by the time-direction encoder and geometric information output by the space-direction encoder is also possible. Therefore, the technology according to this embodiment makes it possible to output text information that accurately and finely granularly represents the position, behavior, gaze, etc., of a person included in the video. This fine-grained spatiotemporal verbalization technology can be used to verbalize traffic safety, purchasing behavior, and the behavior of factory workers, and can also be used for further detailed analysis. In this way, this embodiment improves video-based estimation technology.

[0044] Furthermore, according to this embodiment, the control unit 11 performs a cross-attention operation on the first temporal feature based on the linguistic information for the second predetermined period. That is, the cross-attention operation for the first predetermined period is performed while taking into account the output for the previous second predetermined period. In this way, the output for the current time is adjusted taking into account the output information for the previous time, so according to this embodiment, the chronological continuity of the output content can be ensured.

[0045] In this embodiment, the first predetermined period is the period from t-Δt to t. The second predetermined period is the period from t-Δ2t to t-Δt. In this way, the time width (Δt) of the first predetermined period and the second predetermined period is the same. Furthermore, the first predetermined period is the period immediately following the second predetermined period. By doing so, it is possible to improve the reliability of ensuring the chronological continuity of the output content.

[0046] Although the present disclosure has been described based on the drawings and examples, it should be noted that those skilled in the art may make various modifications and alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are included in the scope of the present disclosure. For example, the functions included in each component or step can be rearranged so as not to be logically inconsistent, and multiple components or steps can be combined or divided into one.

[0047] For example, in step S17, the control unit 11 calculates the distance between geometry information corresponding to a first predetermined period projected onto the feature space and the spatial feature acquired in step S12. In other words, in step S17, the control unit 11 calculates the distance between geometry information and spatial feature corresponding to video of the same period, but this is not limited to this. That is, the periods of information for which the distance is calculated in step S17 do not have to be the same. For example, the control unit 11 may calculate the distance between geometry information corresponding to a second predetermined period projected onto the feature space and the spatial feature acquired in step S17. In this case, in step S18, the control unit 11 may train the time direction encoder and the spatial direction encoder based on the distance and the loss calculated in step S15. Alternatively, the control unit 11 may calculate both the distance between geometry information corresponding to the first predetermined period projected onto the feature space and the spatial feature acquired in step S17, and the distance between geometry information corresponding to the second predetermined period projected onto the feature space and the spatial feature acquired in step S17. In this case, in step S18, the control unit 11 may learn the time direction encoder and the space direction encoder based on these two distances and the loss calculated in step S15. In this way, the distance over the previous second predetermined period can also be reflected in the learning process of the time direction encoder and the space direction encoder so as to be optimized.

[0048] Furthermore, for example, in the above-described embodiment, an example has been described in which the estimation process is performed using the time direction encoder and the space direction encoder trained by the method of steps S11 to S18, but the learning method for the time direction encoder and the space direction encoder is not limited to the method of steps S11 to S18. The information processing device 10 may perform the estimation process in steps S21 to S22 using the time direction encoder and the space direction encoder trained by any method.

[0049] Furthermore, for example, in the above-described embodiment, an embodiment is also possible in which the configuration and operation of the information processing device 10 are distributed among a plurality of other computers that can communicate with each other. In other words, the functional blocks related to the above-described time direction encoder 101, spatial direction encoder 103, cross-attention processing unit 104, language processing model 105, extraction unit 106, text encoder 107, text encoder 108, and memory bank 110 may be distributed as appropriate among the information processing device 10 and a plurality of other devices. [Explanation of symbols]

[0050] 10. Information processing equipment 11 Control section 12 Storage section 13 Input section 14 Output section 15 Communications Department 100 videos 101 Time-direction encoder 102 Instance Information 103 Spatial Direction Encoder 104 Cross-attention processing unit 105 Language Processing Model 106 Extraction part 107 Text Encoder 108 Text Encoder 109 Text Information 110 Memory Bank 120 videos 121 Text Information

Claims

1. An information processing device including a control unit, The control unit inputting video of a first predetermined period into a time direction encoder to obtain a first temporal feature; extracting instance information from the video, inputting the instance information to a spatial direction encoder to obtain spatial features; performing a cross-attention operation based on linguistic information for a second predetermined period on the first temporal feature to obtain a second temporal feature; inputting the second temporal feature amount and the spatial feature amount into a language processing model to output text information corresponding to the video of the first predetermined period; Information processing device.

2. 2. The information processing device according to claim 1, wherein the time direction encoder and the spatial direction encoder are trained based on a loss based on the text information and training information corresponding to the video in the first predetermined period, and a distance between the spatial feature and geometry information extracted from the text information and projected onto a feature space.

3. The information processing device according to claim 1 , wherein the first predetermined period and the second predetermined period have the same time duration, and the first predetermined period is a period immediately following the second predetermined period.

4. A method executed by an information processing device, inputting video of a first predetermined period into a time direction encoder to acquire a first temporal feature; extracting instance information from the video, and inputting the instance information into a spatial direction encoder to obtain spatial features; performing a cross-attention operation based on linguistic information for a second predetermined period on the first temporal feature to acquire a second temporal feature; inputting the second temporal feature amount and the spatial feature amount into a language processing model and outputting text information corresponding to the video of the first predetermined period; A method comprising:

5. On the computer, inputting video of a first predetermined period into a time direction encoder to acquire a first temporal feature; extracting instance information from the video, and inputting the instance information into a spatial direction encoder to obtain spatial features; performing a cross-attention operation based on linguistic information for a second predetermined period on the first temporal feature to acquire a second temporal feature; inputting the second temporal feature amount and the spatial feature amount into a language processing model and outputting text information corresponding to the video of the first predetermined period; A program that executes the following.

Citation Information

Patent Citations

  • Description generation device, method, and program

    JP2023053742A

  • Method, apparatus, device and medium for generating captioning information of multimedia data

    US20220014807A1

  • Systems and methods for video and language pre-training

    US20230154146A1