Information processing device, method, and program

The information processing apparatus improves video estimation by using time-direction and spatial-direction encoders to enhance the accuracy of verbalizing global and local video content, enabling precise representation of positions, actions, and gaze.

JP7893273B2Active Publication Date: 2026-07-22TOYOTA JIDOSHA KK
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2024-04-11
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Conventional video estimation technologies struggle with insufficient accuracy in verbalizing global and local information simultaneously, failing to accurately represent positions, actions, gazes, and motion changes in fine granularity.

Method used

An information processing apparatus that utilizes a time-direction encoder and spatial-direction encoder to extract time-specific and spatial features from video, followed by a cross-attention operation and language processing model to output text information, enabling simultaneous consideration of global and local video content.

Benefits of technology

Enhances the accuracy of verbalizing video content by accurately representing positions, actions, and gaze with fine granularity, improving estimation technology for applications like traffic safety and detailed analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893273000001
    Figure 0007893273000001
  • Figure 0007893273000002
    Figure 0007893273000002
  • Figure 0007893273000003
    Figure 0007893273000003
Patent Text Reader

Abstract

To improve estimation technique based upon a video.SOLUTION: There is provided an information processing unit that comprises a control part, which is configured to: input a video in a first predetermined period to a time-direction encoder to acquire a first time feature quantity; extract instance information from the video; input the instance information to a space direction encoder to acquire a space feature quantity; execute cross-attention operation, based upon language information in a second predetermined period, on the first time feature quantity to acquire a second time feature quantity; and input the second time feature quantity and the space feature quantity to a language processing model and output text information corresponding to the video in the first predetermined period.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, method, and program for generating text information based on video.

Background Art

[0002] Conventionally, an estimation technique for estimating the content of such video and generating text information based on the video is known. For example, Non-Patent Document 1 discloses a technique for estimating the content of a video and verbalizing it into text information.

Prior Art Documents

Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, verbalization can be performed either globally for the entire scene or locally for an instance. However, in the conventional technology, the accuracy of the output information was not sufficient. For example, in the conventional technology, it was not possible to verbalize global and local information simultaneously. Therefore, in the conventional technology, it was not possible to accurately verbalize the positions, actions, gazes, etc. of people, objects, etc. included in the video, or to verbalize them in fine granularity. Also, in the conventional technology, it was difficult to accurately verbalize motion changes and geometric information over time series. Thus, there was room for improvement in the estimation technology based on video.

[0005] In light of these circumstances, the purpose of this disclosure is to improve video-based estimation technology. [Means for solving the problem]

[0006] An information processing apparatus according to one embodiment of this disclosure is An information processing device comprising a control unit, The control unit, The video for the first predetermined period is input to a time-direction encoder to obtain the first time-specific feature, Instance information is extracted from the aforementioned video, and the instance information is input to a spatial directional encoder to obtain spatial features. A second time feature is obtained by performing a cross-attention operation based on language information for a second predetermined period on the first time feature, The second temporal feature and the spatial feature are input to a language processing model to output text information corresponding to the video for the first predetermined period.

[0007] A method according to one embodiment of this disclosure is: A method executed by an information processing device, The process involves inputting video footage from a predetermined period into a time-direction encoder to obtain a first time-direction feature, The process involves extracting instance information from the aforementioned video, inputting the instance information into a spatial directional encoder to obtain spatial features, A second time feature is obtained by performing a cross-attention operation based on language information for a second predetermined period on the aforementioned first time feature, The second temporal feature and the spatial feature are input to a language processing model to output text information corresponding to the video for the first predetermined period. Includes.

[0008] A program according to one embodiment of this disclosure is On the computer, The process involves inputting video footage from a predetermined period into a time-direction encoder to obtain a first time-direction feature, The process involves extracting instance information from the aforementioned video, inputting the instance information into a spatial directional encoder to obtain spatial features, A second time feature is obtained by performing a cross-attention operation based on language information for a second predetermined period on the aforementioned first time feature, The second temporal feature and the spatial feature are input to a language processing model to output text information corresponding to the video for the first predetermined period. Make it run. [Effects of the Invention]

[0009] According to one embodiment of this disclosure, the estimation technique based on video is improved. [Brief explanation of the drawing]

[0010] [Figure 1] This figure shows a schematic configuration of an information processing device according to one embodiment of the present disclosure. [Figure 2] This is a flowchart showing the operation of an information processing device during its learning process. [Figure 3] This is an overview diagram of the functional blocks in the learning process. [Figure 4] This diagram shows an overview of the process for extracting geometric information from text information. [Figure 5] This is a flowchart showing the operation of the information processing device during the estimation process. [Modes for carrying out the invention]

[0011] The embodiments of this disclosure will be described below.

[0012] (Summary of the embodiment) First, referring to FIG. 1, the outline of this embodiment will be described. The information processing apparatus 10 inputs the video of the first predetermined period into the time direction encoder to obtain the first time feature amount. Further, the information processing apparatus extracts instance information from the video, and inputs the instance information into the spatial direction encoder to obtain the spatial feature amount. The information processing apparatus 10 performs a cross-attention operation based on the language information of the second predetermined period on the first time feature amount to obtain the second time feature amount. Then, the information processing apparatus 10 inputs the second time feature amount and the spatial feature amount into the language processing model to output the text information corresponding to the video of the first predetermined period.

[0013] As described above, according to this embodiment, by using the time direction encoder and the spatial direction encoder, feature amounts are obtained based on the video and instance information respectively, and the text information corresponding to the video is output by the language processing model using such feature amounts. Therefore, the estimation technology based on the video is improved in that text information considering the video and instance information can be output, that is, in that estimation considering the entire video (global) and instance information (local) simultaneously can be performed.

[0014] (Configuration of Information Processing Apparatus) Next, each configuration of the information processing apparatus 10 will be described in detail. As shown in FIG. 1, the information processing apparatus 10 includes a control unit 11, a storage unit 12, an input unit 13, an output unit 14, and a communication unit 15.

[0015] The control unit 11 includes at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU (central processing unit) or a GPU (graphics processing unit), or a dedicated processor specialized for specific processing. The dedicated circuit is, for example, an FPGA (field-programmable gate array) or an ASIC (application specific integrated circuit). The control unit 11 executes processes related to the operation of the information processing device 10 while controlling each part of the information processing device 10.

[0016] The storage unit 12 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these. The semiconductor memory is, for example, a RAM (random access memory) or a ROM (read only memory). The RAM is, for example, a SRAM (static random access memory) or a DRAM (dynamic random access memory). The ROM is, for example, an EEPROM (electrically erasable programmable read only memory). The storage unit 12 functions as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 12 stores data used for the operation of the information processing device 10 and data obtained by the operation of the information processing device 10.

[0017] The input unit 13 includes at least one input interface. The input interface may be, for example, a physical key, a capacitive key, a pointing device, or a touchscreen integrated with a display. Alternatively, the input interface may be, for example, a sound sensor that accepts voice input, or a camera that accepts gesture input. The input unit 13 accepts operations to input data used for the operation of the information processing device 10. Instead of being integrated into the information processing device 10, the input unit 13 may be connected to the information processing device 10 as an external input device. Any connection method can be used, for example, USB (Universal Serial Bus), HDMI (registered trademark) (High-Definition Multimedia Interface), or Bluetooth (registered trademark).

[0018] The output unit 14 includes at least one output interface. The output interface is, for example, a display that outputs information as video, or a speaker that outputs information as sound. The display is, for example, an LCD (liquid crystal display) or an organic EL (electroluminescence) display. The output unit 14 displays and outputs data obtained by the operation of the information processing device 10. Instead of being provided in the information processing device 10, the output unit 14 may be connected to the information processing device 10 as an external output device. Any connection method can be used, for example, USB, HDMI (registered trademark), or Bluetooth (registered trademark).

[0019] The communication unit 15 includes at least one external communication interface. The communication interface may be either a wired or wireless communication interface. In the case of wired communication, the communication interface may be, for example, a LAN (Local Area Network) interface or a USB (Universal Serial Bus) interface. In the case of wireless communication, the communication interface may be, for example, an interface compatible with mobile communication standards such as LTE (Long Term Evolution), 4G (4th generation), or 5G (5th generation), or an interface compatible with short-range wireless communication such as Bluetooth (registered trademark). The communication unit 15 receives data used for the operation of the information processing device 10 and transmits data obtained by the operation of the information processing device 10.

[0020] The functions of the information processing device 10 are realized by executing a program according to this embodiment on a processor corresponding to the information processing device 10. In other words, the functions of the information processing device 10 are realized by software. The program causes the computer to perform the operations of the information processing device 10, thereby causing the computer to function as the information processing device 10. That is, the computer functions as the information processing device 10 by performing the operations of the information processing device 10 according to the program.

[0021] In this embodiment, the program can be recorded on a computer-readable recording medium. The computer-readable recording medium includes non-temporary computer-readable media, such as magnetic recording devices, optical discs, magneto-optical recording media, or semiconductor memory. The program can be distributed, for example, by selling, transferring, or lending portable recording media such as DVDs (digital versatile discs) or CD-ROMs (compact disc read-only memory) on which the program is recorded. Alternatively, the program may be distributed by storing it on the storage of an external server and transmitting it from the external server to other computers. The program may also be provided as a program product.

[0022] Some or all of the functions of the information processing device 10 may be implemented by a dedicated circuit corresponding to the control unit 11. In other words, some or all of the functions of the information processing device 10 may be implemented by hardware.

[0023] (Operation of information processing device) Referring to Figures 2, 3, and 4, the operation of the information processing device 10 according to this embodiment during the learning process will be described.

[0024] Step S11: The control unit 11 of the information processing device 10 inputs the learning video for a first predetermined period to the time encoder to obtain the first time feature. In this embodiment, the first predetermined period is the period from t-Δt to t. Δt may be, for example, a few seconds. For example, Δt may be 3 seconds.

[0025] Figure 3 shows an overview of the functional blocks in the learning process. The functional blocks in the learning process of this embodiment include a temporal visual encoder 101, a spatial visual encoder 103, a cross-attention processing unit 104 (Q-Former / cross-attention), a Large Language Model (LLM) 105, a Spatial semantic extractor 106, a text encoder 107, a text encoder 108, and a memory bank 110. In this embodiment, the programs, data, etc., necessary for these functional blocks are stored in, for example, a storage unit 12, and the control unit 11 accesses the storage unit 12 to execute processing related to these functions.

[0026] The control unit 11 inputs the training video 100 from t-Δt to t to the time-direction encoder 101. The time-direction encoder 101 outputs a first time feature corresponding to the input. Specifically, the time-direction encoder 101 analyzes the temporal changes between consecutive image frames of the video, encodes the data based on those changes, and outputs a first time feature. The control unit 11 acquires this first time feature.

[0027] Step S12: The control unit 11 extracts instance information from the training video for a first predetermined period and inputs the instance information to the spatial encoder to obtain spatial features. Instance information is information obtained by cropping the video of objects that are to be verbalized. Objects include any object. For example, objects include people, vehicles, etc. As shown in Figure 3, the control unit 11 extracts instance information 102 from the training video 100 for the period from t-Δt to t. Any method can be used to extract instance information from the video. In the example in Figure 3, the objects to be verbalized are a woman and a black car. In the example in Figure 3, there are two objects to be verbalized, but this is not limited to this, and the number of target objects may be one, three or more. The control unit 11 inputs the instance information and necessary information to the spatial encoder 103. The necessary information here is mask information, etc. Mask information is information that represents the positional relationship of the instance information, etc. To obtain mask information, the control unit 11 may, for example, input the video 100 to the spatial encoder 103. The spatial directional encoder 103 outputs spatial features corresponding to the input. Specifically, the spatial directional encoder 103 analyzes the spatial information within the image frame and outputs spatial features. More specifically, the spatial directional encoder 103 outputs spatial features using the similarity, patterns, etc., of the pixel values ​​of the input information. The control unit 11 acquires these spatial features.

[0028] Step S13: The control unit 11 obtains the second time feature by performing a cross-attention operation on the first time feature based on language information for a second predetermined period. The second predetermined period is a period prior to the first predetermined period. In this embodiment, the second predetermined period is the period from t-Δ2t to t-Δt. As shown in Figure 3, the control unit 11 inputs the first time feature output by the time-direction encoder 101 to the cross-attention processing unit 104. The cross-attention processing unit 104 outputs the second time feature by performing a cross-attention operation on the first time feature. Specifically, the cross-attention processing unit 104 receives two different datasets as input (i.e., the first time feature and language information for a second predetermined period), calculates the degree of influence that an element of one dataset has on an element of the other dataset, performs a cross-attention operation on the first time feature, and outputs the second time feature. The control unit 11 obtains the second time feature.

[0029] Here, the cross-attention processing unit 104 performs a cross-attention operation based on the language information for the second predetermined period, as described above. The language information for the second predetermined period is information generated by the language processing model 105, the extraction unit 106, and the text encoder 107. The language processing model 105 outputs text information corresponding to the second predetermined period based on the input corresponding to the second predetermined period. The extraction unit 106 extracts geometry information from the text information corresponding to the second predetermined period. Geometry information is information related to the position, orientation, etc. of an object from the text information. Figure 4 shows an overview of the process of extracting geometry information from text information. In Figure 4, the target video is video 120. The text information 121 is text-format information corresponding to a predetermined period of this video 120. Specifically, the text information 121 is information corresponding to the video for the period from the start time (35.221 seconds) to the end time (37.223 seconds) within the entire duration of the video (1 minute 17 seconds). The text information 121 includes geometry information, attention information, behavior information, context information, etc. The extraction unit 106 extracts the geometry information from this.Specifically, for example, the text information reads, ``The pedestrian, a male in his 20s, stood perpendicular to the vehicle and to the left. He was positioned diagonally to the right, in front of the vehicle, at a close distance. Slowly looking around, the pedestrian's line of sight was fixed on the vehicle. As for the environment, the weather was cloudy, and the brightness of the surroundings was dim. The road surface conditions were dry on the level asphalt road, which was classified as a residential road with two-way traffic. Notably, there were no sidewalks or roadside strips on both sides of the road, but there were street lights illuminating the "area." (The pedestrian was a man in his 20s, standing perpendicular to the vehicle on its left side. He was positioned diagonally to the right and in front of the vehicle, quite close to it. He slowly glanced around, but his gaze was fixed on the vehicle. He appeared to have noticed the vehicle and was aware of its presence. In front of him, despite walking in the lane, he intended to continue straight ahead.)The speed was slow, befitting his cautious actions. The weather was cloudy, and the surroundings were dim. The road surface was a flat, dry asphalt road, classified as a two-way residential street. There were no sidewalks or shoulders on either side of the road, but streetlights illuminated the area.) If this is the case, the extraction unit 106 extracts the part "The pedestrian, a male in his 20s, stood perpendicular to the vehicle and to the left. He was positioned diagonally to the right, in front of the vehicle, at a close distance" as geometric information.

[0030] The text encoder 107 encodes the extracted geometry information to generate language information for a second predetermined period. The cross-attention processing unit 104 adjusts the weighting based on this language information for the second predetermined period. In other words, the cross-attention processing unit 104 adjusts the parameters for the cross-attention operation of the first predetermined period while considering the output of the previous period, the second predetermined period.

[0031] Step S14: The control unit 11 inputs the second time feature and spatial feature into the language processing model to obtain text information corresponding to the video for the first predetermined period. At this time, the control unit 11 also inputs text related to prompt questions related to the language processing model into the language processing model as appropriate.

[0032] As shown in Figure 3, the prompt question is, for example, "Please describe the scene with the following conditions: XXX". The conditions in the prompt question can be set arbitrarily. The prompt question is input to the text encoder 108 and encoded into text in a format that can be input to the language processing model 105. The language processing model 105 outputs text information (Caption) 109 corresponding to the video 100 for the first predetermined period, based on the second time and spatial features corresponding to the first predetermined period and the text of the prompt question. The content of the text information 109 is, for example, "A woman is seen walking along the sidewalk and starts to cross the crossroad while a black car is turning left through the traffic lights…". The control unit 11 stores the text information corresponding to the video for each period output by the language processing model 105 in the memory bank 110.

[0033] Step S15: The control unit 11 calculates a loss (Loss1) based on the text information and training information corresponding to the video for the first predetermined period. Loss1 is also called the differential loss. Any method may be used in this calculation process. For example, the control unit 11 may calculate the differential loss between the text information and training information corresponding to the video for the first predetermined period based on a predetermined loss function.

[0034] Step S16: The control unit 11 extracts geometry information from the text information corresponding to the video for the first predetermined period and projects it onto the feature space. Specifically, as shown in Figure 3, the extraction unit 106 extracts geometry information corresponding to the first predetermined period from the text information corresponding to the first predetermined period. The text encoder 107 encodes the geometry information corresponding to the first predetermined period and projects it onto the feature space. In other words, the text encoder 107 encodes the geometry information corresponding to the first predetermined period into features corresponding to that information.

[0035] Step S17: The control unit 11 calculates the distance (Loss2) between the geometric information corresponding to the first predetermined period, projected onto the feature space, and the spatial features. Loss2 is also called the consistency loss. In other words, the control unit 11 calculates the distance between the geometric information corresponding to the first predetermined period, projected onto the feature space, and the spatial features obtained in step S12.

[0036] Step S18: The control unit 11 learns the time-direction encoder and the spatial-direction encoder based on the loss (Loss1) calculated in step S15 and the distance (Loss2) calculated in step S17. In other words, the control unit 11 trains the time-direction encoder and the spatial-direction encoder so that the loss and distance are optimized. Any method can be used for optimizing the loss and distance. For example, the time-direction encoder and the spatial-direction encoder may be learned so that both the loss and distance are minimized.

[0037] Once the time-direction encoder and spatial-direction encoder are trained, the information processing device 10 can estimate text information corresponding to the video based on the trained time-direction encoder and spatial-direction encoder. The operation of the information processing device 10 in the estimation process according to this embodiment will be described below with reference to Figure 5.

[0038] Step S21: The control unit 11 of the information processing device 10 inputs video for a first predetermined period into the time-direction encoder to obtain first time features. Specifically, the control unit 11 inputs the video to be estimated for the period from t-Δt to t into the trained time-direction encoder 101. The trained time-direction encoder 101 outputs first time features corresponding to the input. The control unit 11 acquires these first time features.

[0039] Step S22: The control unit 11 extracts instance information from the video for a first predetermined period and inputs the instance information to a trained spatial directional encoder to obtain spatial features. Specifically, the control unit 11 extracts instance information from the video to be estimated for the period from t-Δt to t. The control unit 11 inputs the instance information and other necessary information to the trained spatial directional encoder 103. The necessary information here is mask information, etc. To obtain mask information, the control unit 11 may, for example, input the video to be estimated to the trained spatial directional encoder 103. The trained spatial directional encoder 103 outputs spatial features corresponding to the input. The control unit 11 obtains these spatial features.

[0040] Step S23: The control unit 11 performs a cross-attention operation on the first time feature based on language information for a second predetermined period to obtain the second time feature. Specifically, the control unit 11 inputs the first time feature output by the trained time-direction encoder 101 to the cross-attention processing unit 104. The cross-attention processing unit 104 performs a cross-attention operation on the first time feature and outputs the second time feature. The control unit 11 obtains the second time feature.

[0041] Step S24: The control unit 11 inputs the second time feature and spatial feature into the language processing model to obtain text information corresponding to the video for the first predetermined period. At this time, the control unit 11 also inputs text related to prompt questions related to the language processing model into the language processing model as appropriate. The control unit 11 outputs the acquired text information. Any method can be used to output the information. For example, the control unit 11 may present the information through a user interface displayed by the output unit 14.

[0042] As described above, the information processing device 10 according to this embodiment uses a time-direction encoder and a spatial-direction encoder to acquire feature quantities based on video and instance information, respectively, and uses these feature quantities to output text information corresponding to the video using a language processing model.

[0043] With this configuration, estimation can be performed by simultaneously considering the entire video (global) and instance information (local). Furthermore, estimation can be performed by simultaneously considering information related to time-series motion changes output by the time-direction encoder and geometric information output by the spatial-direction encoder. For this reason, according to the technology of this embodiment, it is possible to output text information that represents the position, actions, gaze, etc. of people included in the video with high accuracy and fine granularity. This fine-grained spatiotemporal languageization technology can be used for languageization related to traffic safety, purchasing behavior, languageization of the actions of factory workers, etc., and can also be used for further detailed analysis. In this way, the estimation technology based on video is improved according to this embodiment.

[0044] Furthermore, according to this embodiment, the control unit 11 performs a cross-attention operation on the first time feature based on the language information of the second predetermined period. In other words, the cross-attention operation for the first predetermined period is performed while considering the output of the previous period, the second predetermined period. Because the output to the current time is adjusted by considering the output information of the previous time, according to this embodiment, the continuity of the output content in the time series can be ensured.

[0045] In this embodiment, the first predetermined period is defined as the period from t-Δt to t. The second predetermined period is defined as the period from t-Δ2t to t-Δt. Thus, the time width (Δt) of the first and second predetermined periods is the same. Furthermore, the first predetermined period is the period immediately following the second predetermined period. By doing so, the certainty of ensuring the time-series continuity of the output content can be improved.

[0046] While this disclosure has been described based on the drawings and embodiments, it should be noted that those skilled in the art may make various modifications and alterations based on this disclosure. Therefore, it should be noted that these modifications and alterations are within the scope of this disclosure. For example, the functions, etc., included in each component or step can be rearranged in a logically consistent manner, and multiple components or steps can be combined into one or divided into two.

[0047] For example, in step S17, the control unit 11 calculates the distance between the geometry information corresponding to a first predetermined period projected onto the feature space and the spatial features acquired in step S12. In other words, in step S17, the control unit 11 calculates the distance between the geometry information and spatial features corresponding to the same period of video, but is not limited to this. That is, the periods of the information for which the distance is calculated in step S17 do not have to be the same. For example, the control unit 11 may calculate the distance between the geometry information corresponding to a second predetermined period projected onto the feature space and the spatial features acquired in step S17. In this case, in step S18, the control unit 11 may learn the time-direction encoder and the spatial-direction encoder based on the said distance and the loss calculated in step S15. Alternatively, the control unit 11 may calculate both the distance between the geometry information corresponding to a first predetermined period projected onto the feature space and the spatial features acquired in step S17, and the distance between the geometry information corresponding to a second predetermined period projected onto the feature space and the spatial features acquired in step S17. In this case, the control unit 11 may, in step S18, learn the time-direction encoder and the spatial-direction encoder based on these two distances and the loss calculated in step S15. By doing so, the learning process of the time-direction encoder and the spatial-direction encoder can be reflected in optimizing the distance over the previous period, the second predetermined period.

[0048] Furthermore, although the above embodiment describes an example in which estimation processing is performed using the time-direction encoder and spatial-direction encoder learned by the method of steps S11 to S18, the method of learning the time-direction encoder and spatial-direction encoder is not limited to the method of steps S11 to S18. The information processing device 10 may perform the estimation processing in steps S21 to S22 using the time-direction encoder and spatial-direction encoder learned by any method.

[0049] Furthermore, in the embodiment described above, it is also possible to distribute the configuration and operation of the information processing device 10 to multiple other computers that can communicate with each other. In other words, the functional blocks related to the time-direction encoder 101, spatial-direction encoder 103, cross-attention processing unit 104, language processing model 105, extraction unit 106, text encoder 107, text encoder 108, and memory bank 110 may be appropriately distributed to the information processing device 10 and multiple other devices. [Explanation of symbols]

[0050] 10 Information Processing Devices 11 Control Unit 12 Storage section 13 Input section 14 Output section 15 Communications Department 100 Videos 101 Time Direction Encoder 102 Instance Information 103 Spatial Direction Encoder 104 Cross-Attention Processing Unit 105 Language Processing Models 106 Extraction part 107 Text Encoder 108 Text Encoders 109 Text Information 110 Memory Bank 120 Videos 121 Text Information

Claims

1. An information processing device comprising a control unit, The control unit, The video for the first predetermined period is input to a time-direction encoder to obtain the first time-specific feature, Instance information is extracted from the aforementioned video, and the instance information is input to a spatial directional encoder to obtain spatial features. A second time feature is obtained by performing a cross-attention operation on the first time feature based on language information generated based on geometry information extracted from text information corresponding to the video of the second predetermined period, which is the period immediately preceding the first predetermined period. The second temporal feature and the spatial feature are input to a language processing model to output text information corresponding to the video for the first predetermined period. Information processing device.

2. The information processing apparatus according to claim 1, wherein the time-direction encoder and the spatial-direction encoder are learned based on the loss based on the text information and the training information corresponding to the image in a first predetermined period, and the distance between the spatial features and the geometry information extracted from the text information and projected onto the feature space.

3. The information processing apparatus according to claim 1 or 2, wherein the time widths of the first predetermined period and the second predetermined period are the same.

4. A method executed by an information processing device, The process involves inputting video footage from a predetermined period into a time-direction encoder to obtain a first time-direction feature, The process involves extracting instance information from the aforementioned video, inputting the instance information into a spatial directional encoder to obtain spatial features, A second time feature is obtained by performing a cross-attention operation on the first time feature based on language information generated based on geometry information extracted from text information corresponding to the video of the second predetermined period, which is the period immediately preceding the first predetermined period. The second temporal feature and the spatial feature are input to a language processing model to output text information corresponding to the video for the first predetermined period. A method that includes this.

5. On the computer, The process involves inputting video footage from a predetermined period into a time-direction encoder to obtain a first time-direction feature, The process involves extracting instance information from the aforementioned video, inputting the instance information into a spatial directional encoder to obtain spatial features, A second time feature is obtained by performing a cross-attention operation on the first time feature based on language information generated based on geometry information extracted from text information corresponding to the video of the second predetermined period, which is the period immediately preceding the first predetermined period. The second temporal feature and the spatial feature are input to a language processing model to output text information corresponding to the video for the first predetermined period. A program that executes the command.