Low-latency captioning system
The low-latency captioning system optimizes caption generation by training a timing detector and generator to use a fraction of video frames, achieving high-quality captions in real-time systems.
Patent Information
- Application Number
- JP2025519308
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-04
- Filing Date
- 2023-06-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-06-16
AI Technical Summary
Conventional video captioning methods are not practical for real-time monitoring and surveillance systems as they require access to all video frames, making them unsuitable for generating captions quickly and accurately.
A low-latency captioning system that trains a timing detector and caption generator to produce captions using only a small portion of video frames, mimicking a pre-trained model and optimizing output timing for high caption quality.
The system achieves 94% of the quality provided by a pre-trained transformer using only 28% of the frames, enabling quick and accurate event description in real-time systems.
Smart Images

Figure 2025522640000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application is a continuation - in - part of Patent Application No. 17 / 384,234, filed on July 23, 2021, the content of which is incorporated herein by reference.
[0002] The present invention generally relates to scene captioning and dialogue based on scene captioning, and more specifically, to end - to - end scene captioning that utilizes a latency optimization function.
Background Art
[0003] At any moment, countless events occurring in the real world are captured by sensors such as cameras, microphones, LiDAR, mmWave, etc. and stored as a large amount of sensor data resources. In order to effectively acquire such records, scene captioning, which has the ability to understand a scene and describe events in natural language, is an essential technology whether in an offline setting or an online setting. Furthermore, a robot can interact with a person or another robot to determine the next action based on understanding the scene through online scene captioning.
[0004] Scene captioning technology has been actively studied in the field of computer vision for video captioning. Deep recurrent neural networks (RNNs) have been applied to video captioning, and RNNs are trained to convert a sequence of image features extracted from a video clip into a sequence of words that describe the video content. The goal is to generate a video description (caption) about the objects and events within any video clip. Transformer models have become more popular in recent years as a model for video captioning compared to RNNs, due to reasons such as improving captioning performance.
[0005] Furthermore, not only image features but also audio features are extracted from video clips and used to improve caption quality, and attention-based multimodal fusion techniques have been introduced to effectively fuse image features and audio features according to the video content (U.S. Patent No. 10,417,498).
[0006] Conventional methods for video captioning are basically assumed to function in an offline manner, and since each video clip is provided before captioning, the system can access all frames of the video clip to generate captions. However, such conventional methods are not practical in real-time monitoring, surveillance systems, car navigation, and scene understanding-based dialogue systems for robots. This is because in such cases, it is essential to generate captions as soon as possible not only to accurately describe events but also to quickly discover and report events and take the next action. Therefore, low-latency captioning for online systems is required to achieve such a function. The system needs to determine the appropriate timing for generating correct captions using only the limited number of frames that the system has received so far. In addition, such a low-latency captioning function enables the robot to determine the next action as soon as possible for interacting with people and other robots. Summary of the Invention
[0007] Some embodiments of the present disclosure are based on the recognition that scene captioning is an essential technique for understanding scenes and describing events in natural language. To apply this to a real-time monitoring system, the system needs to not only accurately describe events but also generate captions as soon as possible. Low-latency captioning is required to achieve such a function, but the research field regarding such online scene captioning has not yet been achieved. This specification proposes a new method for optimizing the output timing of each caption based on the trade-off between latency and caption quality. The audio-visual transformer is trained to generate ground-truth captions using only a very small portion of all scenes and is also trained to mimic the output of a pre-trained transformer given all scenes. The CNN-based timing detector is also trained to detect the appropriate output timing, in which case the captions generated by the two transformers are close enough to each other. By using the jointly trained transformer and timing detector, captions can be generated at the initial stage of an event as soon as the event occurs or when the event is predictable.
[0008] An object of some embodiments of the present invention is to provide a system and method for end-to-end scene captioning that enables an online / offline supervision system and scene-aware interaction (dialogue) so that when the system recognizes an event, it can understand the event in natural language as soon as possible.
[0009] The present disclosure includes a low-latency scene captioning system trained to optimize the output timing for each caption based on a trade-off between latency and caption quality. The present invention can train a low-latency caption generator according to the following strategies. (1) Generate a ground truth caption using a low-latency caption generator that recognizes only a very small portion of all scenes acquired by a sensor or only a very small portion of all signals. (2) Imitate the output of a pre-trained caption generator. The output is generated using the entire scene. (3) Train a timing detector to find the best timing for outputting a caption such that the caption finally generated by the low-latency caption generator is close enough to its ground truth caption or close enough to the caption generated by the pre-trained caption generator using the entire scene. The low-latency caption generator based on (1) and (2) and the timing detector of (3) are jointly trained.
[0010] The jointly trained low-latency caption generator and timing detector can generate captions at an early stage of the scene acquired by the sensor as soon as an event occurs. Additionally, this framework can be applied to predict future events during low-latency captioning. Furthermore, by combining multimodal detection information, events can be recognized at an earlier timing triggered by the earliest cue in one of the modalities, without waiting for other cues in other modalities. In particular, by using an audio-visual transformer constructed as a low-latency caption generator according to an embodiment of the present invention, captions can be generated earlier than the timing of the visual cue based on the timing of the audio cue. Such a low-latency video captioning system using multimodal detection information can contribute not only to quickly retrieving events but also to responding to the scene earlier.
[0011] Some embodiments are based on the recognition that experiments using the ActivityNet Captions dataset have shown that a system according to the present invention can achieve 94% of the upper limit caption quality provided by a pre-trained transformer that uses only 28% of the beginning portion of the frames out of the entire video clip.
[0012] According to some embodiments of the present invention, a scene captioning system can be provided. In this case, the scene captioning system includes an interface configured to obtain a stream of signals captured by a multimodal sensor for captioning a scene, and a computer-executable scene captioning model including a multimodal sensor feature extractor, a multimodal sensor feature encoder, a timing detector, and a scene caption decoder, wherein the multimodal sensor feature encoder is shared by the timing detector and the scene caption decoder, and the scene captioning system further includes a processor associated with the memory, and the processor is configured to extract multimodal sensor features from the multimodal sensor signals by using the multimodal sensor feature extractor, encode the multimodal sensor features by using the multimodal sensor feature encoder, determine the timing for generating a scene caption by using the timing detector, the timing being determined at an initial stage of the stream of the multimodal sensor signals, and the processor is further configured to execute a step of generating the scene caption that describes an event based on the multimodal sensor features by using the scene caption decoder according to the timing.
[0013] Furthermore, some embodiments of the present invention are based on the recognition that a computer-executable training method for training a multimodal sensor feature encoder, a timing detector, and a scene caption decoder is provided. The method includes providing a training data set including a set of multimodal sensor signals and a set of ground truth scene captions, converting the multimodal sensor signals into a feature vector sequence using a feature extractor, and excluding future frames from the feature vector sequence using a future frame eliminator, where the future frame eliminator takes in a first feature vector from the feature vector sequence and removes other feature vectors to generate a future-excluded feature vector sequence. The method may further include encoding the future-excluded feature vector sequence into a hidden activation vector sequence, training a low-latency scene caption generator by calculating a loss value, calculating the loss value based on a posterior probability distribution and a ground truth caption, and training the timing detector by calculating a supervision signal based on a scene caption similarity between a scene caption for the future-excluded feature vector sequence and the ground truth caption.
[0014] Furthermore, when the scene captioning system is configured as a video captioning system, the video captioning system may include an interface configured to obtain a stream of audio-visual signals including image data and audio data, and a memory for storing a computer-executable video captioning model including an audio-visual feature extractor, an audio-visual feature encoder, a timing detector, and a video caption decoder. The audio-visual encoder is shared by the timing decoder, the timing detector, and the video caption decoder, and the video captioning system may further include a processor associated with the memory. The processor is configured to perform steps of extracting audio features and visual features from the audio-visual signal by using the audio-visual extractor, encoding the audio features and visual features from the audio-visual signal by using the audio-visual encoder, and determining a timing for generating a video caption by using the timing detector, where the timing is determined at an initial stage of the stream of audio-visual signals, and the processor is further configured to perform a step of generating the video caption that describes the audio-visual scene based on the audio features and the visual features by using the video caption decoder according to the timing.
[0015] According to some embodiments of the present invention, an artificial intelligence (AI) low-latency processing system is provided. The low-latency processing system may include a processor and a memory storing instructions. When executed by the processor, the instructions cause the low-latency processing system to collect a sequence of frames, where the sequence of frames jointly includes information distributed among at least some of the frames in the sequence of frames. The instructions further cause the execution of a timing neural network trained to identify a subsequence of an initial frame of the sequence of frames that includes at least a portion of the information indicating the information, and the execution of a decoding neural network trained to decode the information from a portion of the information in the subsequence of the frames. The timing neural network is co-trained with the decoding neural network to repeatedly identify a minimum number of sub-frames from the beginning of a training sequence of frames that includes a portion of the training information sufficient to decode the training information.
[0016] Furthermore, an embodiment of the present invention provides a computer-implemented method for an artificial intelligence (AI) low-latency processing system. The low-latency processing system includes a processor and a memory storing instructions of a method implemented by a computer to execute steps using the processor. The computer-implemented method includes a step of collecting a sequence of frames. The sequence of frames jointly includes information distributed among at least some of the frames in the sequence of frames. Further, executing a timing neural network trained to identify a subsequence of an initial frame in the sequence of frames that includes at least a part of the information indicating the information, and executing a decoding neural network trained to decode information from a part of the information in the subsequence of frames. The timing neural network is co-trained with the decoding neural network to repeatedly identify a minimum number of sub-frames from the beginning of a training sequence of frames that includes a part of the training information sufficient to decode the training information.
Brief Description of the Drawings
[0017] Embodiments of the present disclosure will be further described with reference to the accompanying drawings. The drawings shown are not necessarily to scale and generally focus on explaining the principles of the embodiments of the present disclosure.
[0018]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11A
Figure 11B
Best Mode for Carrying Out the Invention
[0019] The above drawings illustrate embodiments of the present disclosure, but other embodiments are also contemplated as described below. The present disclosure presents specific embodiments for purposes of illustration rather than limitation. Those skilled in the art will be able to devise numerous other variations and embodiments that fall within the scope and spirit of the principles of the embodiments of the present disclosure.
[0020] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments will provide those skilled in the art with a feasible description for implementing one or more exemplary embodiments. Various changes can be contemplated with respect to the functions and configurations of elements without departing from the spirit and scope of the subject matter disclosed as set forth in the appended claims.
[0021] In the following description, specific details are provided to provide a complete understanding of the embodiments. However, those skilled in the art will understand that the embodiments can be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary details. In other instances, well-known processes, structures, and techniques may be shown without unnecessary details to avoid obscuring the embodiments. Further, like reference numbers and names in the various drawings indicate like elements.
[0022] Also, individual embodiments may be described as processes shown as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. A flowchart may describe operations as a sequential process, but many of the operations may be executed in parallel or simultaneously. Additionally, the order of the operations may be rearranged. A process may end when its operations are completed, but may also have additional steps not depicted or included in the figure. Further, not all of the operations in any particular process described will be performed in all embodiments. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the end of the function may correspond to the return of the function to the calling function or main function.
[0023] Furthermore, embodiments of the disclosed subject matter may be realized, at least in part, manually or automatically. Manual or automatic realizations may be executed or at least assisted by using a machine, hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. When realized in software, firmware, middleware, or microcode, program code or code segments for performing the necessary tasks may be stored on a machine-readable medium. A processor may execute the necessary tasks.
[0024] The modules and networks illustrated in this disclosure may be a computer program, software, or instruction code that can execute instructions using one or more processors. The modules and networks may be stored in one or more storage devices, or may be stored in a computer-readable medium, such as a storage medium, a computer storage medium, etc., or a (removable and / or non-removable) data storage device, such as a magnetic disk, an optical disk, or a tape, etc. The computer-readable medium is accessible from one or more processors for executing instructions.
[0025] A computer storage medium can include volatile and non-volatile media, removable and non-removable media, implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer storage medium may be RAM, ROM, EEPROM or flash memory, CD-ROM, digital versatile disk (DVD) or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, or any other medium that can be used to store the desired information and is accessible by an application, a module, or both using one or more processors. Such any computer storage medium may be part of a device or may be accessible or connectable to the device. Any application or module described herein may be implemented using computer-readable instructions / computer-executable instructions that may be stored or held by such a computer-readable medium. (Low Latency Captioning)
[0026] According to some embodiments of the present disclosure, a scene captioning system can provide low-latency scene captioning for understanding a scene and describing events in natural language from real-time monitoring of a scene or online / offline scene streaming.
[0027] In some cases, the scene captioning system can be configured as a video captioning system, and the video captioning system may be referred to as a low-latency video captioning method / system.
[0028] Hereinafter, a low-latency video scene captioning method / system is provided as an exemplary description of a low-latency scene captioning method / system.
[0029] For example, a stream of signals obtained by a unimodal sensor or a multimodal sensor can be converted into captions using natural language. When a single sensor is applied, the unimodal is included in the multimodal. Multimodal sensors capture information such as visual images, audio signals, 3D localization, thermography, Wi-Fi, etc. In some cases, the sensor can be an image sensor, a microphone, an audio-visual signal sensor, LiDAR, mmWave, a thermal sensor, an olfactory sensor, a tactile sensor, or any combination thereof.
[0030] When a low-latency scene captioning method / system is applied to a real-time or offline video stream, the multimodal sensor signal can be a stream of audio-visual signals, the scene encoder can be an audio-visual encoder, and the characteristics of the unimodal sensor / multimodal sensor can be audio characteristics and visual characteristics. Further, the multimodal sensor signal may be a signal / information acquired by a multimodal sensor. For example, the signal can be any of the detection information such as video frames, audio data, 3D localization data, thermal sensor data, olfactory sensor data, tactile sensor data, or any combination thereof.
[0031] FIG. 1 is a schematic diagram showing a process of low-latency video captioning according to an embodiment of the present invention. A video stream 101 is given as X, an event occurs between time s and time T, and it is assumed that the ground truth caption 102 as Y can be "One of the workers was splashed by a bulldozer" regarding the event. The caption is typically a sentence that describes the event in natural language. A video captioning system based on the prior art generates such a caption for a video clip using the entire frame x of the video stream X, where x s:T is the video frame x s:T and x s is the video frame x s+1, …, x T-1 , x T shows a sequence of, each corresponding to a video frame for each time index. Thus, such a conventional video captioning system is assumed to function only offline.
[0032] However, such a conventional method is not practical in a real-time monitoring or surveillance system where it is essential not only to accurately describe events but also to generate captions as soon as possible. For this purpose, the low-latency video captioning module 110 uses a partial video frame 104, i.e., the video frame x from the start time s to the current time t s:t to have a timing detector 103 that determines whether the current time t is an appropriate timing for outputting a caption. Only when the timing detector 103 detects an appropriate timing, the low-latency caption generator 105 generates a caption 106 based on the partial video frame 104. The start time s can be determined based on the timing at which the system generated a previous caption or the timing at which a specific change in pixel intensity within the video is detected (not shown in FIG. 1).
[0033] Another remaining problem is that there is no sufficient dataset to train the timing detector 103 and the low-latency caption generator 105. In this case, the dataset must be annotated with appropriate early timings and their respective captions for various videos. A dataset exists only for offline video captioning, i.e., the dataset is annotated at the timing when the event has already ended in the video. Generally, it is extremely costly to correct a large amount of new datasets using such early timings annotations.
[0034] To solve this problem, embodiments of the present invention train the timing detector 103 and the low-latency caption generator 105 using only a dataset for offline video captioning. In this case, the detector and the generator are optimized not only to accurately describe the event but also to generate captions as soon as possible. Therefore, low-latency video captioning is achieved by detecting the appropriate timing for generating correct captions using only the partial video frame 104 up to the current time t. The system can generate a caption 106 equal to the ground truth 102 "One of the workers was splashed by a bulldozer" even though it uses only the partial video frame 104. On the other hand, in the conventional method of training the generator for offline video captioning, since the generator is trained to generate captions for the events occurring in each video clip, there is a possibility of generating an incorrect caption 107 "The worker is walking" for the partial video frame 104. In the example of the low-latency video captioning 100, the event "One of the workers was splashed by a bulldozer" has not yet occurred at time t. Therefore, it is difficult for a system based on the prior art to generate a correct caption, but the system according to the present invention can generate a correct caption because the generator is trained to utilize some signs regarding future events in the partial video frame for generating the correct caption. (Low-Latency Captioning-Based Dialogue System)
[0035] FIG. 2 is a block diagram showing a low-latency video captioning system 200 including a low-latency caption generation training module 300 composed of a training data set 301, a feature extractor 302, a future frame eliminator 303, a pre-trained caption generator 320, a caption generation loss calculator 309, a caption similarity checker 310, and a timing detection loss calculator 311, a low-latency captioning module 110 composed of a timing detector module 103 and a low-latency caption generator 105, and a dialogue module 201 composed of a caption understanding module 202 and a response generator 203.
[0036] The low-latency video captioning system 200 includes a human machine interface (HMI) 210 connectable to a keyboard 211 and a pointing device / media 212, one or more processors 220, a memory 240, a network interface controller (NIC) 250 connectable to a network 290 including a local area network and an Internet network, a display and / or speaker interface 260 connectable to a display and / or speaker device 261, a machine interface 262 connectable to a machine actuator 263, and a multimodal sensor interface 271 including respective audio interfaces 273 and visual interfaces 275 connectable to an input device including multimodal sensors 272 such as a microphone device 274 and a camera device 276, and a printer interface 280 connectable to a printing device 285. The memory 240 may be one or more memory units. The low-latency video captioning-based interaction system 200 can receive multimodal detection data 295 via the network 290 connected to the NIC 250. The storage device 230 includes a feature extractor 302, a low-latency caption generation training module 300, a low-latency captioning module 110, and an interaction module 201. In some cases, the feature extractor 302 may be configured as a multimodal sensor feature extractor 302 when the system is configured as a scene captioning system. Further, the feature extractor 302 may be configured as an audio-visual feature extractor 302 when the system is configured as a video captioning system.
[0037] To perform low-latency captioning, instructions can be sent to the low-latency video captioning system 200 using the keyboard 211, the pointing device / media 212, or via a network 290 connected to another computer (not shown). The system 200 receives the instructions via the HMI 210 and executes instructions for performing low-latency video captioning using the processor 220 associated with the memory 240 by loading the low-latency captioning module 110.
[0038] The low-latency video captioning module 110 outputs a token sequence as a captioning result for a given multimodal detection feature sequence obtained by the feature extractor 302, and transmits the token sequence to the display / speaker device 265 via the display / speaker interface 260, to the printer device 285 via the printer interface 280, or to another computer (not shown) via the network 290. Each token in the token sequence can be a single word, a single character, a single glyph, or a word fragment in text form. (Training Procedure for Low-Latency Video Captioning System)
[0039] FIG. 3 is a block diagram showing a training module 300 included in a low-latency video captioning system 200 that includes a low-latency captioning module 110 that utilizes a pre-trained caption generator 320 to jointly optimize the low-latency caption generator 103 and the timing detector 105 according to an embodiment of the present invention.
[0040]
Number
[0041]
Number
[0042] The timing detector 103 may include an encoder 307 and a timing detection module 308. The encoder 307 encodes the future-excluded feature vector sequence Z′ into the hidden activation vector sequence H′. The timing detector 103 may use the encoder 304 of the low-latency caption generator 105 instead of the encoder 307. In this case, the timing detector 103 may not have the encoder 307 and may receive the output of the encoder 304, that is, H′ may be obtained as H′ = H.
[0043]
Number
[0044]
Number
[0045] When the encoder 304 and the decoder 305 are designed as differentiable functions such as neural networks, the parameters of the neural network can be trained to minimize the cross-entropy loss using the backpropagation algorithm. Minimizing the cross-entropy loss means causing the low-latency caption generator 105 to generate an appropriate caption as close as possible to the ground truth caption. Further, when the system is configured as a scene captioning system, the encoder 304 can be configured as a multimodal sensor feature encoder 304. When the system is configured as a video captioning system, the encoder 304 can be configured as an audio-visual feature encoder 304.
[0046]
Number
Number
[0047]
Number
[0048] Similar to the low-latency caption generator 105, the timing detector 103 can be trained using the backpropagation algorithm when the encoder 307 and the timing detection module 308 are designed as differentiable functions such as neural networks.
[0049]
Number
[0050]
Number
[0051]
Number
[0052]
Number
[0053]
Number
[0054] This mechanism prompts the timing detector 103 to detect the timing at which the low-latency caption generator 105 can generate a caption that is close not only to the ground truth caption Y but also to the caption Y' generated by the pre-trained caption generator 203 for the entire feature vector sequence Z. This also avoids relying solely on the similarity with the ground truth caption Y for the supervision of the timing detector 103, thereby hopefully improving the robustness of the timing detector 103. (Some embodiments based on audio-visual transformers)
[0055] FIG. 4 is a schematic diagram showing an audio-visual transformer 400 constructed as a low-latency caption generator combined with a timing detector for low-latency video captioning according to an embodiment of the present invention. The audio-visual transformer 400 for low-latency video captioning includes a feature extractor 401, an encoder 402, a decoder 403, and a timing detection module 404.
[0056] In the case of video stream 405, the feature extractor 401 extracts VGGish features 406 and I3D features 407 from the audio signal 408 and the visual signal 409 from the audio track and the visual track of the video stream 405, respectively. Here, the audio signal 408 can be an audio waveform, and the visual signal 409 can be a sequence of images. The frame rate for feature extraction can be different for each track. The encoder 402 encodes the feature vectors of VGGish 406 and I3D 407. In this case, the sequence of audio features and visual features from the start point to the current time is supplied to the encoder 402 and is converted into a sequence of hidden activation vectors through the self-attention layer 410, the bimodal attention layer 411, and the feed-forward layer 412. Typically, this encoder block 413 is repeated N times (e.g., N = 6 or more). The last sequence of hidden activation vectors is obtained via the Nth encoder block.
[0057]
Number
[0058]
Number
[0059]
Number
[0060] The self-attention layer 410 extracts the temporal dependencies within each modality, where all the arguments for MHA() are the same, i.e., A n-1 or V n-1It is. The bimodal attention layer 411 further extracts cross-modal dependencies between audio features and visual features, in which case it takes keys and values from other modalities. Then, a feed-forward layer 412 is applied pointwise. The encoded representations for the audio features and visual features are A N and V N obtained as such.
[0061] The timing detection module 404 receives the encoded hidden vector sequence based on the audio-visual information available at that time. The role of the timing detection module 404 is to estimate whether the system should generate a caption for a given encoded feature. The timing detection module 404 first processes the encoded vector sequences from each modality where the 1D convolutional layer 414 is stacked as follows.
Number
[0062] Next, each temporal convolutional sequence is aggregated into a single vector through the operations of pooling 415 and concatenation 416.
Number
[0063]
Number
[0064] When the timing detection module 404 provides a probability higher than a threshold, for example, P(d = 1|X A , X V ) > 0.5, the decoder generates a caption based on the encoded hidden vector sequence (A N , V N ).
[0065] The decoder 403 is the start word ( <sos>) repeatedly predicts the next word. At each iteration step, the decoder receives the previously generated word 419 and estimates the posterior probability distribution of the next word 420 by applying word embedding 421, M decoder blocks 422, a linear layer 423, and a softmax operation 424.
[0066] [Number]
[0067] [Number]
[0068] [Number]
[0069] Select multiple words with the highest probabilities and consider multiple candidates for the caption according to the beam search technique in the search module 306. (Training)
[0070] The multimodal encoder, the timing detector, and the caption decoder are jointly trained so that the model achieves caption quality comparable to that of the complete video even if the given video is shorter than the original video by truncating the latter part. [Number] (Evaluation of low-latency video captioning quality)
[0071] Figure 5 shows the evaluation results obtained by running the video captioning benchmark according to an embodiment of the present invention.
[0072] The proposed low-latency caption generation was tested using the ActivityNet Captions dataset, which consists of 100k caption sentences associated with temporary localization information based on 20k YouTube (R) videos. The dataset was split into 50%, 25%, and 25% for training, validation, and testing. The validation set was split into two subsets, and the performance regarding the two subsets was reported. The average duration of video clips is 35.5 seconds, 37.7 seconds, and 40.2 seconds for the training set and validation subsets 1 and 2, respectively. VGGish features and I3D features were used. The VGGish features were configured to form a 128-dimensional vector sequence for the audio track of each video, where each audio frame corresponds to a non-overlapping segment of 0.96 seconds. The I3D features were configured to form a 1024-dimensional vector sequence for the video track, where each visual frame corresponds to a non-overlapping segment of 2.56 seconds.
[0073] First, the multimodal transformer was trained with the entire video clips and their ground-truth captions. This model was used as the baseline and teacher model. N = 2 was used for the encoder blocks, M = 2 was used for the decoder blocks, and the number of attention heads was 4. The vocabulary size was 10172, and the dimension of the word embedding vector was 300.
[0074] The proposed model for online captioning was trained using incomplete video clips according to the steps of the present invention. In the training process, α = β = γ = 1 / 3 was used for the loss function. The dimensions of the hidden activations in the audio attention layer and the visual attention layer were 128 and 1024 respectively. The timing detector had two stacked 1D convolutional layers with ReLU non-linearity in between. The performance was measured by BLEU3 score, BLEU4 score, and METEOR score.
[0075] The latency ratio indicates the ratio of the video duration used for captioning to the duration of the original video clip. In the baseline model, the latency ratio was always 1, which means that all frames were used to generate the caption.
[0076] Figure 5 compares captioning methods in terms of METEOR score for validation subset 1. The model selected for evaluation was trained with (empirically determined) S = 0.6 and had the best METEOR score for validation subset 2. The latency was controlled at detection threshold F. As shown in Figure 5, the proposed method at 55% latency achieves a METEOR score of 10.45, which is only a slight degradation and corresponds to 98% of the baseline score of 10.67. It also achieves a METEOR score of 10.00, corresponding to 94% of the baseline, at 28% latency. A naive method was tested where video frames were taken from the start at a fixed ratio to the original video length and baseline captioning was performed on the shortened video clip. The results show that the proposed approach is clearly superior to the naive method at equivalent latencies.
[0077] The table also includes the results using a unimodal transformer that receives only visual features. These results indicate that the proposed method functions only for visual features and the performance degrades due to the lack of audio features. This result shows that audio features are essential even for the proposed low-latency method.
[0078] The low-latency caption is input into a dialogue system trained using pairs of captions and action commands to understand the scene and determine the next action. The dialogue system may use a sequence-to-sequence neural network model that converts the received caption into a sequence of words for responding to people or a sequence of action commands for controlling a robot.
[0079] Furthermore, some embodiments are based on the recognition that scene-aware interaction technology enables a machine to interact based on shared knowledge obtained by recognizing and understanding humans and their surroundings using various types of sensors. Audio-visual scene-aware dialog (AVSD) is one of the scene-aware interaction technologies. According to a certain recognition based on AVSD, the end-to-end approach can better handle flexible conversations between users and systems by training models on large conversation datasets. Such an approach is extended to conduct conversations about objects and events occurring around a machine or a user based on the understanding of dynamic scenes captured by multi-modal sensors such as video cameras and microphones. This extension enables users to ask questions about what is happening around them. For example, this framework can be applied to visual question answering (VQA) for one-shot QA regarding still images, visual dialog where an AI agent conducts a meaningful conversation with a human regarding a still image using natural conversation language, video QA for one-shot QA regarding video clips, and AVSD where a multi-turn dialog task is based on generating responses to users' questions regarding video clips of daily life. Video QA is a QA (question answering) task based on answering a single question regarding video clips as seen on YouTube (registered trademark) and is standardized in the MSVD (Microsoft Research Video Description Corpus) and MSRVTT (Microsoft Research Video to Text) datasets. AVSD is a multi-turn dialog task based on generating responses to users' questions regarding video clips of daily life, and Dialog Systems Technology Challenges (DSTC) addresses this.The AVSD system is trained to create a text that answers a user's question about a video clip. In this case, the system needs to understand what events were done, when, how, and by whom based on the audio-visual features in time series in order to provide a correct answer.
[0080] At any moment, countless events occurring in the real world are captured by cameras and stored as a large amount of video data resources. Since the S2VT (Sequence to Sequence -- Video to Text) system was first proposed, video captioning has been actively studied in the field of computer vision using an end-to-end sequence-to-sequence model to effectively retrieve such records regardless of whether it is an offline setting or an online setting. Its goal is to generate video descriptions (captions) about objects and events within a video clip. To further utilize audio features to identify events, the multi-modal attention method enables the fusion of audio and visual features such as VGGish (a voice classification model similar to the Visual Geometry Group (VGG)) and I3D (Interactive Three Dimensional) to generate video captions. Such video clip captioning technology has been extended to offline video stream captioning technologies such as high-density video captioning and progressive video description generators. In this case, all the events of interest in the video stream are localized in time, and captions triggered by the events are generated in a multi-threaded manner. All video captioning technologies have been based on LSTM so far, but the transformer can be successfully applied in combination with the audio-visual attention framework.
[0081] Current dialog systems for video QA typically process offline video data and generate answers to questions after the clip ends. In order for the dialogue system to enable real-time conversations about scenes for monitoring and robot applications, it is necessary to understand the scene and events and quickly respond to the user's queries about the scene. In that research, an audio-visual transformer was tested using the ActivityNet Captions dataset within an offline video captioning system and achieved the best performance for the high-density video captioning task.
[0082] However, such scene-aware interaction tasks assume that the system has access to all frames of the video clip when predicting answers to questions. This assumption is not realistic for a real-time dialog system monitoring an ongoing audio-visual scene, and for this system, it is essential not only to accurately predict answers but also to quickly discover question-related events in the online video stream and respond to the user as soon as possible by generating appropriate answers. Such capabilities require the development of new low-latency QA technologies.
[0083] In previous research, we proposed a low-latency audio-visual captioning method that can accurately and quickly describe events without waiting for the end of the video clip and optimize the timing of captioning. In parallel, we recently introduced a new AVSD task for the 3rd AVSD challenge in DSTC10, which particularly requires the system to demonstrate temporal reasoning by discovering evidence from the video to support its answers. This is based on a new extension of the DSTC10 AVSD dataset, for which we collected human-generated temporal reasoning data. This study proposes to extend our low-latency captioning approach to the scene-aware interaction task and combine it with our inference AVSD system to develop a low-latency online scene-aware interaction system.
[0084] Some extensions of research on low-latency video captioning can construct new methods that can optimize the timing of generating each answer under the trade-off between the latency of answer generation and the quality of the answer. In the case of video QA, the timing detector now plays the role of discovering the timing of events related to the question, instead of judging when sufficient understanding has been obtained to generate a general caption as in the case of video captioning. The audio-visual scene-aware dialogue system constructed for the 10th Dialog System Technology Challenge can be extended to utilize low-latency capabilities. As an example, experiments with the MSRVTT-QA and AVSD datasets show that our approach achieves 97% - 99% of the answer quality upper bound given by a transformer pre-trained using less than 40% of the frames from the beginning compared to using the entire video clip.
[0085] In a similar vein to the low-latency captioning approach, another system according to another embodiment optimizes the output timing for each answer based on the trade-off between latency and answer quality. We train a low-latency audio-visual transformer composed of (1) a transformer-based response generator that attempts to generate a ground-truth answer after looking at only a small fraction of all video frames and also attempts to mimic the output of a similarly pre-trained response generator that can view the entire video, and (2) a CNN-based timing detector that can find the best timing to output an answer to the input question such that the responses finally generated by the two transformers are close enough to each other. The proposed jointly trained response generator and timing detector can generate a response at the initial stage of the video clip as soon as a relevant event occurs and can even predict future frames. Thanks to the combination of information from multiple modalities, the system has more opportunities to recognize events at an earlier timing by relying on the earliest cue in one of the modalities. Experiments on the MSR-VTT QA and AVSD datasets show that our approach achieves low-latency video QA with answer quality comparable to the offline video QA baseline that uses the entire video frames.
[0086] According to an embodiment of the present invention, video QA different from video captioning is provided, and an appropriate answer to the user's question is generated as early as possible. To that end, we introduce a question encoder to provide question embeddings to the timing detector and extend the text generator to accept the question as context information. Thus, the proposed method uses the same basic mechanism for low-latency processing, but the model is extended for the video QA task. (Low-latency video QA model)
[0087] We construct our proposed model for low-latency video QA on the DSTC10-AVSD system that employs an AV Transformer architecture. For the DSTC10-AVSD challenge, we extended the AV Transformer using student-teacher co-learning and attention multimodal fusion to achieve state-of-the-art performance. Like our low-latency video captioning system, the low-latency QA model receives video and audio features in a streaming manner, and the timing detector determines when to generate an answer response for the feature sequence received by the model up to that instant. FIG. 6 shows a model architecture (low-latency video QA architecture) 600 for low-latency video QA according to some embodiments of the present invention. The model architecture 600 includes functional parts similar to those used in the audio-visual transformer 400. In the drawings, the same numbers are used for the parts (or layers) corresponding to the parts (or layers) used in the audio-visual transformer 400. The low-latency video QA architecture 600 further includes a word embedding 621, a question encoder (text encoder) 613 having a self-attention layer 610 and a feed-forward layer 612, a pooling 615 connected to a contact 416, an AV encoder 402, a timing detector 404, and a response decoder 403, and the AV encoder 402 is shared by the timing detector 404 and the response decoder 403.
[0088] Furthermore, the mathematical formulas or equations used to describe the audio-visual transformer 400 are used in the following description to describe the low-latency video QA architecture 600.
[0089] When a video stream and question text are provided as input, the AV encoder encodes the VGGish and I3D features extracted from the audio and video tracks, and the transformer-based text encoder encodes the question. A sequence of audio and visual features from the starting point to the current point is given to the encoder and transformed into a sequence of hidden vectors through self-attention layers, bimodal attention layers, and feed-forward layers. This encoder block is repeated N times, and the encoded final representation is obtained through the Nth encoder block. The question word sequence is also encoded through a word embedding layer, followed by a transformer with N' blocks.
[0090] [Number]
[0091] [Number]
[0092] [Number]
[0093] [Number]
[0094] When the timing detector outputs a probability higher than the threshold, for example, when P(d = 1|A, V, Q)>0.5, the decoder generates an answer based on the encoded representation.
[0095] [Number] (Training and Inference)
[0096] Train an AV encoder, a question encoder, a response decoder, and a timing detector jointly so that the system achieves a response quality comparable to that of the complete video even when the given video is shorter than the original video by trimming the latter half part.
[0097]
Number
[0098]
Number
[0099]
Number
[0100]
Number
[0101]
Number
[0102]
Number
[0103] Note that T S is assumed to be already determined.
[0104] According to some embodiments of the present invention, an artificial intelligence (AI) low-latency processing system is provided. FIG. 8 shows a computer-implemented method 800 according to an embodiment of the present invention, including process steps executed by a low-latency processing system using a processor and a memory or a memory storage.
[0105] A low-latency processing system may include a processor and a memory storing instructions as a method 800 implemented by a computer. When executed by the processor, the instructions cause the low-latency processing system to collect a sequence of frames. The sequence of frames jointly includes information distributed among at least some of the frames in the sequence of frames. Further, the instructions cause the execution of a timing neural network trained to identify a subsequence of an initial frame of the sequence of frames that includes at least a portion of the information indicating the information, and the execution of a decoding neural network trained to decode the information from a portion of the information in the subsequence of the frames. The timing neural network is co-trained with the decoding neural network to repeatedly identify a minimum number of sub-frames from the beginning of a training sequence of frames that includes a portion of the training information sufficient to decode the training information.
[0106] In some cases, for the features of different subsequences in the sequence of frames, the timing detector neural network is jointly trained with the decoder neural network to minimize a multi-task loss function that includes a time detection loss and an information generation loss. The multi-task loss function can include three losses, and the three losses are: (1) the accuracy of the decoded information, (2) the difference between the information decoded from the subsequence of frames and the information decoded from the entire sequence of frames, and (3) the accuracy of the prediction of the timing detector neural network. Further, the processor may be configured to extract features from each frame in the sequence of frames by executing a feature extractor neural network, encode the extracted features of each frame by executing a feature encoder neural network to generate a sequence of encoded features, provide the sequence of encoded features to the timing detector neural network to identify a subsequence of encoded features representing a subsequence of frames, and provide the subsequence of encoded features to the decoder neural network to decode information. In other cases, the processor may trigger the execution of the modules of the AI low-latency processing system when receiving a new input frame attached to the sequence of frames. The information may be a caption of an audio scene, a video scene, or an audio-video scene. Also, the information is an answer to a question regarding the sequence of frames. In some cases, a frame may include multi-modal information from different sensors of different modalities. Further, the information may be an answer to a question regarding the sequence of frames, and the processor may execute a text encoder neural network to encode the question, provide the encoded question to the timing neural network, and provide the question or the encoded question to the decoder neural network. (Experiment)
[0107] We evaluate our low-latency video QA method using the MSRVTT-QA and AVSD datasets. MSRVTT-QA is based on the MSR-VTT dataset containing 10k video clips and 243k question-answer (QA) pairs. The QA pairs are automatically generated from captions manually annotated for each video clip, where the questions are sentences and the answers are single words. We follow the data split in the MSR-VTT dataset, with 65% for training, 5% for validation, and 30% for testing. AVSD is a set of text-based dialogues about short videos from the Charades dataset, consisting of untrimmed multi-action videos each containing an audio track. In AVSD, as two participants, a named questioner and an answerer conduct a dialogue regarding the events in the provided video. The job of the answerer, who has already watched the video, is to answer the questions asked by the questioner. We follow the AVSD challenge setting, where the training, validation, and test sets consist of 7.7k, 1.8k, and 1.8k dialogues respectively, each dialogue containing 10-turn QA pairs, where both the questions and answers are sentences. The length of the video clips ranges from 10 to 40 seconds.
[0108] The VGGish features were configured to form a 128-dimensional vector sequence for the audio track of each video, where each audio frame corresponds to a non-overlapping 0.96-second segment. The I3D features were configured to form a 2048-dimensional vector sequence for the video track, where each visual frame corresponds to a non-overlapping 2.56-second segment.
[0109] We first trained a multimodal transformer on all video clips and QA pairs. This model was used as the baseline and teacher model. We used N = 2 audio-visual encoder blocks, N' = 4 question encoder blocks, M = 4 decoder blocks, and set the number of attention heads to 4. The vocabulary sizes were 7,599 for MSRVTT-QA and 3,669 for AVSD. The dimension of the word embedding vectors was 300.
[0110] The proposed model for low-latency video QA was trained on incomplete video clips following the steps in Section 3.2. The architecture was the same as the baseline / teacher model except for the addition of the timing detector. In the training process, we consistently used α = β = γ = 1 / 3 in the loss function and the threshold S = 0.9 in Equation (24). We set the dimensions of the hidden activations in the audio and visual attention layers to 256 and 1024 respectively, set the dropout rate to 0.1, and applied label smoothing techniques. The timing detector consisted of two stacked 1D convolutional layers with ReLU non-linearity in between. The performance was measured by the answer accuracy of MSRVTT-QA and the BLEU4 and METEOR scores of AVSD.
[0111] Figure 9 shows the relationship between the latency ratio and the answer accuracy for MSRVTT-QA. The latency ratio represents the ratio of the frames actually used (from the beginning) to the entire video frames. The baseline results were obtained by simply omitting future frames at various ratios using the baseline (teacher) model. The results of the proposed model were obtained by changing the detection threshold F. The accuracy of MSRVTT-QA indicates the ratio of one-word answers that match the ground truth. This result demonstrates that the method we proposed realizes low-latency video QA with a much smaller accuracy degradation compared to the baseline. Our approach achieves 97% of the upper limit of the answer quality given by the transformer pre-trained using the entire video clip by using only 40% of the frames from the beginning.
[0112] Figure 10 shows the comparison of the quality of the answer sentences for the AVSD task. The latency ratio was controlled by setting the detection threshold F so that average ratios of 1.0, 0.5, and 0.2 were obtained for the proposed method. Future frames with the above fixed ratios were excluded for the baseline system. As shown in the table, the method we proposed is slightly better than the baseline even at a latency of 1.0. This may be due to the increased robustness of the model by training with randomly shortened videos. In addition, the proposed method maintains the same level of BLEU4 and METEOR scores at a latency of 0.5, achieves competing scores even at a latency of 0.2 with slight degradation, and reaches 98% - 99% of the scores under the condition of a latency of 1.0.
[0113] Figures 11A and 11B show the distribution of QAs over latency for detection thresholds F = 0.3 and F = 0.4 for the AVSD-DSTC7 task, which correspond to the results at latencies 0.2 (left) and 0.5 (right) in Figure 10. These results indicate that most of the QAs for AVSD require either only the initial frame or all of the frames to generate the correct answer. The reasons for the polarized distribution were investigated. The most frequent pattern of questions leading to early decisions is "How does the video start?" Furthermore, there are some fixed answers in the training data, such as "one" for "How many people are there in the video?" Such common language patterns can also cause early decisions. Cases where the decision is delayed include patterns such as "How does the video end?" Such questions are natural for questioners who need to generate video captions through 10 QAs without watching the entire video.
[0114] According to some of the above embodiments, the low-latency video QA method can accurately and quickly answer the user's questions without waiting for the end of the video clip. The proposed method optimizes the output timing of each answer based on the trade-off between latency and answer quality. The above system can generate answers at the initial stage of the video clip using the MSRVTT-QA and AVSD datasets, achieving 97% - 99% of the answer quality upper limit provided by a transformer pre-trained using less than 40% of the frames from the beginning to the entire video clip.< / sos>
Claims
1. An artificial intelligence (AI) low-latency processing system, wherein the low-latency processing system comprises a processor and a memory storing instructions, and when the instructions are executed by the processor, the low-latency processing system is caused to execute collecting a sequence of frames, the sequence of frames jointly containing information distributed among at least some of the frames in the sequence of frames, and further execute a timing neural network trained to identify a subsequence of an initial frame of the sequence of frames that includes at least a portion of the information indicating the information, and execute a decoding neural network trained to decode the information from a portion of the information in the subsequence of the frames, and The timing neural network is co-trained with the decoding neural network to repeatedly identify a minimum number of sub-frames from the start of a training sequence of frames that includes a portion of the training information sufficient to decode the training information. An AI low-latency processing system.
2. The timing detector neural network is co-trained with the decoder neural network for features of different subsequences of the sequence of frames to minimize a multi-task loss function including a time detection loss and an information generation loss. The AI low-latency processing system according to claim 1.
3. The multi-task loss function includes three losses, and the three losses are (1) the accuracy of the decoded information, (2) the difference between the information decoded from the subsequence of the frames and the information decoded from the whole sequence of the frames, and (3) the prediction accuracy of the timing detector neural network. The AI low-latency processing system according to claim 2.
4. The processor extracts features from each frame in the sequence of frames by executing a feature extraction neural network, encodes the extracted features of each frame by executing a feature encoder neural network to generate a sequence of encoded features, Provide the sequence of the encoded features to the timing detector neural network to identify a subsequence of the encoded features representing a subsequence of the frame, The AI low-latency processing system according to claim 1, which is configured to provide the subsequence of the encoded features to the decoder neural network to decode the information.
5. The AI low-latency processing system according to claim 1, wherein when the processor receives a new input frame attached to the sequence of the frames, the execution of the modules of the AI low-latency processing system is triggered.
6. The AI low-latency processing system according to claim 1, wherein the information is a caption of an audio scene, a video scene, or an audio-video scene.
7. The AI low-latency processing system according to claim 1, wherein the information is an answer to a question regarding the sequence of the frames.
8. The information is an answer to a question regarding the sequence of the frames, The processor, Executes a text encoder neural network to encode the question, Provide the encoded question to the timing neural network, The AI low-latency processing system according to claim 4, which is configured to provide the question or the encoded question to the decoder neural network.
9. The AI low-latency processing system according to claim 1, wherein the frame includes multimodal information incoming from different sensors of different modalities.
10. A computer-implemented method of an artificial intelligence (AI) low-latency processing system, wherein the low-latency processing system includes a processor and a memory storing instructions of the method implemented by the computer to execute steps using the processor, and the steps include: Collecting a sequence of frames, the sequence of frames jointly including information distributed among at least some of the frames in the sequence of frames, and further, Executing a timing neural network trained to identify a subsequence of an initial frame in the sequence of frames that includes at least a part of the information indicating the information. Executing a decoding neural network trained to decode the information from a part of the information in the subsequence of the frame, The timing neural network is co-trained with the decoding neural network to repeatedly identify the minimum number of sub-frames from the beginning of the training sequence of the frame, including a part of the training information sufficient to decode the training information, which is a method implemented by a computer. **Claim 11** The method implemented by a computer according to claim 10, wherein the timing detector neural network is co-trained with the decoder neural network to minimize a multi-task loss function including a time detection loss and an information generation loss for features of different subsequences in the sequence of the frame. **Claim 12** The method implemented by a computer according to claim 11, wherein the multi-task loss function includes three losses, and the three losses are: (1) the accuracy of the decoded information; (2) the difference between the information decoded from the subsequence of the frame and the information decoded from the whole sequence of the frame; and (3) the prediction accuracy of the timing detector neural network. **Claim 13** The processor Extracts features from each frame in the sequence of the frame by executing a feature extractor neural network, Encodes the extracted features of each frame by executing a feature encoder neural network to generate a sequence of encoded features, Provides the sequence of encoded features to the timing detector neural network to identify a subsequence of encoded features representing the subsequence of the frame, The method implemented by a computer according to claim 10, which is configured to provide the subsequence of encoded features to the decoder neural network to decode the information. **Claim 14** The method implemented by a computer according to claim 10, wherein when the processor receives a new input frame attached to the sequence of the frame, it triggers the execution of the modules of the AI low-latency processing system. **Claim 15** The method according to claim 10, wherein the information is a caption of an audio scene, a video scene, or an audio-video scene.
16. The method according to claim 10, wherein the information is an answer to a question regarding the sequence of the frames.
17. The method according to claim 10, wherein the frame includes multi-modal information incoming from different sensors of different modalities.
18. The information is an answer to a question regarding the sequence of the frames, the processor is configured to execute a text encoder neural network to encode the question, provide the encoded question to the timing neural network, and provide the question or the encoded question to the decoder neural network, according to the method realized by the computer according to claim 13.
Citation Information
Patent Citations
Method and device for generating description information of multimedia data, equipment and medium
CN111723937A
System and method for attention-based configurable convolutional neural network (abc-CNN) for visual question answering
JP2017091525A
Method of generating summary of medial file that comprises a plurality of media segments, program and media analysis device
JP2018124969A
Method, apparatus, device and medium for generating captioning information of multimedia data
US20220014807A1