Neuro-symbolic action transducer for video question answering

By using a neural symbolic action transformer and employing situational graphs and hypergraph data structures for logical reasoning, the limitations of existing VQA AI systems in answering complex questions are overcome, enabling logical reasoning and automated responses to input video data.

CN115700589BActive Publication Date: 2026-07-03INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2022-07-19
Publication Date
2026-07-03

Smart Images

  • Figure CN115700589B_ABST
    Figure CN115700589B_ABST
Patent Text Reader

Abstract

This disclosure relates to a neural symbolic action transformer for video question answering. A mechanism for performing AI-based video question answering is provided. A video parser parses an input video data sequence to generate situational data structures, each including data elements corresponding to entities and a first relation between the entities, said entities and the first relation being identified by the video parser as existing in images of the input video data sequence. A first machine learning computer model operates on the situational data structures to predict a second relation between the situational data structures. A second machine learning computer model executes on a received input question to predict an executable program to be executed to answer the received question. This program executes on the situational data structures and the predicted second relation. An answer to the question is output based on the result of executing the program.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Progressively Extending Conversation Scope in Multi-User Messaging Platform

    US20190149489A1