Individual inference path-based video question and answer device and method therefor
The individual inference path-based video question-answering device addresses the challenge of adapting large-scale models to diverse domains by extracting domain-independent features, reducing computational needs and enabling efficient domain adaptation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- KYUNGPOOK NAT UNIV IND ACADEMIC COOP FOUND
- Filing Date
- 2025-11-26
- Publication Date
- 2026-06-04
AI Technical Summary
Large-scale pre-trained models for video question answering require significant computational resources and time for fine-tuning, making them difficult to adapt to various domains effectively, especially in environments with limited resources.
An individual inference path-based video question-answering device that extracts domain-independent motion and object features from audio and text features using a pre-trained model, employing a dictionary database, feature decomposer, and global expression derivation units to derive answers without full fine-tuning.
Enables effective adaptation to multiple domains with reduced computational requirements, leveraging pre-trained model features efficiently.
Smart Images

Figure KR2025019822_04062026_PF_FP_ABST
Abstract
Description
Individual inference path-based video question answering device and method
[0001] The present invention relates to a video question-answering device and a method thereof, and more specifically, to a video question-answering device and a method thereof based on individual inference paths using a pre-trained model that encodes audio features and text features from multi-domain input data.
[0002] Video question answering is a core challenge in deep learning, requiring systems to understand visual content in response to natural language queries. Recently, the adoption of large-scale pre-trained models in the field of video question answering has become the standard due to their ability to adapt to various downstream tasks. However, the size of these large-scale pre-trained models and their massive computational resource requirements present a major challenge that makes effective fine-tuning difficult in environments with limited resources.
[0003] In particular, it is difficult to leverage the potential of large-scale pre-trained models without access to sufficient computational resources. Furthermore, these large-scale pre-trained models do not adapt well to other domains without fine-tuning the entire framework. This fine-tuning process is not only computationally intensive but also requires significant resources and time.
[0004] Nevertheless, large-scale pre-trained models inherently possess a wealth of features acquired through vast training data. Therefore, it is essential to seek a computationally low-cost approach that can leverage the potential of these pre-trained features without fine-tuning the entire model.
[0005] The present invention has been devised to solve the above-mentioned problems, and the objective of the present invention is to provide an individual inference path-based video question-answering device and method using a pre-trained model that encodes audio features and text features from multi-domain input data.
[0006] The problems of the present invention are not limited to those mentioned above, and other unmentioned problems will be clearly understood by those skilled in the art from the description below.
[0007] A video question-answering device based on individual inference paths using a pre-trained model that encodes audio features and text features from input data of multiple domains according to an embodiment of the present invention includes a motion and object feature extraction unit that extracts domain-independent motion features and object features from the audio features and text features; a global expression derivation unit that derives global expressions by processing the extracted motion features and object features through individual inference paths; and an initial answer output unit that outputs an initial answer by summing the similarity between the global expressions derived through the individual inference paths.
[0008] At this time, the input data includes video data, audio data, and text data, wherein the text data may include question-and-answer data and subtitle data.
[0009] Additionally, the action and object feature extraction unit may include a dictionary database comprising a dataset composed of predefined action labels and object labels, a text embeddinger pre-trained to encode the action labels and object labels, and a feature decomposer that calculates the similarity between the embedded action labels and object labels and the audio features and text features to distinguish and extract the domain-independent action features and object features from the audio features and text features.
[0010] In addition, the feature decomposer can calculate the cosine similarity between the audio feature and each of the embedded action label and object label to distinguish and extract audio-action features and audio-object features from the audio feature, and calculate the cosine similarity between the text feature and each of the embedded action label and object label to distinguish and extract text-action features and text-object features from the text feature.
[0011] In addition, the above dictionary database further includes a dataset composed of verb labels and noun labels, and the feature decomposer can further extract subtitle-action features and subtitle-object features from the subtitle data by distinguishing them based on the verb labels and noun labels.
[0012] In addition, the global representation derivation unit can transmit each of the audio-action feature, audio-object feature, text-action feature, text-object feature, subtitle-action feature, and subtitle-object feature to an individual inference path composed of a fully connected layer and an LSTM (Long Short-Term Memory) module, and derive a global representation composed of an audio-action vector, an audio-object vector, a text-action vector, a text-object vector, a subtitle-action vector, and a subtitle-object vector by aggregating from each individual inference path.
[0013] Meanwhile, the individual inference path-based video question-answering device may further include a final answer output unit that outputs a final answer by summing the initial answer to the answer provided by a pre-prepared classifier.
[0014] A video question answering method based on an individual inference path, performed by a video question answering device based on an individual inference path using a pre-trained model that encodes audio features and text features from input data of multiple domains according to an embodiment of the present invention, comprises: a motion and object feature extraction step for extracting domain-independent motion features and object features from the audio features and text features; a global expression derivation step for deriving global expressions by processing the extracted motion features and object features through an individual inference path; an initial answer output step for outputting an initial answer for the input data by summing the similarity between the global expressions derived through the individual inference path; and a final answer output step for outputting a final answer by summing the initial answer with an answer provided by a pre-prepared classifier.
[0015] According to one aspect of the present invention described above, by distinguishing and extracting domain-independent action features and object features from audio features and text features encoded by a pre-trained model, and deriving an answer based thereon, a video question-answering device capable of effectively adapting to various domains without fine-tuning the pre-trained model can be provided.
[0016] The effects of the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description in the claims.
[0017] FIG. 1 is a block diagram illustrating the schematic configuration of a video question-and-answer device according to one embodiment of the present invention.
[0018] FIG. 2 is a block diagram illustrating the specific configuration of an operation and object feature extraction unit according to an embodiment of the present invention.
[0019] FIG. 3 is a block diagram illustrating the specific configuration of an initial answer output unit according to one embodiment of the present invention.
[0020] FIG. 4 is a conceptual diagram illustrating a video question-and-answer device according to one embodiment of the present invention.
[0021] FIG. 5 is a conceptual diagram illustrating an operation and object feature extraction unit according to an embodiment of the present invention.
[0022] FIG. 6 is a conceptual diagram illustrating an initial answer extraction unit according to an embodiment of the present invention.
[0023] FIG. 7 is a flowchart of a video question-and-answer method according to one embodiment of the present invention.
[0024] The following detailed description of the invention refers to the accompanying drawings, which illustrate specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention. It should be understood that various embodiments of the invention are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the invention in relation to one embodiment. It should also be understood that the location or arrangement of individual components within each disclosed embodiment may be changed without departing from the spirit and scope of the invention. Accordingly, the following detailed description is not intended to be limiting, and the scope of the invention is limited only by the appended claims, including all equivalents to those claimed therein, provided they are appropriately described. Similar reference numerals in the drawings refer to the same or similar functions across various aspects.
[0025] The components according to the present invention are defined by functional distinction rather than physical distinction, and can be defined by the functions each performs. Each component may be implemented as hardware or as program code and processing units that perform each function, and the functions of two or more components may be included and implemented in a single component. Therefore, it should be noted that the names assigned to the components in the following embodiments are not intended to physically distinguish each component but are assigned to imply the representative function performed by each component, and that the technical concept of the present invention is not limited by the names of the components.
[0026] Preferred embodiments of the present invention will be described in more detail below with reference to the drawings.
[0027] FIG. 1 is a schematic block diagram of an individual inference path-based video question-answering device (10) according to one embodiment of the present invention.
[0028] A video question-answering device (10) according to one embodiment of the present invention may include a motion and object feature extraction unit (100), a global expression derivation unit (200), an initial answer output unit (300), and a final answer output unit (400). In addition, software (application) for performing an individual inference path-based video question-answering method may be installed and executed on the video question-answering device (10), and the motion and object feature extraction unit (100), the global expression derivation unit (200), the initial answer output unit (300), and the final answer output unit (400) may be controlled by the software (application) for performing an individual inference path-based video question-answering method.
[0029] At this time, the video question-answering device (10) may be a separate terminal or a part module of the terminal. Additionally, the configuration of the operation and object feature extraction unit (100), the global expression derivation unit (200), the initial answer output unit (300), and the final answer output unit (400) may be formed as an integrated module or composed of one or more modules. However, conversely, each configuration may be composed of a separate module.
[0030] Additionally, the video question-and-answer device (10) may be mobile or fixed. This video question-and-answer device (100) may be in the form of a server or an engine, and may be referred to by other terms such as device, apparatus, terminal, user equipment (UE), mobile station (MS), wireless device, or handheld device. Furthermore, the video question-and-answer device (10) may execute or create various software based on an operating system (OS), that is, a system. Here, the operating system is a system program that enables software to use the hardware of the device, and may include all mobile computer operating systems such as Android OS, iOS, Windows Mobile OS, Bada OS, Symbian OS, BlackBerry OS, etc., as well as computer operating systems such as Windows family, Linux family, Unix family, MAC, AIX, HP-UX, etc.
[0031] The pre-trained model (20) can encode audio features and text features from input data of multiple domains.
[0032] At this time, the input data may include video data, audio data, and text data, and the text data may include question-and-answer data and subtitle data.
[0033] The pre-training model (20) can encode the video data, audio data, and question-and-answer data into audio features (Audio Features, D), and encode the video data, subtitle data, and question-and-answer data into text features (Text Features, T).
[0034] At this time, the above question-and-answer data may consist of multiple-choice question information and answer option information. In addition, the above subtitle data may include subtitle information corresponding to the video data and audio data.
[0035] These pre-trained models (20) may be large-scale pre-trained deep learning models such as FrozenBiLM or Merlot Reserve, but are not limited thereto.
[0036] Hereinafter, specific configurations of a video question-and-answer device (10) according to an embodiment of the present invention will be described with reference to FIGS. 2 to 6.
[0037] FIG. 2 is a block diagram illustrating the specific configuration of an operation and object feature extraction unit according to an embodiment of the present invention.
[0038] The motion and object feature extraction unit (100) can extract domain-independent motion features and object features from the audio features (D) and text features (T) encoded by the pre-trained model (10).
[0039] To this end, the motion and object feature extraction unit (100) may include a dictionary database (110), a text embedder (120), and a feature decomposer (130).
[0040] The dictionary database (110) may include a dataset consisting of predefined action labels and object labels.
[0041] And, the text embedder (120) may be prepared by being pre-trained to encode the action label and object label.
[0042] And, the feature decomposer (130) can calculate the similarity between the embedded action label and object label and the audio feature (D) and text feature (T) to extract the domain-independent action feature and object feature from the audio feature (D) and text feature (T).
[0043] More specifically, as illustrated in FIG. 5, the feature decomposer (130) calculates the cosine similarity between the audio feature (D) and each of the embedded action label and object label to derive an audio-action feature ( ) and audio-object features( ) is extracted by distinguishing, and the cosine similarity between the text feature (T) and each of the embedded action label and object label is calculated to obtain text-action features ( ) and text-object features( ) can be distinguished and extracted. At this time, the feature decomposer (130) can identify K action labels and object labels with the highest calculated cosine similarity.
[0044] Meanwhile, the above dictionary database (110) may further include a dataset consisting of verb labels and noun labels.
[0045] And, the feature decomposer (130) derives subtitle-action features ( ) features and subtitle-object features( It is possible to further extract by distinguishing ). These subtitle-action features ( ) and subtitle-object features( ) is extracted based on subtitle data independently input to the feature decomposer (130), so it may be a feature unrelated to the question or answer.
[0046] The global representation derivation unit (200) can derive global representations by processing the extracted operation features and object features through individual inference paths.
[0047] More specifically, as illustrated in FIG. 4, the global expression derivation unit (200) is the audio-operation feature ( ), audio-object features( ), Text-action features( ), text-object features( ), Subtitle-Motion Features( ) and subtitle-object features( Each is passed to individual inference paths composed of fully connected layers (FC) and Long Short-Term Memory (LSTM) modules, and audio-action vectors ( ), audio-object vector( ), text-action vector( ), text-object vector( ), subtitle-motion vector( ) and subtitle-object vector( It can be derived by aggregating global expressions composed of ).
[0048] The output of the motion and object feature extraction unit (100) can be passed to linear layers to form individual path features. The fully connected layer (FC) can be used to encode motion features and object features, which can be represented by Equation 1 below.
[0049] [Mathematical Formula 1]
[0050]
[0051] The above audio-action vector ( ), audio-object vector( ), text-action vector( ), text-object vector( ), subtitle-motion vector( ) and subtitle-object vector( The fully connected layer where ) is transmitted to each is, represents the learning parameters including the weights and biases of each fully connected layer. And represents the k-th action label with the highest cosine similarity, and represents the k-th object label with the highest cosine similarity.
[0052] In addition, the above LSTM module may be a bidirectional LSTM module. The efficient cells and hidden vectors of the LSTM module can maintain and update relevant information, enabling effective learning even in a very limited feature space, and thus are suitable for a compact feature space. A bidirectional LSTM module ( and ) can be used to aggregate and derive global representations for each operation path where operation features are processed and object path where object features are processed, and this can be expressed by the following mathematical formula 2.
[0053] [Mathematical Formula 2]
[0054]
[0055] And the initial answer output unit (300) can output an initial answer (P) by summing the similarity between the global expressions derived through individual inference paths.
[0056] To this end, as illustrated in FIGS. 3 and 6, the initial answer output unit (300) may include a similarity deriver (310), an attention weight deriver (320), and an initial answer output unit (330).
[0057] The similarity deriver (310) is the audio-action vector ( ) and text-action vector( ) respectively the above subtitle-action vector( Deriving similarity by comparing with ), and the above audio-object vector ( ) and text-object vector( ) respectively subtitle-object vector( Similarity can be derived by comparing with ). In this case, the similarity can be calculated through cosine similarity as shown in Equation 3 below.
[0058] The attention weight deriver (320) derives the attention weights to be applied to each of the derived similarities ( ) can be derived.
[0059] More specifically, the attention weight deriver (320) can derive attention weights for each individual inference path using a fully connected layer (FC).
[0060] And the initial answer output device (330) can output an initial answer (P) by multiplying the similarity derived above by the attention weights and summing them. This process can be represented by the following mathematical formulas 3 and 4.
[0061] [Mathematical Formula 3]
[0062]
[0063] [Mathematical Formula 4]
[0064]
[0065] The final answer output unit (400) can output a final answer by adding the initial answer to the answer provided by a pre-prepared classifier.
[0066] Preferably, the classifier utilizes the text features (T) encoded in the prior training model, and a linear hierarchical classifier ( It can be provided as a text classifier that outputs an answer (Q) through ), which can be represented by the following mathematical formula 5.
[0067] [Mathematical Formula 5]
[0068]
[0069] Finally, the probability (F) of the final answer can be calculated by summing the initial answer (P) output based on the individual inference path and the answer (Q) output by the pre-prepared classifier, which can be expressed by the following mathematical formula 6.
[0070] [Mathematical Formula 6]
[0071]
[0072] Meanwhile, a video question-answering device (10) according to another embodiment of the present invention may further include an audio classifier that outputs an answer by utilizing the audio feature (D) encoded in the prior learning model, and may output a final answer by adding at least one of the answer output from the text classifier and the answer output from the audio classifier to the initial answer.
[0073] FIG. 7 is a flowchart illustrating an individual inference path-based video question-answering method according to an embodiment of the present invention. Since the video question-answering method according to an embodiment of the present invention proceeds on substantially the same configuration as the video question-answering device (10) shown in FIG. 1, the same reference numerals are assigned to the same components as those in the video question-answering device (10) of FIG. 1, and repetitive descriptions are omitted.
[0074] A video question answering method based on an individual inference path, performed by a video question answering device (10) based on an individual inference path using a pre-learned model (20) that encodes audio features and text features from input data of multiple domains according to an embodiment of the present invention, may include a motion and object feature extraction step (S100) for extracting domain-independent motion features and object features from the audio features and text features; a global expression derivation step (S200) for deriving global expressions by processing the extracted motion features and object features through an individual inference path; an initial answer output step (S300) for outputting an initial answer for the input data by summing the similarity between the global expressions derived through the individual inference path; and a final answer output step (S400) for outputting a final answer by summing the initial answer to an answer provided by a pre-prepared classifier.
[0075] The video question-and-answer method of the present invention as described above can be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either individually or in combination.
[0076] The program instructions recorded on the above-mentioned computer-readable recording medium may be those specifically designed and configured for the present invention, or they may be those known and available to those skilled in the art of computer software.
[0077] Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions such as ROM, RAM, and flash memory.
[0078] Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware device may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.
[0079] Although various embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the invention as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present invention.
Claims
1. A video question answering device based on individual inference paths using a pre-trained model that encodes audio features and text features from multi-domain input data, A motion and object feature extraction unit that extracts domain-independent motion features and object features from the above audio features and text features; A global expression derivation unit that derives global expressions by processing the above-mentioned extracted motion features and object features through individual inference paths; and An individual inference path-based video question answering device comprising: an initial answer output unit that outputs an initial answer by summing the similarities between the global expressions derived through the individual inference paths.
2. In Paragraph 1, The above input data is, An individual inference path-based video question-answering device comprising video data, audio data and text data, wherein the text data includes question-answering data and subtitle data.
3. In Paragraph 2, The above operation and object feature extraction unit is, A dictionary database containing a dataset composed of predefined action labels and object labels; A text embedder pre-trained to encode the above action labels and object labels; and An individual inference path-based video question answering device comprising: a feature decomposer that calculates the similarity between embedded action labels and object labels and the audio features and text features to distinguish and extract the domain-independent action features and object features from the audio features and text features.
4. In Paragraph 3, The above feature decomposer is, Calculate the cosine similarity between the above audio feature and each of the above embedded action label and object label to distinguish and extract audio-action features and audio-object features from the above audio feature, and An individual inference path-based video question answering device that calculates the cosine similarity between the text feature and each of the embedded action label and object label to distinguish and extract text-action features and text-object features from the text feature.
5. In Paragraph 4, The above dictionary database is, It further includes a dataset consisting of verb labels and noun labels, and The above feature decomposer is, A video question answering device based on individual inference paths that further extracts subtitle-action features and subtitle-object features from the subtitle data based on the verb labels and noun labels.
6. In Paragraph 5, The above global expression derivation unit is, Each of the above audio-action features, audio-object features, text-action features, text-object features, subtitle-action features, and subtitle-object features is passed to an individual inference path composed of a fully connected layer and an LSTM (Long Short-Term Memory) module, An individual inference path-based video question answering device that aggregates and derives a global representation composed of an audio-action vector, an audio-object vector, a text-action vector, a text-object vector, a subtitle-action vector, and a subtitle-object vector from each individual inference path.
7. In Paragraph 6, The above initial answer output unit is, A similarity deriver that derives similarity by comparing the above audio-motion vector and text-motion vector with the above subtitle-motion vector, respectively, and derives similarity by comparing the above audio-object vector and text-object vector with the above subtitle-object vector, respectively; An attention weight generator that derives attention weights to be applied to each of the similarities derived above; and An individual inference path-based video question answering device comprising: an initial answer outputter that outputs an initial answer by multiplying each of the derived similarities by the attention weights and summing them up.
8. In Paragraph 1, An individual inference path-based video question-answering device further comprising a final answer output unit that outputs a final answer by summing the initial answer to the answer provided by a pre-prepared classifier.
9. An individual inference path-based video question answering method performed by an individual inference path-based video question answering device using a pre-trained model that encodes audio features and text features from multi-domain input data, wherein A motion and object feature extraction step for extracting domain-independent motion features and object features from the above audio features and text features; A global expression derivation step for deriving global expressions by processing the extracted operation features and object features through individual inference paths; An initial answer output step for outputting an initial answer for the input data by summing the similarities between the global expressions derived through the individual inference paths; and A video question answering method based on individual inference paths, comprising: a final answer output step of outputting a final answer by summing the initial answer to the answer provided by a pre-prepared classifier.