Visual dialogue methods and devices based on dynamic routing interaction and hybrid graph reasoning

By employing a visual dialogue method that combines dynamic routing interaction and hybrid graph reasoning, the heterogeneity gap between multimodal information is addressed, enabling flexible cross-modal interaction and reliable answer reasoning. This enhances the collaborative representation and semantic understanding capabilities of human-computer interaction, and improves the real-time performance and accuracy of the visual dialogue model.

CN117217313BActive Publication Date: 2025-10-28TONGJI UNIV

Patent Information

Application Number
CN202311130501.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-01
Publication Date
2025-10-28
Estimated Expiration
2043-09-01

AI Technical Summary

Technical Problem

Existing technologies struggle to bridge the heterogeneous gap between multimodal information and cannot fully model the semantic relationships between dialogues, leading to limitations in human-computer interaction understanding. Furthermore, existing methods primarily rely on expert knowledge and experience feedback, limiting intermodal interaction to static forms and failing to flexibly bridge the heterogeneous gap between different modalities.

Method used

A visual dialogue method based on dynamic routing interaction and hybrid graph reasoning is adopted. The dynamic routing interaction module realizes the filtering, extraction and alignment of cross-modal features, and the hybrid graph reasoning module performs multi-step historical dialogue semantic association reasoning. The structured reference graph and context-aware temporal graph are constructed and fused into a hybrid graph for multi-step graph reasoning. Finally, the answer to the visual dialogue is inferred through the answer decoder.

Benefits of technology

It improves the collaborative representation capability of human-computer interaction, fully explores semantic dependencies, realizes more comprehensive semantic relationship modeling between multi-turn dialogues, enhances the real-time performance and response speed of visual dialogue models, solves the problem of referential resolution, and enhances the semantic and contextual understanding capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117217313B_ABST
    Figure CN117217313B_ABST
Patent Text Reader

Abstract

This invention relates to a visual dialogue method and device based on dynamic routing interaction and hybrid graph reasoning. The method includes the following steps: acquiring image features and text features; performing cross-modal interaction of filtering, extraction, and alignment based on a dynamic routing interaction module to obtain potentially aligned cross-modal features; performing multi-step historical dialogue semantic association reasoning based on a hybrid graph reasoning module for the cross-modal features to obtain visually guided text features; inputting the text features into a decoder, and obtaining the answer to the visual dialogue through reasoning. Compared with existing technologies, this invention can fully exploit semantic dependencies, has strong collaborative representation capabilities, and provides more accurate and reliable multi-turn visual dialogues.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual reasoning technology, and in particular to a visual dialogue method and device based on dynamic routing interaction and hybrid graph reasoning. Background Technology

[0002] Multimedia data is diverse in type, widely sourced, and complex in its relationships. How to quickly and efficiently utilize multimodal data to serve people is a current hot research topic. Currently, the application of multimodal data is still in its nascent stage, and related theories and technologies are constantly evolving. Further enhancing the automatic perception and understanding capabilities of multimodal data, and exploring the semantic relationships across modal data, is a challenge with significant application and research value. Visual dialogue is an important cross-modal understanding task that enables AI agents to complete multiple rounds of continuous question-and-answer sessions based on visual information and historical dialogue information. It is a crucial component of intelligent human-computer interaction, and this technology can be applied to intelligent customer service interactions, entertainment games, smart homes, healthcare, and intelligent manufacturing, showing broad application prospects.

[0003] In visual dialogue tasks, existing technologies struggle to bridge the heterogeneous gap between multimodal information and adequately model the semantic relationships between dialogues. On one hand, current methods for multimodal interaction largely rely on expert knowledge and experiential feedback, limiting interaction forms to manually defined or static formats, failing to flexibly bridge the heterogeneous gap between different modalities. On the other hand, current work mostly captures semantic features or relationships from a single modality or a single dialogue modeling structure in an attempt to address issues such as coreference resolution in visual dialogue. However, this semantic modeling approach increases the limitations of dialogue semantic understanding, preventing the model from comprehensively and accurately understanding the dialogue intent. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a visual dialogue method and device based on dynamic routing interaction and hybrid graph reasoning, thereby improving human-computer interaction capabilities through flexible cross-modal information interaction and reliable answer reasoning. This invention can achieve its objective through the following technical solutions:

[0005] One aspect of the present invention provides a visual dialogue method based on dynamic routing interaction and hybrid graph reasoning, comprising the following steps:

[0006] Image and text features are acquired, and cross-modal interaction of filtering, extraction, and alignment is performed based on the dynamic routing interaction module to obtain potential aligned cross-modal features;

[0007] For the cross-modal features, multi-step historical dialogue semantic association reasoning is performed based on the hybrid graph reasoning module to obtain visually guided text features;

[0008] The text features are input into the decoder, and the answer to the visual dialogue is obtained through reasoning.

[0009] As a preferred technical solution, the cross-modal interaction based on the dynamic routing interaction module for filtering, extraction, and alignment includes the following steps:

[0010] Based on the image and text features, dynamic interaction between different levels of modalities is achieved using a filtering-extraction-alignment function block with a dynamic router, resulting in cross-modal features for potential alignment.

[0011] As a preferred technical solution, the filtering-extraction-alignment functional block includes:

[0012] The FILTER block is used to filter out redundant and irrelevant visual information.

[0013] The INTRA block is used to obtain contextual information of a local region based on a self-attention mechanism.

[0014] The INTER block is used to achieve implicit alignment between images and text based on a self-attention mechanism.

[0015] As a preferred technical solution, both the INTRA block and the INTER block include an FFN network.

[0016] As a preferred technical solution, the dynamic routing interaction module, the hybrid graph inference module, and the decoder are set in the same single-stage network model.

[0017] As a preferred technical solution, the multi-step historical dialogue semantic association reasoning based on the hybrid graph reasoning module includes the following steps:

[0018] Based on the aforementioned cross-modal features, a structured reference graph and a context-aware temporal graph are constructed.

[0019] The structured reference graph and the context-aware temporal graph are fused into a hybrid graph, and multi-step graph reasoning is performed based on the hybrid graph to obtain visually guided text features.

[0020] As a preferred technical solution, the structured reference graph is used to parse dialogues including pronouns in a sparse manner, the context-aware temporal graph is used to supplement the inherent temporal structure in the dialogue that is destroyed by sparse operations, and the hybrid graph is obtained by fusion through gate operations.

[0021] As a preferred technical solution, the acquisition of image features includes the following steps:

[0022] Image features are extracted based on regression boxes and a pre-trained visual classification network.

[0023] As a preferred technical solution, the acquisition of text features includes the following steps:

[0024] Text features are extracted based on a pre-trained LSTM text classification network.

[0025] In another aspect, an electronic device is provided, comprising: one or more processors and a memory, wherein the memory stores one or more programs, the one or more programs including instructions for executing the above-described visual dialogue method based on dynamic routing interaction and hybrid graph reasoning.

[0026] In another aspect, the present invention provides a computer-readable storage medium including one or more programs executable by one or more processors of an electronic device, said one or more programs including instructions for performing the above-described visual dialogue method based on dynamic routing interaction and hybrid graph reasoning.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] (1) Strong collaborative representation capability: This invention inputs reliable cross-modal interaction features into the visually guided dialogue text structure graph, and then infers the answer through the answer decoder to achieve the purpose of human-like visual dialogue. By adopting a dynamic routing interaction module to generate various flexible cross-modal interaction modes, it overcomes the problem of static and insufficient cross-modal interaction. Functional blocks are designed to capture valuable image features and text features, promote implicit alignment between modalities, and enhance the collaborative representation capability of the visual dialogue model.

[0029] (2) Fully exploit semantic dependencies: This invention uses a hybrid graph reasoning module to infer the structural and temporal relationships between multiple turns of dialogue, thereby overcoming the problem of insufficient mining of potential semantic dependencies. This module provides more comprehensive semantic relationships between multiple turns of dialogue and solves the problem of referential resolution to a certain extent, thus promoting the reasoning ability of the visual dialogue model.

[0030] (3) Fine-grained association of information: This invention employs a multi-head attention mechanism, which simultaneously focuses on different attention subspaces to achieve more comprehensive and accurate information interaction and representation. The multi-head attention mechanism can model and associate dialogue text and visual features in a more fine-grained manner, thereby improving the model's ability to understand semantics and context.

[0031] (4) Effectively improve reasoning efficiency: The present invention adopts a unified single-stage network model, which integrates different processing steps into one model, avoiding complex multi-stage processes and information transmission problems between modules, and improving the real-time performance and response speed of the visual dialogue system. Attached Figure Description

[0032] Figure 1 This is a flowchart of the visual dialogue method based on dynamic routing interaction and hybrid graph reasoning in the embodiment;

[0033] Figure 2 This is a schematic diagram of the unified single-stage visual dialogue framework in the embodiment. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0035] Example 1

[0036] To address the problems of the aforementioned existing technologies, this embodiment provides a visual dialogue method based on dynamic routing interaction and hybrid graph reasoning. This method extracts image region features and text semantic features, and predicts the answer to the current question through a novel end-to-end framework of dynamic routing interaction and hybrid graph reasoning. The end-to-end model includes a feature encoding module, a dynamic routing interaction module, a hybrid graph reasoning module, and an answer decoding module. The feature encoding module is used to extract visual and textual information. The dynamic routing interaction module can automatically determine the optimal modal interaction mode suitable for different questions; this module consists of three interaction function blocks. The hybrid graph reasoning module explores the full semantic association between dialogues from multiple perspectives, where the hybrid graph is formed by fusing a structured referential graph and a context-aware temporal graph. The answer decoder is input into the visually guided text features after multi-step graph reasoning to infer a reliable answer.

[0037] The construction and training process of the dynamic routing interaction module and the hybrid graph inference model includes the following steps:

[0038] Step 1: Extract visual features from the image and text features from the dialogue. Visual features include region detection features, and text features include historical dialogue features and current question features.

[0039] The extraction of visual features specifically involves:

[0040] A large-scale object detection network based on bounding boxes and classification extracts object features from high-confidence regions. A large-scale visual classification network based on pre-training extracts high-level semantic features from images.

[0041] The extraction of text information specifically involves:

[0042] We initialize text semantic features based on a pre-trained large-scale text classification network. Then, we encode high-level text features using a Long Short-Term Memory (LSTM) network.

[0043] Step 2: Based on the Dynamic Interaction Module (DIM), the extracted visual and text features are dynamically routed and interacted across modal information, and then trained and optimized on the visual dialogue dataset.

[0044] The dynamic routing interaction module feeds the extracted visual features of single-image regions and high-level textual features into three cross-modal interaction functional blocks to achieve dynamic interaction between different levels of modality and obtain potentially aligned cross-modal features. The three cross-modal interaction functional blocks are: FILTER for filtering redundant visual information, INTRA for representing local region contextual information through a self-attention mechanism, and INTER for promoting implicit alignment between vision and language. Simultaneously, each functional block is assigned a dynamic router to flexibly select the most effective cross-modal interaction mode.

[0045] Step 3: Based on the Hybrid Graph Reasoning Module (HGR), perform historical dialogue semantic association reasoning on cross-modal features after full interaction via DIM, and train and optimize on the visual dialogue dataset.

[0046] The hybrid graph reasoning module leverages cross-modal features gleaned from extensive interaction with the dynamic routing module to construct a structured reference graph and a context-aware temporal graph. These are then fused into a hybrid graph to perform multi-step iterations for accurate and reliable reasoning. Specifically, the structured reference graph, context-aware temporal graph, and hybrid graph are constructed as follows: a structured reference graph is built to sparse complex dialogues involving pronouns. A context-aware temporal graph is constructed to supplement the inherent temporal structure of the dialogue, which is disrupted by sparse operations. The structured reference graph and the context-aware temporal graph are then fused through gate operations to form a hybrid graph for multi-step graph reasoning.

[0047] Step 4: The visually guided text features obtained through HGR multi-step reasoning are input into the answer decoder for answer reasoning. The visually guided text features obtained through multi-step graph reasoning are fed into the answer decoder to infer the accurate answer.

[0048] like Figure 1As shown, in a specific implementation process, this method can be divided into the following steps:

[0049] S1. Extract image region features and text semantic features.

[0050] The feature encoder encodes visual and textual features into a common vector space to produce higher-level representations suitable for cross-modal interaction. Extracting image region features and textual semantic features includes: extracting visual region detection features, extracting historical dialogue text features, and extracting current question text features.

[0051] In this embodiment, the model first extracts target-level visual features using the Faster-RCNN framework pre-trained on the Visual Genome dataset, and encodes these features using average pooling. In the first round of dialogue, the model initializes the text sentences using the GloVe word embedding method and encodes the words in the current question as word vectors using a Long Short-Term Memory (LSTM) network. Similarly, the historical dialogue text features from the second round are encoded by LSTM. These three feature sets constitute the input for subsequent model processing.

[0052] S2. Construct a dynamic routing interaction module.

[0053] The dynamic routing interaction module provides a variety of more flexible interaction modes for cross-modal interaction. For example... Figure 2 As shown, building a dynamic routing interaction module includes: designing interactive function blocks and executing dynamic routing.

[0054] S21. To facilitate intra-modal and inter-modal interactions, three different types of interaction blocks were designed: FILTER, INTRA, and INTER. Specifically, the purpose of FILTER is to filter out redundant visual features. It takes visually encoded features as input, filters irrelevant visual information through a function, and outputs visual features. The process is shown in the following formula:

[0055]

[0056] The purpose of INTRA is to reflect the potential dependencies between local regions. This embodiment uses self-attention units in the Transformer to facilitate intramodal interactions. The role of the self-attention unit is to collect relevant target-level visual features under the guidance of the current problem. Specifically, given a learnable projection matrix W... q W k and W v Calculate the query matrix M q =XW q Key matrix M k =XWk Sum matrix M v =XW v The scaling factor is calculated using the Softmax(·) function. The dot product attention feature F is obtained, and this process is shown in the following equation:

[0057]

[0058] To facilitate obtaining richer feature representations from multiple subspaces in this example, a multi-head attention mechanism is used to further aggregate the features of h parallel subspaces. This process is shown in the following equation:

[0059] MultiHead(X)=Concat(F1,...,F h W o +X

[0060] Among them, F h The attention feature of the h-th subspace head is represented by Concat(·), which represents the join operation. o It is a linear matrix. Based on the multi-head attention mechanism, after processing by the feedforward network FFN(·), the output features of INTRA are... This can be further expressed as:

[0061]

[0062] The purpose of INTER is to bridge the heterogeneity gap between modalities and obtain rich features after implicit alignment. INTER is designed similarly to INTRA. Specifically, given text features Y (Y:=Q or Y:=H), its visually guided contextual information F′ can be calculated from the following formula:

[0063]

[0064] Among them, M′ q =YW′ q Similar to INTRA, this example still uses a multi-head attention mechanism to mine complementary information between different modalities, which can be represented as:

[0065] MultiHead(X,Y)=Concat(F′1,...F′ h W O′ +X

[0066] Among them, F′ h W is the cross-modal attention feature of the h-th subspace head. O′ This is the corresponding linear matrix. Then, both visual and textual features are fed into the feedforward network FFN(·), which outputs the final cross-modal interaction representation. The process is as follows:

[0067]

[0068] S22. To execute the dynamic routing process, interactive function blocks are stacked in depth and width to establish a complete routing space. After stacking the interactive function blocks and establishing a dense path space, a soft conditional gate function is assigned to each layer to learn a vector representation related to the input. Specifically, given the visual features processed by the i-th interactive functional block of the l-th layer... The vector representation can be obtained by using a soft-conditional gate function consisting of the average pooling function AP(·), the multilayer perceptron function MLP(·), the activation function Tanh(·), and the activation function ReLU(·). The process is shown in the following formula:

[0069]

[0070] Specifically, in activators With the help of [the relevant mechanism], this example can dynamically select the interaction path, promoting efficient cross-modal interaction. The process of executing dynamic routing is shown in the following equation:

[0071]

[0072]

[0073] Among them, O l-1 is the output feature of the interactive function block in layer (l-1), and B is the number of interactive function blocks in each layer, which is set to 3 in this example.

[0074] Finally, the output of the last layer of each interactive functional block is aggregated into visually guided text features, and the calculation process is shown in the following formula:

[0075]

[0076]

[0077] in, The current problem characteristics are guided by visual information. It is the historical dialogue feature under the visual guidance of the r-th round of dialogue. The Flatten(·) function flattens the input data into a vector representation, and LN(·) is the layer normalization function.

[0078] S3. Construct a hybrid graph reasoning module.

[0079] The hybrid graph reasoning module can infer more reliable semantic relationships in dialogue. For example... Figure 2As shown, the hybrid graph reasoning module includes: constructing a structured reference graph, constructing a context-aware sequence graph, and constructing a hybrid graph.

[0080] S31. In order to describe the relationship between multiple inputs from different perspectives, this example establishes two types of graph structures, where the nodes of each graph correspond to each round of dialogue, and the edges represent the semantic relationship between each round of dialogue.

[0081] Specifically, the first step is to construct a structured reference diagram. To implicitly interpret ambiguous pronouns in dialogue, among which, The set of nodes in the graph, ε c Represents the edge set of a graph. Adjacent nodes. and The similarity can be calculated as Where k and k′ represent the index of the graph node, z kk′ ∈{0,1} are the coefficients obtained after the sparsification operation.

[0082] S32. Next, construct a context-aware sequence graph. The aim is to alleviate the limitations of single-view semantic relation modeling and to supplement the temporal information destroyed by sparsification operations. The set of nodes in the graph, ε t The edge set of a graph, and the adjacency matrix A of a context-aware temporal graph. t It is constructed by sequentially connecting each round of dialogue.

[0083] S33. Finally, the relationships between nodes in different types of graph structures are propagated through message passing mechanisms, and the contextual features of nodes are updated through graph neural networks. After multi-step graph inference, two types of outputs are generated: the output of the structured reference graph. and the output of context-aware sequence graphs This example uses gate operations to further fuse different types of graph structures. The new graph after fusion via gate operations is called a hybrid graph, and the output feature X of the hybrid graph is... hybrid It can be calculated using the following formula:

[0084]

[0085]

[0086] Where [;] represents a connection operation, δ(·) represents the activation function, and W g and W h It is a learnable parameter matrix, where ⊙ denotes element-wise multiplication.

[0087] S4. Use the answer decoder to deduce the answer.

[0088] To fully utilize the visual-text features in hybrid graph reasoning, a cross-modal answer decoder is employed to infer reasonable answers, such as... Figure 2 As shown. In this example, a discriminative answer decoder was selected, and the probability distribution of the candidate answers can be calculated using the following formula:

[0089]

[0090] in, The candidate answer features are obtained through shared LSTM encoding. is the hybrid graph node feature of the r-th round of dialogue, and p is the posterior probability of each candidate answer.

[0091] During the training phase, this example employs multiple loss functions to jointly optimize all modules and infer reliable answers. Specifically, the standard cross-entropy loss function is used first, as shown in the following equation:

[0092]

[0093] Among them, y a It is the encoded vector of the true answer, p a It is the categorical distribution of candidate answers.

[0094] Then, a routing regularization method is used to guide the selection of dynamic routing paths and ensure the consistency of semantic paths. Specifically, given an input image I... x Its path vector This is obtained by concatenating the outputs of all routers in the dynamic routing module. Then, this example calculates the input image I. x With other images (such as I) y The semantic similarity between images is used to measure the similarity of interaction paths between these images. This process can be represented by the following formula:

[0095]

[0096] in, It is a collection of images, and H0 is the semantic feature of each image. Finally, all loss functions are fused together through a balancing parameter λ, and the final loss function used during training is...

[0097] S5. Evaluate the overall performance of the visual dialogue model based on retrieval metrics.

[0098] To verify the performance of the method in this application, the following experiments were designed.

[0099] Similar to experiments with other methods, the HGDI method of this invention was validated on the visual dialogue benchmark datasets VisDial v0.9 and VisDial v1.0, and its performance was evaluated in a discriminative setting.

[0100] The visual dialogue model is trained using a benchmark dataset. The specific process of training the visual dialogue model includes:

[0101] S51: Given an image, a question, a historical dialogue, and 100 candidate answers;

[0102] S52: Extract visual and textual features using a feature encoder;

[0103] S53: Input the visual and text features from step S52 into the dynamic routing module to obtain cross-modal interaction features;

[0104] S54: Input the cross-modal interaction features from step S53 into the hybrid graph inference module to obtain the visually guided text features after multi-step graph inference;

[0105] S55: Input the visually guided text features from step S54 into the answer decoder, and train the modules in steps S53 and S54 end-to-end under the action of the loss function to obtain the reasoning answer;

[0106] S56: Repeat steps S51, S52, S53, S54 and S55 multiple times to complete the multi-turn dialogue answer reasoning for all images in the dataset.

[0107] The evaluation metrics for visual dialogue models are Mean Rank, Mean Reciprocal Rank (MRR), Recall (R@k, k = 1, 5, 10), and Normalized Discounted Cumulative Gain (NDCG). Mean Rank, Mean Reciprocal Rank, and Recall are based on a single relevant answer, while NDCG is based on multiple semantically similar and correct results. A lower Mean Rank value is better, while higher values ​​for the other evaluation metrics are better. It is worth noting that NDCG is only applicable to the VisDial v1.0 dataset and is used to evaluate the generalization ability of the dialogue model, while MRR evaluates the accuracy of the inferred answers. Average, as the average of MRR and NDCG, is used to comprehensively consider the model's generalization and accuracy. The experimental results of the method in this invention on the VisDial v0.9 and VisDial v1.0 datasets are shown in Tables 1 and 2.

[0108] Table 1 shows the performance of different models on the VisDial v0.9 dataset.

[0109]

[0110]

[0111] Table 2 shows the performance of different models on the VisDial v1.0 dataset.

[0112]

[0113] Compared with the prior art, the present invention has the following beneficial effects:

[0114] I. This invention proposes a dynamic routing interaction module to generate various flexible cross-modal interaction modes to solve the problem of static and insufficient cross-modal interaction. It designs three functional blocks to capture valuable visual-text features, promote implicit alignment between modalities, and enhance the collaborative representation capability of the visual dialogue model.

[0115] Second, this invention proposes a hybrid graph reasoning module to infer the structural and temporal relationships between multi-turn dialogues, thereby addressing the problem of insufficient mining of potential semantic dependencies. This module provides a more comprehensive understanding of semantic relationships between multi-turn dialogues and, to some extent, solves the problem of referential resolution, thus enhancing the reasoning capabilities of visual dialogue models.

[0116] Third, this invention designs a unified single-stage framework to realize visual dialogue, which can realize a single-stage cross-modal interaction-reasoning mode. Reliable cross-modal interaction features are input into the dialogue text structure graph under visual guidance, and then the answer is inferred through the answer decoder to achieve the purpose of human-like visual dialogue.

[0117] Example 2

[0118] This embodiment provides a visual dialogue device (i.e., an electronic device), including: one or more processors and a memory, wherein the memory stores one or more programs, the one or more programs including instructions for performing a visual dialogue method as described in Embodiment 1, which involves dynamic routing interaction and hybrid graph reasoning.

[0119] Example 3

[0120] This embodiment provides a computer-readable storage medium including one or more programs executable by one or more processors of an electronic device, the one or more programs including instructions for performing a visual dialogue method as described in Embodiment 1, which involves dynamic routing interaction and hybrid graph reasoning.

[0121] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A visual dialogue method based on dynamic routing interaction and hybrid graph reasoning, characterized in that, Includes the following steps: Image and text features are acquired, and cross-modal interaction of filtering, extraction, and alignment is performed based on the dynamic routing interaction module to obtain potential aligned cross-modal features; For the cross-modal features, multi-step historical dialogue semantic association reasoning is performed based on the hybrid graph reasoning module to obtain visually guided text features; The text features are input into the decoder, and the answer to the visual dialogue is obtained through reasoning. The cross-modal interaction based on the dynamic routing interaction module for filtering, extraction, and alignment includes the following steps: Based on the aforementioned image and text features, a filtering-extraction-alignment function block with a dynamic router is used to achieve dynamic interaction between different levels of modalities, thereby obtaining cross-modal features for potential alignment. The filtering-extraction-alignment function block includes: The FILTER block is used to filter out redundant visual information; The INTRA block is used to obtain contextual information of a local region based on a self-attention mechanism. The INTER block is used to achieve implicit alignment between images and text based on a self-attention mechanism. The multi-step historical dialogue semantic association reasoning based on the hybrid graph reasoning module includes the following steps: Based on the aforementioned cross-modal features, a structured reference graph and a context-aware temporal graph are constructed. The structured reference graph and the context-aware temporal graph are fused into a hybrid graph, and multi-step graph reasoning is performed based on the hybrid graph to obtain visually guided text features.

2. The visual dialogue method based on dynamic routing interaction and hybrid graph reasoning according to claim 1, characterized in that, The dynamic routing interaction module, hybrid graph inference module, and decoder are all housed in the same single-stage network model.

3. The visual dialogue method based on dynamic routing interaction and hybrid graph reasoning according to claim 1, characterized in that, The structured reference graph is used to parse dialogues including pronouns in a sparse manner; the context-aware temporal graph is used to supplement the inherent temporal structure in the dialogue that is destroyed by sparse operations; and the hybrid graph is obtained by fusion through gate operations.

4. The visual dialogue method based on dynamic routing interaction and hybrid graph reasoning according to claim 1, characterized in that, The acquisition of image features includes the following steps: Image features are extracted based on regression boxes and a pre-trained visual classification network.

5. A visual dialogue method based on dynamic routing interaction and hybrid graph reasoning according to claim 1, characterized in that, The acquisition of the text features includes the following steps: Text features are extracted based on a pre-trained LSTM text classification network.

6. An electronic device, characterized in that, include: One or more processors and a memory, wherein the memory stores one or more programs, the one or more programs including instructions for executing the visual dialogue method based on dynamic routing interaction and hybrid graph reasoning as described in any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, It includes one or more programs that are executed by one or more processors of an electronic device, the one or more programs including instructions for performing the visual dialogue method based on dynamic routing interaction and hybrid graph reasoning as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Rich semantic dialogue generation method fusing visual situation

    CN115964467A

  • News event search method and system based on multi-level image-text semantic alignment model

    WO2023093574A1

Cited By

  • Pay-off crossing frame and pay-off construction method

    CN121216299A