Human-machine multi-turn interaction method and device for visual image

Through the visual image and text feature update system with optimal matching mechanism and cross-attention mechanism, the problem of insufficient semantic information inheritance in multi-round interactions is solved, the semantic clue tracking and dynamic understanding of visual focus of the multi-round dialogue system are improved, and more stable wide-area visual understanding is achieved.

CN120353959BActive Publication Date: 2025-10-10启元实验室
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510855555.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-10
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

Existing methods for fusing visual images and language cannot effectively inherit the semantic information of historical rounds during multi-round interactions, resulting in unstable performance of the model when processing complex reasoning and continuous question-answering, semantic drift or reference errors, and a lack of a feature update mechanism for dynamic context evolution.

Method used

A bimodal contextual feature updating system for text and visual images based on the optimal matching mechanism is constructed. Through local image feature extraction, updating and cross-attention mechanism, cross-round visual attention area extraction, fusion and updating are realized. Combined with the global-local fusion strategy, the ability to track semantic clues and the dynamic understanding of visual focus are improved.

Benefits of technology

It significantly improves the performance of multi-round dialogue systems in wide-area visual understanding, enhances target consistency and context coherence, alleviates the heterogeneity between multimodal inputs, and improves the efficiency of semantic fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353959B_ABST
    Figure CN120353959B_ABST
Patent Text Reader

Abstract

The application relates to a man-machine multi-round interaction method for visual images, comprising the following steps: extracting local image features related to current multi-round dialogue text features from global image information; updating the local image features according to current historical local image features to obtain updated local image features; adopting a cross attention mechanism to determine visual image features according to the updated local image features and global image features corresponding to the global image information; and inputting the visual image features into a multi-modal large model for processing. The application constructs a text and visual image double-modal context feature updating system based on an optimal matching mechanism, can have the abilities of being updateable, compressible and fusable in both text and image modes, significantly improves the tracking ability of a model to semantic clues and the dynamic understanding ability of the model to visual focus in multi-round dialogue, and promotes performance breakthroughs of a multi-round image-text dialogue system in wide-area visual understanding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the theoretical and system research fields of large-scale edge sensing and computing models for wide-area scene understanding, and in particular to a multi-round human-computer interaction method and device for visual images. Background Art

[0002] Multi-round human-computer interaction refers to the process of multiple rounds of question-and-answer or command interaction between a user and an intelligent system through continuous rounds of language or voice input. During this process, the system must not only understand the current input but also reference the history of one or more previous rounds of interaction to achieve context-consistent semantic reasoning and response generation. Compared to single-round dialogue, multi-round interaction better meets the task continuity, semantic progression, and referential diversity requirements of real-world application scenarios, placing higher demands on the system's context modeling capabilities and depth of semantic understanding.

[0003] In recent years, research on multimodal dialogue systems has focused on the deep integration of visual images and language, becoming a core area for intelligent perception and reasoning. With the rapid development of large multimodal models (such as pre-trained visual image and language models and multi-round image-text question-answering systems), they have demonstrated strong capabilities in scenarios such as wide-area visual image understanding, semantic question-answering, object localization, and task reasoning. However, in image understanding tasks involving large scenes, multiple entities, and complex contexts, current systems face significant challenges, including fragmented semantic expression, weak contextual awareness, and discontinuous connections between visual images during multi-round interactions.

[0004] Especially in typical scenarios of "wide-area visual images + multi-round human-computer dialogue", such as security monitoring, human-computer collaboration, remote command, interactive retrieval and other applications, natural language queries initiated by users often involve cross-round entity references, fuzzy descriptions or dynamic focus migration, which puts higher requirements on the model's semantic memory, visual understanding and cross-round state updates. Summary of the Invention

[0005] The inventors discovered that existing methods for fusing visual images and language primarily rely on matching text extracted independently from each round of conversation with visual image features. This fails to effectively inherit the semantic information accumulated in previous rounds, resulting in unstable model performance when handling complex reasoning and continuous question-answering, and even causing semantic drift or reference errors. Specifically, existing technologies suffer from the following problems: the visual side relies solely on the current frame image or static region representation, failing to perceive the local image regions that users have focused on historically; and multimodal fusion is mostly based on one-time splicing or simple attention mechanisms, lacking a feature update mechanism tailored to the dynamic evolution of context.

[0006] To address the above problems, the present invention proposes a multi-round human-computer interaction method for visual images, and constructs a dual-modal context feature update system for text and visual images based on the optimal matching mechanism. It can have the capabilities of "updatable, compressible, and fusible" in both text and image modalities, significantly improving the model's ability to track semantic clues and dynamically understand visual focus in multi-round conversations, thereby promoting performance breakthroughs in multi-round text-image dialogue systems in wide-area visual understanding.

[0007] According to a first aspect of the present application, a human-computer multi-round interaction method for visual images is provided, characterized by comprising:

[0008] Extract local image features related to the current multi-round dialogue text features from the global image information;

[0009] Updating the local image features according to the current historical local image features to obtain updated local image features;

[0010] Determining visual image features based on the updated local image features and global image features corresponding to the global image information using a cross-attention mechanism; and

[0011] The visual image features are input into a multimodal large model for processing.

[0012] According to a second aspect of the present application, a human-computer multi-round interaction device for visual images is provided, characterized by comprising:

[0013] An extraction module is used to extract local image features related to the current multi-round dialogue text features from the global image information;

[0014] An acquisition module, configured to update the local image features according to the current historical local image features and acquire the updated local image features;

[0015] a determination module, configured to determine visual image features based on the updated local image features and global image features corresponding to the global image information using a cross-attention mechanism; and

[0016] An input module is used to input the visual image features into a multimodal large model for processing.

[0017] According to a third aspect of the present application, an electronic device is provided, including:

[0018] processor; and

[0019] The memory stores computer instructions, and when the computer instructions are executed by the processor, the processor is caused to perform the method described in the first aspect.

[0020] According to a fourth aspect of the present application, a non-transitory computer storage medium is provided, storing a computer program, which, when executed by multiple processors, enables the processors to execute the method described in the first aspect.

[0021] According to the human-computer multi-round interaction method and device for visual images provided by this application, the extraction, fusion and update of visual focus areas across rounds are effectively realized through the optimal matching mechanism and the cross-attention mechanism, which significantly improves the system's target consistency and context coherence in multi-round tasks; in addition, this application proposes a global-local fusion strategy to achieve detail focus while retaining the overall visual structure, taking into account spatial perception and semantic density; in addition, this application can effectively alleviate the heterogeneity between multimodal inputs and improve semantic fusion efficiency through dual-channel semantic compression and update of text and visual images. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without exceeding the scope of protection required by this application.

[0023] Figure 1 It is a schematic block diagram of a human-computer multi-round interaction system for visual images according to an embodiment of the present application.

[0024] Figure 2 This is a schematic block diagram of a multi-round dialogue information feature updating system based on an optimal matching mechanism according to an embodiment of the present application.

[0025] Figure 3 This is a flowchart of a multi-round human-computer interaction method for visual images according to an embodiment of the present application.

[0026] Figure 4 This is a flowchart of a multi-round human-computer interaction method for visual images according to another embodiment of the present application.

[0027] Figure 5 Schematic diagram of a human-computer multi-round interaction device for visual images according to an embodiment of the present application.

[0028] Figure 6 It is a schematic diagram of a human-computer multi-round interaction device for visual images according to another embodiment of the present application.

[0029] Figure 7 This is a structural diagram of an electronic device provided by this application. DETAILED DESCRIPTION

[0030] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0031] In multi-turn human-computer interaction systems, achieving continuous understanding and semantic tracking of visual information is a key challenge in improving the intelligence of multimodal dialogue systems. Given the complex object distribution and dynamic contextual semantic changes in large-scale imagery, single-shot static perception is no longer sufficient for continuous understanding. This proposed multi-turn human-computer interaction solution for visual images utilizes semantically guided local visual extraction, dynamic cross-turn visual feature updates, and collaborative fusion of global and local features to construct a temporally consistent and context-adaptive visual representation. This approach is particularly suitable for multi-turn human-computer interaction scenarios involving large-scale, wide-area visual images. This solution first extracts semantically relevant local visual regions from the image based on the current multi-turn dialogue textual features. Local visual information from past dialogues is dynamically enhanced through an optimal matching mechanism. Subsequently, the updated local features are fused with global image information to construct a unified, multi-granular visual representation using a cross-attention mechanism. Finally, the fused visual image and textual features are fed into a large multimodal model to facilitate understanding and reasoning on complex downstream tasks, providing multi-turn dialogue systems with continuous, accurate, and context-consistent visual perception capabilities.

[0032] In this application, text features refer to the conversion of textual information in natural language (such as user-entered dialogue statements) into model-processable vector representations, thereby serving as intermediate expressions for semantic modeling, matching, and reasoning. In this application, text features can include semantic information extracted from the language input in the current round, and can also include global semantic representations formed through context fusion and feature updates, which are used to guide visual information extraction, dialogue state modeling, and multimodal feature fusion. Visual image features refer to feature vectors extracted from images or video frames to represent visual content, typically extracted using convolutional neural networks, visual transformers, or other perceptual models. In this application, visual image features can be divided into local visual image features and global visual image features, corresponding to the representation of the local image region of interest to the user in the current dialogue and the semantic representation of the entire image, respectively. After being fused with text features, visual image features can be used to construct a multimodal joint representation to support tasks such as image-text comprehension, target location, region reference, and command execution in multi-round dialogues.

[0033] Figure 1is a schematic block diagram of a human-machine multi-round interaction system for visual images according to an embodiment of the present application. As shown in Figure 1 The system includes a current multi-round dialogue text feature acquisition stage, a local image feature extraction and update stage, an image feature compression and fusion stage, and a multi-modal large model processing stage.

[0034] For the current multi-round dialogue text feature acquisition stage, existing known technologies can be used to extract the current multi-round dialogue text feature from the current text feature and the historical dialogue text feature. For example, the full concatenation method can be used to directly concatenate the historical text to the current input and then uniformly encode; the historical dialogue text feature can be pruned through a rule-based summary, window interception, or pruning strategy based on attention weight, etc.

[0035] In an embodiment, a multi-round dialogue text feature update method based on an optimal matching mechanism can be used to perform semantic enhancement on the text input of the current round. This method can use any context compression encoding mechanism to process the current input through a self-attention mechanism, and then combine the historical text information to perform optimal matching, extract key information and fusion, thereby obtaining a new round of text feature representation containing historical context semantics. The current multi-round dialogue text feature will be used in subsequent visual focus extraction and multi-modal fusion. The multi-round dialogue text feature acquisition method based on the optimal matching mechanism can also be used, which will be described in detail according to Figure 2 .

[0036] In the local image feature extraction and update stage, the system uses a local image correlation extractor to extract the most relevant local region feature from the global image information of the current frame based on the current multi-round dialogue text feature. The extractor performs content-aware focusing on the image by calculating the semantic similarity between the image region and the text representation, thereby forming a local image embedding representation.

[0037] In addition, to enhance the context consistency of the local image representation, a feature update mechanism on the visual side can be further introduced. This mechanism can be a multi-round dialogue image feature update method based on an optimal matching mechanism: the local visual feature extracted in the current round is semantically matched with the visual focus information accumulated in the historical dialogue, the highly relevant historical visual feature is extracted, and the cross-attention mechanism is used for fusion and update. Finally, a local visual state representation with historical context fusion is formed, which has a trackable and inheritable cross-round perception capability. The specific implementation of the multi-round dialogue image feature update method based on the optimal matching mechanism can be understood with reference to Figure 2 .

[0038] During the image feature compression and fusion phase, to improve the global consistency and task expressiveness of visual representation, this application introduces a global image feature fusion mechanism based on local visual updating. This mechanism can bridge the gap between local and global visual features by introducing learnable vectors as fusion guides.

[0039] Specifically, a global visual representation is extracted from the original full-image information and fed into the cross-attention module along with the updated local visual features. Furthermore, a learnable guidance vector can be used to guide the system to focus on key targets or regions within the local area, while supplementing their spatial and structural information within the full-image context, thereby generating a unified fused visual representation. This representation combines local visual focus with the preservation of the overall semantic structure of the image, making it particularly suitable for complex conversational scenarios that require precise focus on multiple target regions within a large image and reasoning based on contextual background.

[0040] During the multimodal large model processing phase, after multiple rounds of feature updates and fusion across text and visual image modalities, the final multimodal embedding (including updated text and image features) is fed into the multimodal large model for downstream processing. The multimodal large model can be a pre-trained cross-modal encoder, a unified Transformer architecture, or a visual language generation model, supporting a variety of downstream tasks such as image-text question answering, object description generation, visual command execution, and interactive search.

[0041] Thanks to the aforementioned update mechanism, the input multimodal embedding can not only reflect the current input semantics, but also dynamically incorporate information from historical rounds. It can also improve the ability to model the context while maintaining compressibility, significantly enhancing the expressiveness and generalization capabilities of large models in complex human-computer dialogue tasks.

[0042] Figure 2 This is a schematic block diagram of a multi-round dialogue information feature updating system based on an optimal matching mechanism according to an embodiment of the present application. Figure 2 The embodiment provides a method for obtaining characteristics of current multi-round dialogue information, which can be applied to Figure 1 In a visual image-based human-computer multi-turn interaction system, this method serves as a specific implementation for obtaining information features of the current multi-turn dialogue. Furthermore, this method for obtaining information features of the current multi-turn dialogue can also be performed independently to achieve corresponding technical effects. This embodiment uses the independent implementation as an example for illustration. In one embodiment, information features may include text features, image features, and / or voice features.

[0043] like Figure 2 The system can mainly include three stages: a stage of preparing turn-based dialogue and historical dialogue features, a stage of optimal matching, and a stage of updating historical information.

[0044] To improve the information compression and fusion capabilities of multi-turn dialogue systems when faced with long contextual information, this application first performs feature enhancement processing on the input text of the current turn and the conversation content of previous historical turns. During the turn-to-conversation and historical conversation feature preparation phase, upon or after receiving the input information of the current turn (which may include text information, image information, audio information, etc.), the system or model determines the information features of the current conversation turn corresponding to the content of the current conversation turn and the information features of the historical conversations corresponding to the current conversation turn (hereinafter referred to as "the information features of the current historical conversation"). Based on the information features of the current conversation turn and the information features of the current historical conversation, it determines the updated features of the current conversation turn and the updated features of the historical conversations, respectively.

[0045] In one embodiment, the information features of the current round of dialogue can be processed according to the self-attention mechanism to determine the updated features of the current round of dialogue.

[0046] The self-attention mechanism module encodes the information features of the current conversation. By building a self-attention structure within a sentence, it captures the interdependencies between the various pieces of information in the current input sentence, such as dependencies between sentences, images, or speech, thereby improving the accuracy of the current information representation. After processing, a representation of the current conversation that understands contextual information is obtained, which determines the updated features of the current conversation.

[0047] In one embodiment, the information features of the current historical conversation can be processed according to the self-attention mechanism to determine the updated features of the historical conversation.

[0048] The self-attention mechanism module acts on the information features of the current historical conversation, uses the self-attention mechanism to mine the important information structure within the historical conversation, extracts and strengthens the fragments that may provide information support for the current round, and enables them to participate in subsequent screening and fusion.

[0049] After determining the updated features of the current round of dialogue and the updated features of the historical dialogue, a fully connected layer can be used to determine the corresponding query vector matrix of the current round of dialogue based on the updated features of the current round of dialogue, and to determine the corresponding historical dialogue key matrix and historical dialogue value matrix based on the updated features of the historical dialogue.

[0050] The self-attention mechanism is a method for dynamically filtering the content of historical rounds of dialogue in multi-round dialogue scenarios. It can compress redundant historical content during the information understanding process and retain only the key information that contributes most to the current dialogue round, thereby improving the system's ability to fine-tune context modeling. It is particularly suitable for scenarios with dense or redundant information in long contexts.

[0051] Through the processing of the above self-attention mechanism module, two enhanced information representations can be obtained: a vector feature used to express the information of this round, and a set of information features containing the context information of historical rounds.

[0052] Next, in order to avoid the interference of redundant content in historical information on current information modeling and to fully utilize potential valuable information, the present invention designs an optimal matching mechanism based on information relevance. The main process of this mechanism may include:

[0053] Information relevance assessment: The system first matches the current round's information representation with each unit in the historical information set, and calculates the information relevance score between them. This score is used to measure the contribution of each piece of historical information to the current information.

[0054] Highly relevant screening: Based on the aforementioned relevance scores, the historical conversation information features are sorted and a preset number of historical segments with the highest scores are screened. These segments are considered the historical information that the current information is most dependent on.

[0055] Constructing a set of valid key-value pairs: The selected historical information is further organized into a "key-value pair" structure for context memory input in the subsequent information fusion module, thereby constructing a highly targeted information context subset.

[0056] This stage achieves effective compression of historical content, avoiding the redundancy and misleading caused by indiscriminately inputting the entire conversation history into the model.

[0057] After filtering key historical information, the system needs to fuse it with the information features of the current round of dialogue to generate a new round of dialogue information representation for downstream tasks or for the next round of dialogue. To this end, this stage uses a cross-attention mechanism to achieve deep fusion of historical and current information. This fusion process mainly includes the following steps:

[0058] Input the information features of this round of conversation as query information;

[0059] Input the filtered historical information key-value pairs as context information;

[0060] Through the cross-attention mechanism, the information features of this round of dialogue are weighted updated, enabling it to extract valuable supplementary information from the historical context.

[0061] The resulting new information representation combines the current input with key points from the previous context. This is a context-compressed representation that combines information integrity with concise representation. This fused feature can be used as input for subsequent tasks, such as text prompts, dialogue state tracking, or response generation modules in multimodal models.

[0062] In summary, this application constructs a complete multi-turn conversation context compression encoding process through three key steps: conversation feature preparation, optimal matching, and historical information updating. This process not only improves the efficiency and accuracy of multi-turn information modeling, but also has good adaptability and scalability, suitable for multiple natural language processing scenarios such as question-answering systems, conversation understanding, virtual assistants, and visual language understanding.

[0063] exist Figure 2 Based on the system shown, a method for updating multi-round dialogue information features based on an optimal matching mechanism can be provided. The method includes the following steps:

[0064] Step 1: According to the information features of the current round of dialogue and the information features of the current historical dialogue, the updated features of the current round of dialogue and the updated features of the historical dialogue are determined respectively.

[0065] In one embodiment, to ensure that the multi-round conversation feature update method based on the optimal matching mechanism can specifically select information from historical data and integrate it into the current round of conversation, it is necessary to first preprocess the information features of the current round of conversation and the information features of the historical conversations. In some embodiments, this preprocessing process primarily utilizes a self-attention mechanism to process the information features of the current round of conversation and the information features of the historical conversations to determine the updated features of the current round of conversation and the updated features of the historical conversations, respectively.

[0066] In one embodiment, the information features may include text features, image features, and / or voice features.

[0067] In a specific embodiment, for each round of dialogue information features, a self-attention mechanism can be used to calculate its own importance, as shown in equation (1):

[0068] (1)

[0069] in, 、 and Represent the information characteristics of this round of dialogue The query matrix, key matrix and value matrix of is the scaling factor of the feature dimension, Indicates the updated features of this round of dialogue. represents the normalized exponential function, and T represents the matrix transpose.

[0070] In a specific embodiment, for the information features of the current historical conversation, a self-attention mechanism can be used to calculate its own importance, as shown in equation (2):

[0071] (2)

[0072] in, 、 and Respectively represent the information features of the current historical conversation The query matrix, key matrix and value matrix of is the scaling factor of the feature dimension, Indicates the historical conversation update feature, represents the normalized exponential function, and T represents the matrix transpose.

[0073] In an optional embodiment, step one may include:

[0074] Process the information features of this round of dialogue using the self-attention mechanism to determine the updated features of this round of dialogue; and

[0075] The information features of the current historical conversation are processed according to the self-attention mechanism to determine the updated features of the historical conversation.

[0076] Step 2: determining a corresponding query vector matrix for this round of dialogue based on the updated features of this round of dialogue;

[0077] Step three: determining the corresponding historical conversation key matrix and historical conversation value matrix according to the historical conversation update feature.

[0078] In one embodiment, Figure 1 As shown in the figure, after determining the updated features of the current round of dialogue and the updated features of the historical dialogue, a fully connected layer can be used to determine the corresponding query vector matrix of the current round of dialogue based on the updated features of the current round of dialogue, and to determine the corresponding historical dialogue key matrix and historical dialogue value matrix based on the updated features of the historical dialogue.

[0079] Step 4: Based on the query vector matrix of the current round of conversation and the historical conversation key matrix, an optimal matching mechanism is adopted to obtain, from the historical conversation key matrix and the historical conversation value matrix, a filtered historical conversation key matrix and a filtered historical conversation value matrix whose correlation with the information features of the current round of conversation meets preset conditions.

[0080] In order to efficiently integrate the information features of this round of conversation Information features that dialogue with current history , an optimal matching mechanism is introduced to obtain the historical conversation key vector and historical conversation value vector from the historical conversation key matrix and the historical conversation value matrix, whose correlation with the information features of the current round of conversation meets the preset conditions, so as to screen and strengthen the key historical features.

[0081] In one embodiment, a correlation score is calculated between each vector in the historical conversation key matrix and the query vector matrix of the current conversation. Based on this score, a preset number of the most informative historical conversation features are selected. A preset number of historical conversation key vectors with the highest correlation scores can be determined from the historical conversation key matrix. Subsequently, key-value pairs corresponding to these high-weighted historical features are extracted. Specifically, a preset number of historical conversation value vectors corresponding to the preset number of historical conversation key vectors can be determined. The key-value pairs corresponding to the high-weighted historical features can be incorporated into subsequent processing steps to ensure that the model focuses on the most informative contextual information, thereby enhancing conversation comprehension capabilities.

[0082] In a specific embodiment, in order to effectively integrate the information features of the current historical conversation, a cross-round feature correlation evaluation method can be used, as shown in equation (3):

[0083] (3)

[0084] in, and Indicates that the features are updated by this round of dialogue Update features of historical dialogue The generated query vector matrix for this round of dialogue and the historical dialogue key matrix, represents the normalized exponential function, T represents the matrix transpose, and S represents the relevance score matrix.

[0085] History Conversation Key Matrix Typically, multiple historical conversation key vectors are included. Accordingly, the relevance score matrix S includes multiple scores. Based on the relevance score matrix S, a preset number of historical conversation key vectors with the highest relevance scores are determined from the historical conversation key matrix. After determining the preset number of historical conversation key vectors, the corresponding preset number of historical conversation value vectors are determined based on the correspondence between the historical conversation key vectors and the historical conversation value vectors, thereby forming a historical key-value pair matrix, which can be marked as and ,in, Represents the filtered historical conversation key matrix, Represents the filtered historical conversation value matrix.

[0086] In an optional embodiment, step four may include:

[0087] Determining, based on the query vector matrix of the current conversation and the historical conversation key matrix, a relevance score corresponding to each vector in the historical conversation key matrix;

[0088] Determining a preset number of historical conversation key vectors having the highest correlation scores from the historical conversation key matrix based on the correlation scores, and generating the filtered historical conversation key matrix; and

[0089] The preset number of historical conversation value vectors corresponding to the preset number of historical conversation key vectors are determined to generate the filtered historical conversation value matrix.

[0090] This application introduces an optimal matching mechanism to dynamically filter and compress multi-round historical information, significantly improving the accuracy and efficiency of information comprehension in multi-round dialogue systems over long historical scenarios. Compared to traditional splicing input or window-based compression strategies, this application uses relevance scoring and filtering to eliminate redundant and interfering content, retaining only key information, resulting in highly robust information compression capabilities.

[0091] Step 5: Obtain updated dialogue information features based on the current dialogue query vector matrix, the filtered historical dialogue key matrix, and the filtered historical dialogue value matrix.

[0092] In one embodiment, after obtaining the historical key-value pair matrix, updated conversation information features may be obtained based on the current conversation query vector matrix, the filtered historical conversation key matrix, and the filtered historical conversation value matrix.

[0093] In a specific embodiment, a cross-attention mechanism can be used to update the information features of this round of dialogue, as shown in equation (4):

[0094] (4)

[0095] in, Indicates that the feature is updated by this round of dialogue The generated query vector matrix for this round of dialogue, Represents the filtered historical conversation key matrix, Represents the filtered historical dialogue value matrix, is the scaling factor of the feature dimension, represents the normalized exponential function, T represents the matrix transpose, Represents the updated conversation information feature, which combines the information features of historical conversations and the information features of the current round of conversations and can be used for subsequent information feature expression, such as visual models.

[0096] In an optional embodiment, step five may include:

[0097] Based on the query vector matrix of the current round of dialogue, the filtered historical dialogue key matrix, and the filtered historical dialogue value matrix, a cross-round cross-attention mechanism is adopted to determine updated dialogue information features.

[0098] This application uses cross-attention to achieve deep alignment of current round and historical information, improving the context consistency modeling capability.

[0099] According to another embodiment, the method for updating multi-round dialogue information features based on the optimal matching mechanism may further include:

[0100] Step 6: In response to receiving the input information, determining the information characteristics of the current round of dialogue; and

[0101] Step 7: Determine the updated conversation information feature as the information feature of the current historical conversation, and return to step 1.

[0102] In one embodiment, upon or after receiving the information input for the current round (e.g., including text information, image information, or voice information), the system or model may determine the information features of the current round of conversation corresponding to the conversation information of the current round through feature extraction processing, and determine the updated conversation information features determined in step five as the information features of the current historical conversation, and return to step one for processing.

[0103] In this application, the updated conversation information features determined in this round are used as the information features of the current historical conversation in the next round of conversation, avoiding repeated modeling of all historical information, significantly reducing computing costs, and being suitable for large-scale deployment and multi-task sharing.

[0104] exist Figure 1 Based on the human-computer multi-round interaction system for visual images shown in FIG, according to one aspect of the present application, a human-computer multi-round interaction method for visual images is provided, such as Figure 3 As shown, the method includes the following steps:

[0105] Step S301: extracting local image features related to the text features of the current multi-round dialogue from the global image information.

[0106] In accordance with e.g. Figure 2 After obtaining the current multi-round dialogue text features, the method shown can use the local image related feature extractor to extract the global image information. Extract the local image features that are most relevant to the current conversation semantics, as shown in Equation (5):

[0107] (5)

[0108] in, represents the local feature extractor, Represents global image information, Represents the current multi-round dialogue text features, Represents local image features.

[0109] In one specific implementation, a method for extracting local image features that are most relevant to the current conversation semantics from global image information may include extracting multiple features from the global image information, then performing correlation calculations on text features and image features, and extracting a preset number of image features with the highest correlation values ​​as local visual features.

[0110] In an optional embodiment, step 301 may include:

[0111] performing correlation calculations on the multiple image features corresponding to the global image information and the text features of the current multiple rounds of conversations, respectively, to determine multiple correlation values ​​accordingly; and

[0112] A preset number of image features with the largest correlation values ​​among the multiple image features are determined as local image features related to the current multi-round dialogue text features.

[0113] Those skilled in the art will appreciate that any existing method may be used to extract local features, and this application does not impose any limitation thereto.

[0114] Step S302: updating the local image features according to the current historical local image features to obtain updated local image features.

[0115] In a specific embodiment, the local visual features are updated using an optimal matching mechanism to fuse historical visual information, as shown in Equation (6):

[0116] (6)

[0117] in, Represents a multi-round dialogue image feature update function or network based on the optimal matching mechanism, Represents the current historical local image features, represents local image features, Represents the updated local image features.

[0118] In one embodiment, it is possible to use Figure 2 The method shown in the multi-round dialogue information feature update system based on the optimal matching mechanism updates local image features. The information features are image features. Accordingly, step S302 may include the following steps:

[0119] (1) Determine the updated features of the current round of dialogue and the updated features of the historical dialogue based on the local image features and the current historical local image features;

[0120] (2) Determine the corresponding query vector matrix of this round of dialogue based on the updated features of this round of dialogue;

[0121] (3) Determining the corresponding historical conversation key matrix and historical conversation value matrix based on the historical conversation update features;

[0122] (4) Based on the current round conversation query vector matrix and the historical conversation key matrix, an optimal matching mechanism is employed to obtain, from the historical conversation key matrix and the historical conversation value matrix, a filtered historical conversation key matrix and a filtered historical conversation value matrix whose correlation with the local image features satisfies a preset condition; and

[0123] (5) Obtaining the updated local image features based on the current round dialogue query vector matrix, the filtered historical dialogue key matrix, and the filtered historical dialogue value matrix.

[0124] According to an optional embodiment, step (4) may include:

[0125] Determining, based on the query vector matrix of the current conversation and the historical conversation key matrix, a relevance score corresponding to each vector in the historical conversation key matrix;

[0126] Determining a preset number of historical conversation key vectors having the highest correlation scores from the historical conversation key matrix based on the correlation scores, and generating the filtered historical conversation key matrix; and

[0127] The preset number of historical conversation value vectors corresponding to the preset number of historical conversation key vectors are determined to generate the filtered historical conversation value matrix.

[0128] According to an optional embodiment, after obtaining the updated local image features, step S302 may further include:

[0129] (6) in response to receiving the input image information, determining the image features of the current round of conversation; and

[0130] (7) Determine the updated local image feature as the current historical local image feature, and return to step (1).

[0131] Those skilled in the art will appreciate that any existing method may be used to obtain updated local image features, and this application does not impose any limitation thereto.

[0132] Step S303 , using a cross-attention mechanism, determines visual image features based on the updated local image features and the global image features corresponding to the global image information.

[0133] In one embodiment, to enhance the overall visual understanding capability of the dialogue system, the global image features corresponding to the global image information and the updated local image features can be further fused. The fusion process can adopt a cross-attention mechanism to ensure effective integration of important information, as shown in equation (7):

[0134] (7)

[0135] wherein, represents a cross-attention mechanism or function for modeling the semantic relationship between the global image features and the updated local image features, represents the global image features, represents the updated local image features, represents the visual image features.

[0136] In one optional embodiment, the fusion process can adopt a cross-attention mechanism and a learnable vector to ensure effective integration of important information, as shown in equation (8):

[0137] (8)

[0138] wherein, represents a learnable vector.

[0139] In one optional embodiment, step S303 can include:

[0140] determining visual image features according to the updated local image features, the global image features, and a learnable vector using a cross-attention mechanism.

[0141] Step S304, inputting the visual image features into a multi-modal large model for processing.

[0142] In one embodiment, the current multi-turn dialogue text features and visual image features are sent into a multi-modal large model for downstream tasks, as shown in equation (9):

[0143] (9)

[0144] wherein, represents a multi-modal large model, represents the current multi-turn dialogue text features, represents the visual image features, represents the output of the multi-modal large model.

[0145] The output of the present application can be directly connected to existing multi-modal pre-training frameworks or large model systems, with good versatility and landing ability.

[0146] Figure 4This is a flowchart of a multi-round human-computer interaction method for visual images according to another embodiment of the present application. Figure 3 compared to, Figure 4 Steps S401 to S404 of the method shown are the same as Figure 3 Steps S301 to S304 are the same except that: Figure 4 The illustrated method further includes:

[0147] Step S405, in response to receiving the input text, determining the text features of the current multi-round dialogue; and

[0148] Step S406: Determine the visual image feature as the current historical local image feature, and return to step S401.

[0149] In one embodiment, when or after receiving the text or conversation content input in the current round, the system or model determines the current multi-round conversation text features corresponding to the conversation content of the current round, and determines the visual image features determined in step S403 as the current historical local image features, and returns to step S401 for processing.

[0150] In this application, the updated visual image features determined in this round are used as the current historical local image features for the next round of dialogue, avoiding repeated modeling of all historical information, significantly reducing computational costs, and being suitable for large-scale deployment and multi-task sharing.

[0151] According to another aspect of the present application, a human-computer multi-round interaction device for visual images is provided, such as Figure 5 As shown, the device includes: an extraction module 501, an acquisition module 502, a first determination module 503 and an input module 504, wherein the extraction module 501 is used to extract local image features related to the current multi-round dialogue text features from the global image information; the acquisition module 502 is used to update the local image features according to the current historical local image features and obtain the updated local image features; the first determination module 503 is used to adopt a cross-attention mechanism to determine the visual image features according to the updated local image features and the global image features corresponding to the global image information; the input module 504 is used to input the visual image features into the multimodal large model for processing.

[0152] In an optional embodiment, the extraction module 501 may be used to:

[0153] performing correlation calculations on the multiple image features corresponding to the global image information and the text features of the current multiple rounds of conversations, respectively, to determine multiple correlation values ​​accordingly; and

[0154] A preset number of image features with the largest correlation values ​​among the multiple image features are determined as local image features related to the current multi-round dialogue text features.

[0155] In an optional embodiment, the acquisition module 502 may include:

[0156] A first determination submodule is configured to determine, based on the local image features and the current historical local image features, the updated features of the current round of conversation and the updated features of the historical conversation respectively;

[0157] A second determination submodule is configured to determine a corresponding query vector matrix for this round of dialogue based on the updated features of this round of dialogue;

[0158] A third determination submodule is configured to determine a corresponding historical conversation key matrix and a historical conversation value matrix according to the historical conversation update feature;

[0159] a first acquisition submodule, configured to employ an optimal matching mechanism based on the current-round conversation query vector matrix and the historical conversation key matrix to acquire, from the historical conversation key matrix and the historical conversation value matrix, a filtered historical conversation key matrix and a filtered historical conversation value matrix whose correlation with the local image features satisfies a preset condition; and

[0160] The second acquisition submodule is used to obtain the updated local image features according to the current round dialogue query vector matrix, the filtered historical dialogue key matrix and the filtered historical dialogue value matrix.

[0161] In an optional embodiment, the first acquisition submodule may be configured to:

[0162] Determining, based on the query vector matrix of the current conversation and the historical conversation key matrix, a relevance score corresponding to each vector in the historical conversation key matrix;

[0163] Determining a preset number of historical conversation key vectors having the highest correlation scores from the historical conversation key matrix based on the correlation scores, and generating the filtered historical conversation key matrix; and

[0164] The preset number of historical conversation value vectors corresponding to the preset number of historical conversation key vectors are determined to generate the filtered historical conversation value matrix.

[0165] In an optional embodiment, the acquisition module 502 may further include:

[0166] A fourth determining submodule is configured to determine image features of the current round of conversation in response to receiving the input image information; and

[0167] A fifth determining sub-module is configured to determine the updated local image feature as a current historical local image feature.

[0168] In an optional embodiment, the first determining module 503 can be configured to:

[0169] The cross-attention mechanism is adopted to determine the visual image feature according to the updated local image feature, the global image feature, and a learnable vector.

[0170] Figure 6 FIG. 1 is a schematic diagram of a human-machine multi-round interaction device for visual images according to another embodiment of the present application. Compared with the device shown in FIG. 1, Figure 5 the device shown in FIG. 1 further comprises: Figure 6 The modules 601-604 of the device shown in FIG. 1 are the same as the modules 501-504 of the device shown in FIG. 1, except that: Figure 5 The device shown in FIG. 1 further comprises: Figure 6 A second determining module 605 is configured to determine a current multi-round dialogue text feature in response to receiving an input text; and

[0171] A third determining module 606 is configured to determine the visual image feature as a current historical local image feature.

[0172] According to the human-machine multi-round interaction method and device for visual images provided in the present application, the optimal matching mechanism and the cross-attention mechanism are adopted to effectively realize extraction, fusion, and updating of visual attention areas across rounds, significantly improve the target consistency and context coherence of the system in multi-round tasks, and have strong dynamic visual focus modeling capability. In addition, the global-local fusion strategy is adopted to reserve the overall visual structure while realizing detail focusing, taking into account spatial perception and semantic density, and realizing local and global semantic integration. Furthermore, the text and visual image double-channel semantic compression updating is adopted to effectively alleviate the heterogeneity between multi-modal inputs, improve the semantic fusion efficiency, and realize cross-modal alignment precision. The output of the present application can be directly connected to the existing multi-modal pre-training framework or large model system, has good universality and landing ability, and can be compatible with mainstream large model interfaces.

[0173] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0174]

[0175] ​It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.

[0176] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical connection or other forms.

[0177] See Figure 7 , Figure 7 An electronic device is provided, comprising a processor and a memory. The memory stores computer instructions or one or more programs. When the computer instructions or one or more programs are executed by the processor, the processor executes the computer instructions to achieve the following Figure 3 and Figure 4 The method and refinement scheme shown.

[0178] It should be understood that the above-described device embodiments are merely illustrative, and the devices disclosed herein may also be implemented in other ways. For example, the division of units / modules described in the above-described embodiments is merely a logical functional division, and actual implementations may employ alternative divisions. For example, multiple units, modules, or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0179] In addition, unless otherwise specified, the functional units / modules in the various embodiments of the present invention may be integrated into a single unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The aforementioned integrated units / modules may be implemented in the form of hardware or software program modules.

[0180] The integrated units / modules, if implemented in the form of hardware, can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor or chip can be any appropriate hardware processor, such as a CPU, a GPU, an FPGA, a DSP, an ASIC, etc. Unless otherwise specified, the on-chip cache, off-chip memory, storage can be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.

[0181] If the integrated units / modules are implemented in the form of software program modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application, essentially or the part that contributes to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for making a computer electronic device (which can be a personal computer, a server, or a network electronic device, etc.) execute all or part of the steps of the method described in various embodiments of the present disclosure. The aforementioned storage medium includes: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various program code storage media.

[0182] The embodiments of the present application also provide a computer readable storage medium, which stores one or more computer programs, and when the one or more computer programs are executed by a plurality of processors, the processors execute the method and detailed solutions shown in the above embodiments. Figure 3 and Figure 4 The embodiments of the present application also provide a computer readable storage medium, which stores one or more computer programs, and when the one or more computer programs are executed by a plurality of processors, the processors execute the method and detailed solutions shown in the above embodiments.

[0183] The embodiments of the present application also provide a computer program product, which contains a computer program, and when the computer program runs on a computer, the computer executes the method of any of the above embodiments.

[0184] References to features, advantages, or similar language throughout this specification do not imply that all features and advantages achievable with this solution are included or embodied in any single implementation thereof. Rather, language referring to features and advantages is understood to mean that a specific feature, advantage, or characteristic described in connection with an embodiment is included in at least one embodiment of this solution. Therefore, discussions of features, advantages, and similar language throughout this specification may, but do not necessarily, refer to the same embodiment.

[0185] Furthermore, the features, advantages, and characteristics of the present invention may be combined in any suitable manner in one or more embodiments. Based on the description herein, one of ordinary skill in the relevant art will recognize that the present invention may be practiced without one or more of the specific features or advantages of a particular embodiment. In other cases, additional features and advantages may be realized in a particular embodiment that is not presented in all embodiments of the present invention.

[0186] The embodiments of the present application are described in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. At the same time, changes or modifications made by those skilled in the art based on the ideas of the present application, the specific implementation methods, and the scope of application of the present application, all fall within the scope of protection of the present application. In summary, the contents of this specification should not be construed as limiting the present application.

Claims

1. A multi-round human-computer interaction method for visual images, characterized in that: include: (a) Extracting local image features related to the current multi-round dialogue text features from the global image information; (b) updating the local image features according to the current historical local image features to obtain updated local image features; (c) determining visual image features based on the updated local image features and global image features corresponding to the global image information using a cross-attention mechanism; as well as (d) inputting the visual image features into a multimodal large model for processing; Wherein, step (b) comprises: (b1) Determine the updated features of the current round of dialogue and the updated features of the historical dialogue based on the local image features and the current historical local image features; (b2) determining a corresponding query vector matrix for this round of dialogue based on the updated features of this round of dialogue; (b3) determining a corresponding historical conversation key matrix and a historical conversation value matrix according to the historical conversation update feature; (b4) using an optimal matching mechanism based on the current round conversation query vector matrix and the historical conversation key matrix, obtaining a filtered historical conversation key matrix and a filtered historical conversation value matrix from the historical conversation key matrix and the historical conversation value matrix, the correlations of which with the local image features satisfy a preset condition; and (b5) obtaining the updated local image features according to the current round conversation query vector matrix, the filtered historical conversation key matrix, and the filtered historical conversation value matrix; Wherein, step (b4) comprises: Determining, based on the query vector matrix of the current conversation and the historical conversation key matrix, a relevance score corresponding to each vector in the historical conversation key matrix; Determining a preset number of historical conversation key vectors having the highest correlation scores from the historical conversation key matrix based on the correlation scores, and generating the filtered historical conversation key matrix; and The preset number of historical conversation value vectors corresponding to the preset number of historical conversation key vectors are determined to generate the filtered historical conversation value matrix.

2. The method according to claim 1, wherein After step (b5), step (b) further comprises: (b6) in response to receiving the input image information, determining image features of the current round of conversation; and (b7) Determine the updated local image feature as the current historical local image feature, and return to step (b1).

3. The method according to claim 1 or 2, wherein: Step (a) comprises: performing correlation calculations on the multiple image features corresponding to the global image information and the text features of the current multiple rounds of conversations, respectively, to determine multiple correlation values ​​accordingly; and A preset number of image features with the largest correlation values ​​among the multiple image features are determined as local image features related to the current multi-round dialogue text features.

4. The method according to claim 1 or 2, wherein: Step (c) comprises: A cross-attention mechanism is adopted to determine visual image features based on the updated local image features, the global image features and the learnable vector.

5. The method according to claim 1 or 2, wherein: After step (d), the method further comprises: (e) determining text features of the current multi-turn dialogue in response to receiving the input text; and (f) Determine the visual image feature as the current historical local image feature and return to step (a).

6. A human-computer multi-round interaction device for visual images, characterized in that: include: An extraction module is used to extract local image features related to the current multi-round dialogue text features from the global image information; An acquisition module, configured to update the local image features according to the current historical local image features and acquire the updated local image features; a determination module, configured to determine visual image features based on the updated local image features and global image features corresponding to the global image information using a cross-attention mechanism; as well as An input module, configured to input the visual image features into a multimodal large model for processing; Wherein, the acquisition module includes: A first determination submodule is configured to determine, based on the local image features and the current historical local image features, the updated features of the current round of conversation and the updated features of the historical conversation respectively; A second determination submodule is configured to determine a corresponding query vector matrix for this round of dialogue based on the updated features of this round of dialogue; A third determination submodule is configured to determine a corresponding historical conversation key matrix and a historical conversation value matrix according to the historical conversation update feature; a first acquisition submodule, configured to employ an optimal matching mechanism based on the current-round conversation query vector matrix and the historical conversation key matrix to acquire, from the historical conversation key matrix and the historical conversation value matrix, a filtered historical conversation key matrix and a filtered historical conversation value matrix whose correlation with the local image features satisfies a preset condition; and a second acquisition submodule, configured to acquire the updated local image features based on the current round conversation query vector matrix, the filtered historical conversation key matrix, and the filtered historical conversation value matrix; Wherein, the first acquisition submodule is used for: Determining, based on the query vector matrix of the current conversation and the historical conversation key matrix, a relevance score corresponding to each vector in the historical conversation key matrix; Determining a preset number of historical conversation key vectors having the highest correlation scores from the historical conversation key matrix based on the correlation scores, and generating the filtered historical conversation key matrix; and The preset number of historical conversation value vectors corresponding to the preset number of historical conversation key vectors are determined to generate the filtered historical conversation value matrix.

7. An electronic device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the method according to any one of claims 1 to 5 when executing the computer program in the memory.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Visual dialogue generation method based on human-like visual perception and language memory network

    CN116303955A

  • Multi-round dialogue processing method and system based on RAG and knowledge graph

    CN118885627A