Iterative optimization and multi-granularity perception-based personal dialogue positioning method and device
By introducing multi-scale feature extraction, early cross-modal fusion, gated network and iterative optimization query optimizer in the embodied dialogue positioning method, the problems of insufficient multi-grained feature perception, insufficient cross-modal fusion and low positioning accuracy in the existing technology are solved, and a more efficient and robust embodied dialogue positioning effect is achieved.
Patent Information
- Application Number
- CN202510321220.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-20
AI Technical Summary
The existing embodied dialogue positioning methods have shortcomings in multi-grained feature perception, cross-modal fusion, iterative optimization mechanism, and generalization capabilities in complex scenarios, resulting in insufficient positioning accuracy and poor generalization capabilities.
Using an embodied dialogue positioning method based on iterative optimization and multi-grained size perception, multi-grained visual features are extracted through a multi-scale feature extraction module, the cross-modal feature fusion module performs early cross-modal fusion at the intermediate layer of the encoder, the gated network performs feature filtering, and the mask query optimizer can learn to query and gradually optimize the predicted coordinates through iterative updates.
It significantly improves the accuracy and robustness of embodied dialogue positioning, can effectively handle complex scenarios, and is suitable for practical application scenarios such as emergency search and rescue, navigation, etc.
Smart Images

Figure CN120180366A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision technology and embodied intelligence technology, and particularly relates to an embodied dialogue localization method and device based on iterative optimization and multi-granularity perception. Background Art
[0002] Embodied dialogue localization refers to the task of determining the target location on a given 2D map through multiple rounds of natural language dialogue. Such tasks have broad application prospects in fields such as emergency search and rescue, intelligent navigation, and robot operation. For example, in an emergency search and rescue scenario, rescue personnel need to quickly locate the position of the trapped person through dialogue with them; in intelligent navigation, users can guide a robot or an autonomous vehicle to a specified location through natural language instructions.
[0003] Currently, existing embodied dialogue localization methods mainly rely on cross-modal fusion of vision and language. Although certain progress has been made in rough localization, there are still significant challenges in precise localization, mainly including:
[0004] (1) Insufficient multi-granularity feature perception: Existing methods usually only focus on feature extraction at a single scale, making it difficult to simultaneously process large-scale regions (such as "bedroom") and small-scale objects (such as "sofa") in the map. This single-scale feature extraction method results in insufficient localization accuracy of the model in complex scenarios, especially when the target object is similar to the surrounding environment, and the localization error is large.
[0005] (2) Inadequate cross-modal fusion: Most existing methods adopt a late fusion strategy, that is, fusing after visual and language features are extracted separately. This strategy is difficult to effectively capture the complex semantic relationships between vision and language, especially when the dialogue involves multi-level semantic information, and the performance of the model is often unsatisfactory.
[0006] (3) Poor accurate positioning ability: Existing methods usually regard embodied dialogue localization as a prediction problem from the original image to the target heat map, lacking a mechanism for gradually optimizing the target position. This one-time prediction method is prone to cumulative localization errors. Especially in multi-round dialogues, the target position may change with the dialogue content, and the model is difficult to dynamically adjust the prediction result.
[0007] (4) Insufficient generalization ability in complex scenarios: Existing methods perform well in known environments, but have poor generalization ability when facing new environments. This is mainly because the model over-relies on features of specific scenarios during training and lacks the ability to adapt to multi-scenarios and multi-tasks.
[0008] In summary, the existing embodied dialogue localization methods have obvious deficiencies in aspects such as multi-granularity feature perception, cross-modal fusion, iterative optimization mechanism, and generalization ability in complex scenarios. Therefore, designing an embodied dialogue localization method that can effectively extract multi-granularity features, achieve early cross-modal fusion, and gradually improve the localization accuracy through iterative optimization has important practical significance and application value. Summary of the Invention
[0009] The purpose of the present invention is to provide an embodied dialogue localization method and device based on iterative optimization and multi-granularity perception to solve the technical problems existing in the prior art in aspects such as multi-granularity feature perception, cross-modal fusion, iterative optimization mechanism, and generalization ability in complex scenarios. The technical solution disclosed by the present invention effectively extracts multi-granularity features, realizes early cross-modal fusion, and gradually improves the localization accuracy through iterative optimization, which can significantly improve the accuracy and robustness of embodied dialogue localization.
[0010] To achieve the above object, the present invention adopts the following technical solutions:
[0011] In the first aspect of the present invention, an embodied dialogue localization method based on iterative optimization and multi-granularity perception is provided, including the following steps:
[0012] Obtain text containing multiple rounds of dialogue and the corresponding 2D map image;
[0013] Based on the obtained text and 2D map image, use the trained embodied dialogue localization model to predict the target position and obtain the predicted coordinates of the target position;
[0014] Among them, the embodied dialogue localization model includes:
[0015] A multi-scale feature extraction module for extracting the text features of the text and the multi-granularity visual features of the 2D map image;
[0016] A cross-modal feature fusion module for early fusion of the text features and the multi-granularity visual features in the middle layer of the encoder and outputting the fused multi-modal features;
[0017] A gating network for filtering the multi-modal features and outputting the filtered features;
[0018] A mask query optimizer for inputting the filtered features, gradually optimizing the predicted coordinates by iteratively updating the learnable query, and using the mask mechanism to narrow the attention range to obtain the query result; generating a confidence level and a heat map according to the query result to obtain the final predicted coordinates of the target position.
[0019] A further improvement of the present invention lies in that, in the multi-scale feature extraction module, the steps of extracting the text features of the text include:
[0020] Encoding the text containing multi-turn conversations into text features using a pre-trained BERT model.
[0021] A further improvement of the present invention lies in that, in the multi-scale feature extraction module, the steps of extracting the multi-granularity visual features of the 2D map image include:
[0022] Using a backbone network to extract initial multi-granularity visual features from the 2D map image;
[0023] Using multiple parallel convolutional branches to enhance the initial multi-granularity visual features to obtain enhanced multi-granularity visual features; wherein each convolutional branch has a different convolutional kernel size.
[0024] A further improvement of the present invention lies in that, in the cross-modal feature fusion module, the steps of performing early fusion of the text features and the multi-granularity visual features in the middle layer of the encoder include:
[0025] Performing cross-attention calculation on the text features and the multi-granularity visual features in each layer of the encoder to obtain aligned multi-modal features;
[0026] Among them, the aligned multi-modal features are expressed as:
[0027]
[0028] In the formula, F i m represents the multi-modal features, i represents the number of the layer in the backbone network, m represents the multi-modal; Q i is the attention query, K i is the attention key, K i =IW i K ; V i is the attention value, V i =IW i V ; C i is the scaling factor, which is the dimension of Q i and K i ; is the multi-granularity visual features; I is the text features; W i Q 、W i K and W i V are projection matrices; ⊙ represents the Hadamard product.
[0029] A further improvement of the present invention lies in that the filtered and processed feature representation output by the gating network is:
[0030] F i e = F i + G(F i m ) · F i m ;
[0031] In the formula, F i e is the filtered and processed feature; F i is the feature of the i-th layer in the multi-granularity visual feature; F i m is the multi-modal feature; G(F i m ) is the feature weighting coefficient;
[0032] G(F i m ) = Tanh(Linear(ReLU(Linear(F i m )))));
[0033] In the formula, Linear(·) represents linear projection; Tanh(·) and ReLU(·) both represent activation functions.
[0034] A further improvement of the present invention lies in that in the mask query optimizer, after performing input filtering and processing on the feature, the learnable query is iteratively updated to gradually optimize the predicted coordinates, and the attention range is narrowed using the mask mechanism to obtain the query result; the steps of generating the confidence and heatmap based on the query result to obtain the final predicted coordinates of the target position include:
[0035] First, downsample the features of different scales in the filtered and processed feature, then perform self-attention calculation, and finally upsample the output back to the original size to obtain the enhanced feature
[0036] First, initialize a group of learnable queries, and then iteratively update the learnable queries through the mask query optimizer to finally obtain the query result; among them, in a total of N loop iterations, each in is decoded by the decoder
[0037]
[0038] Wherein, X is the set of queries output by each layer; represents the binary mask obtained from the last layer; Q is the query of the previous layer after linear projection; K and V are linear projections of
[0039] The final heat map M is generated by a multi-layer perceptron and is expressed as:
[0040]
[0041] A further improvement of the present invention lies in that the overall loss function adopted during the training of the embodied dialogue localization model is expressed as:
[0042]
[0043] Wherein, is the overall loss function; is the divergence loss; is the confidence loss; λ1 and λ2 are the weights of the divergence loss and the confidence loss respectively.
[0044] In the second aspect of the present invention, an embodied dialogue localization system based on iterative optimization and multi-granularity perception is provided, including:
[0045] A data acquisition module for acquiring text containing multi-round dialogues and corresponding 2D map images;
[0046] A position prediction module for predicting the target position based on the acquired text and 2D map images by using the trained embodied dialogue localization model to obtain the predicted coordinates of the target position;
[0047] Wherein, the embodied dialogue localization model includes:
[0048] A multi-scale feature extraction module for extracting the text features of the text and the multi-granularity visual features of the 2D map image;
[0049] A cross-modal feature fusion module for early fusing the text features and the multi-granularity visual features in the middle layer of the encoder and outputting the fused multi-modal features;
[0050] A gating network for filtering the multi-modal features and outputting the filtered features;
[0051] A mask query optimizer for inputting the filtered features, gradually optimizing the predicted coordinates by iteratively updating the learnable query, and using the mask mechanism to narrow the attention range to obtain the query result; generating the confidence and heat map according to the query result to obtain the final predicted coordinates of the target position.
[0052] In a third aspect of the present invention, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the embodied dialogue localization method based on iterative optimization and multi-granularity perception as described in any one of the first aspects of the present invention.
[0053] In a fourth aspect of the present invention, there is provided a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the embodied dialogue localization method based on iterative optimization and multi-granularity perception as described in any one of the first aspects of the present invention.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] The present invention discloses an embodied dialogue localization method based on iterative optimization and multi-granularity perception. Based on a trained embodied dialogue localization model, it can effectively extract multi-granularity features, achieve early cross-modal fusion, and gradually improve the localization accuracy through iterative optimization, significantly improving the accuracy and robustness of embodied dialogue localization. Among them, the embodied dialogue localization model includes: a multi-scale feature extraction module for extracting multi-granularity visual features from a 2D map; a cross-modal feature fusion module for early fusing visual features and text features in the middle layer of the encoder; a gating network for filtering redundant features and enhancing useful features; an iterative optimization decoder for gradually optimizing the predicted coordinates through a masked query optimizer, narrowing the attention range and improving the localization accuracy. Summarily, the technical solution of the present invention significantly improves the accuracy of embodied dialogue localization through a multi-granularity perception and iterative optimization mechanism, especially performs well in complex scenarios, and is applicable to practical application scenarios such as emergency search and rescue, navigation, etc.
[0056] Aiming at the problem of insufficient multi-granularity feature perception in existing methods, the present invention adopts a multi-scale feature extraction module (MFE) to extract visual features of different scales through multiple parallel convolutional branches, which can simultaneously process large-scale regions (such as "bedroom") and small-scale objects (such as "sofa") in the map. In this way, it has the ability of multi-granularity feature perception and can obtain a more accurate localization effect, especially significantly improving the localization accuracy in complex scenarios.
[0057] Aiming at the problem of insufficient cross-modal fusion in existing methods, the present invention adopts a cross-modal feature fusion module (CFP) to perform early fusion of visual and text features in the middle layer of the encoder, and realizes the alignment of vision and language through a cross-attention mechanism. In this way, it has the ability of early cross-modal fusion and can obtain a better semantic alignment effect, improving the model's understanding ability of multi-level semantic information.
[0058] To address the problem of poor accurate positioning ability of existing methods, the present invention adopts an iterative optimization decoder. By means of a Mask Query Optimizer (MQR), it gradually optimizes the predicted coordinates, gradually narrows the attention range, and corrects the prediction error. In this way, it has the ability of iterative optimization, can obtain a gradually accurate positioning effect, and significantly reduces the positioning error, especially for dynamically adjusting the prediction results in multi-turn conversations.
[0059] To address the problem of insufficient generalization ability of existing methods in complex scenarios, the present invention adopts a Gating Network (GN) and a multi-turn conversation processing mechanism. The gating network filters redundant features and enhances useful features, and at the same time supports both single-turn and multi-turn conversation settings. In this way, it has stronger generalization ability, can obtain excellent performance in both known and unknown environments, and is applicable to a variety of practical application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art; obviously, the drawings in the following description are some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0061] Figure 1 is a schematic flowchart of an embodied conversation positioning method based on iterative optimization and multi-granularity perception according to an embodiment of the present invention;
[0062] Figure 2 is a schematic diagram of the processing logic of an embodied conversation positioning model according to an embodiment of the present invention;
[0063] Figure 3 is a schematic diagram for the principle explanation of an embodied conversation positioning model according to an embodiment of the present invention;
[0064] Figure 4 is a schematic diagram of the network structure of an embodied conversation positioning model according to an embodiment of the present invention;
[0065] Figure 5 is a schematic diagram of an iterative optimization decoder according to an embodiment of the present invention;
[0066] Figure 6 is a schematic diagram of an embodied conversation positioning system based on iterative optimization and multi-granularity perception according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the embodiments of the present invention; obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0068] Based on the technical solutions disclosed in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0069] Please refer to Figures 1 to 4 , an embodied dialogue localization method based on iterative optimization and multi-granularity perception provided by the embodiments of the present invention includes the following steps:
[0070] Step 1, obtain the text containing multiple rounds of conversations and the corresponding 2D map images;
[0071] Step 2, based on the text and 2D map images obtained in Step 1, use the trained embodied dialogue localization model to predict the target position and obtain the predicted coordinates of the target position;
[0072] Among them, the embodied dialogue localization model includes:
[0073] A multi-scale feature extraction module for extracting the text features of the text and the multi-granularity visual features of the 2D map images;
[0074] A cross-modal feature fusion module for early fusing the text features and the multi-granularity visual features in the middle layer of the encoder and outputting the fused multi-modal features;
[0075] A gating network for filtering the multi-modal features and outputting the filtered features;
[0076] A mask query optimizer for inputting the filtered features, gradually optimizing the predicted coordinates by iteratively updating the learnable query, using the mask mechanism to narrow the attention range and obtain the query result; generating the confidence and heat map according to the query result to obtain the final predicted coordinates of the target position.
[0077] In the technical solution provided by the embodiments of the present invention, taking the text containing multiple rounds of conversations and the corresponding 2D map image as inputs, and using the trained embodied conversation localization model to predict the target position, the precise coordinates of the target position can be finally obtained. In the embodied conversation localization model of the embodiments of the present invention, multi-scale feature extraction module (MFE) is used to extract multi-granularity features from the 2D map; through cross-modal feature fusion, cross-modal feature fusion is performed in the middle layer of the encoder to perform early fusion of text features and visual features; a gating network is used to filter the fused features to reduce unnecessary attention; a masked query optimizer (MQR) is used to iteratively update the learnable query to gradually optimize the predicted coordinates; a multi-layer perceptron (MLP) is used to generate confidence and heatmaps to output the final localization result, which solves the deficiencies of the prior art in aspects such as multi-granularity feature perception, cross-modal fusion, iterative optimization mechanism, and generalization ability in complex scenarios. Summarily, the embodiments of the present invention propose a solution combining multi-granularity feature extraction, early cross-modal fusion, and iterative optimization mechanism. This method can effectively capture the multi-granularity visual features in the 2D map, and accurately align them with the text information. At the same time, the prediction error is gradually corrected through iterative optimization, significantly improving the accuracy and robustness of conversation localization; experimental results show that the present invention is superior to the existing state-of-the-art methods in both single-conversation and multi-conversation settings, especially showing strong generalization ability in unseen environments, providing an efficient and accurate solution for tasks such as navigation and emergency rescue in practical application scenarios.
[0078] In a specific embodiment of the present invention, the further explanation of the multi-scale feature extraction module (MFE) is as follows:
[0079] Given multiple rounds of conversations and a 2D map, first use the pre-trained BERT model to encode the conversations into text features S represents the maximum token length, D l represents the output dimension; at the same time, use a visual backbone network to extract multi-granularity visual features from the 2D map H and W respectively represent the height and width of the input image, D v represents the output dimension, and i represents the number of a certain layer in the backbone network;
[0080] To further capture the features of different granularities in the 2D map, J parallel convolutional branches are used, each branch having a different convolutional kernel size to capture the local features of different receptive fields;
[0081] Among them, the enhanced visual features are calculated as follows:
[0082]
[0083] Among them, Denote the j-th convolutional kernel, and Conv(·) represents the convolution operation.
[0084] In a specific embodiment of the present invention, the further explanation of the cross-modal feature fusion module is as follows:
[0085] Different from existing previous methods, the embodiment of the present invention performs early cross-modal feature fusion in the middle layer of the encoder; specifically, after obtaining the enhanced visual features and text features I, cross-attention calculation is performed on them at each layer of the encoder to obtain the aligned multi-modal features F i m , and the calculation is as follows:
[0086]
[0087] K i =IW i K
[0088] V i =IW i V
[0089]
[0090] where W i Q , W i K and W i V are projection matrices, and ⊙ represents the Hadamard product; Q i is the attention query used in the i-th layer of the backbone network, K i is the attention key used in the i-th layer of the backbone network, and V i is the attention value used in the i-th layer of the backbone network.
[0091] In an embodiment of the present invention, the further explanation of the gating network is as follows:
[0092] After multi-granularity feature extraction and cross-modal feature fusion, a large number of multi-granularity features aligned with the language are obtained; in order to distinguish which features are useful, the embodiment of the present invention uses a gating network to adjust the multi-modal features F i m ;
[0093] where the specific implementation of the gating network is as follows:
[0094] G(F i m ) = Tanh(Linear(ReLU(Linear(Fi m ))));
[0095] In the formula, Linear(·) represents a linear projection; both Tanh(·) and ReLU(·) represent activation functions;
[0096] Subsequently, the output of the gating network is superimposed on F i m to obtain the final encoder output representation as:
[0097] F i e = F i + G(F i m ) · F i m ;
[0098] In the formula, F i e is the filtered feature; F i is the feature extracted by the original backbone network.
[0099] Please refer to Figure 5 for a further explanation of the mask query optimizer (mask query optimizer) in a specific embodiment of the present invention as follows:
[0100] To enhance the interaction between features of different scales, the present invention implements a cross-scale attention mechanism; wherein, first, the features of different scales are downsampled, then self-attention calculation is performed, and finally the output is upsampled back to the original size to obtain enhanced features
[0101] To further improve the localization accuracy of the model, the present invention proposes a new dialogue localization paradigm, introduces a learnable query set, and iteratively updates these queries through a mask query optimizer (MQR);
[0102] In a specific exemplary technical solution, in the mask query optimizer module, N mask query optimizer loops are used, and each loop uses 3 mask query optimizer layers to respectively decode each in for query refinement to obtain query results, and the queries are all processed by a multi-layer perceptron; specifically, after each layer, the queries are processed by a multi-layer perceptron, then a dot product is performed with Figure 2 and then a sigmoid operation is performed to obtain a heat map, and then the heat
[0103] Among them, each masked query optimizer layer consists of a masked cross-attention layer, a standard self-attention layer, and a feed-forward layer with residual connections. The specific calculation of the masked cross-attention is as follows:
[0104]
[0105] Among them, X is the set of queries output by each layer, represents the binary mask obtained from the last layer. Q is the query of the previous layer after linear projection; K and V are linear projections of;
[0106] The final heatmap M is generated by a multi-layer perceptron (MLP) and can be expressed as:
[0107]
[0108] In a specific embodiment of the present invention, further explanations regarding the training of the embodied dialogue localization model and the loss function are as follows:
[0109] Given the predicted heatmap M and the target distribution the model is trained by minimizing the KL divergence; in addition, Focal Loss is used as the confidence loss; therefore, the total loss is calculated as follows:
[0110]
[0111] Among them, λ1 and λ2 are the weights of the KL divergence and Focal Loss respectively.
[0112] Please refer to Figure 3 and Figure 4 , which shows the main structural framework and its working principle of the embodied dialogue localization method based on iterative optimization and multi-granularity perception proposed by the present invention. As can be seen from Figure 4 the left part, the network consists of multiple key modules, including a multi-scale feature extraction module (MFE), a cross-modal feature fusion module, a gating network, and a masked query optimizer (MQR); among them, the multi-scale feature extraction module captures multi-granularity visual features in the 2D map through convolutional kernels of different scales, while the cross-modal feature fusion module performs early fusion of text features and visual features in the middle layer of the encoder to ensure the precise alignment of multi-modal information; the gating network further filters the fused features to reduce unnecessary attention interference; finally, the masked query optimizer gradually optimizes the predicted coordinates by iteratively updating the learnable queries, and uses the masking mechanism to narrow the attention range, thereby improving the localization accuracy. Figure 4The right part shows the attention distribution of the network at different iteration stages and its gradual optimization process for the localization results. It can be seen that as the number of decoder iterations increases, the model gradually focuses its attention within a smaller area and can correct the errors in the initial prediction, ultimately achieving high-precision object localization.
[0113] For example, in the first round of iteration, the model may disperse its attention over multiple similar semantic regions (such as multiple bedrooms), but as the iteration progresses, the model gradually narrows the attention range and finally accurately locates the target region (such as "the sofa beside the bed"). This process fully demonstrates the advantages of the network in multi-granularity perception and iterative optimization and can effectively handle the dialogue localization task in complex scenarios.
[0114] In the specific verification embodiments of the present invention, experiments were conducted on the WAY dataset. The results show that in both single-shot and multi-shot dialogue settings, the RAMP method of the present invention significantly outperforms the existing state-of-the-art methods. Especially in unseen environments, RAMP improved by 29.34% and 14.13% respectively in the Acc0@valUnseen and Acc5@valUnseen metrics.
[0115] Table 1. Experimental results on the WAY dataset
[0116]
[0117] Specifically and explanatorily, Table 1 shows the experimental results of the method of the present invention on the WAY dataset, comparing the localization accuracies in single-shot and multi-shot dialogue settings respectively. In Table 1, valSeen and valUnseen represent the test results in seen and unseen environments respectively, and Acc0 and Acc5 represent the accuracies where the distance error between the predicted position and the true position is within 0 meters and 5 meters respectively. It can be seen from Table 1 that in the single-shot dialogue setting, compared with the existing state-of-the-art methods (such as LingUNet and DiaLoc), the RAMP method of the present invention improved by 29.34% and 14.13% respectively in the Acc0 and Acc5 metrics in the valUnseen environment. In the multi-shot dialogue setting, the RAMP method also performed excellently, especially in the Acc0 metric in the valUnseen environment, which improved by 36.85% compared with the existing methods. These results indicate that the RAMP method not only performs excellently in seen environments but also demonstrates strong generalization ability in unseen environments and can effectively handle the dialogue localization task in complex scenarios.
[0118] In summary, through iterative optimization and multi-granularity perception, the present invention significantly improves the accuracy of dialogue localization and can be effectively applied to tasks such as navigation and emergency rescue in practical scenarios. Specifically, in the embodiments of the present invention, aiming at the problems of insufficient multi-granularity feature perception, insufficient cross-modal alignment, and low localization accuracy existing in the existing methods in the dialogue localization task, a dialogue localization method (RAMP) based on iterative optimization and multi-granularity perception is proposed. By introducing a multi-scale feature extraction module (MFE) and a cross-modal early fusion mechanism, multi-granularity visual features in the 2D map are effectively captured and precisely aligned with text features. In addition, through the iterative optimization mechanism of the mask query optimizer (MQR), the attention range is gradually narrowed and the prediction error is corrected, significantly improving the localization accuracy. Experimental results show that this method is superior to the existing state-of-the-art methods in both single-dialogue and multi-dialogue settings, especially showing stronger generalization ability in unseen environments. The present invention provides an efficient and accurate solution for the dialogue localization task and has broad practical application prospects in fields such as emergency rescue and intelligent navigation.
[0119] The following is the device embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For details not disclosed in the device embodiment, please refer to the method embodiment of the present invention.
[0120] Please refer to Figure 6 , in the embodiments of the present invention, a grounded dialogue localization system based on iterative optimization and multi-granularity perception is provided, including:
[0121] A data acquisition module, configured to acquire text containing multi-turn dialogues and corresponding 2D map images;
[0122] A position prediction module, configured to perform target position prediction based on the acquired text and 2D map images by using a trained grounded dialogue localization model to obtain target position prediction coordinates;
[0123] Wherein, the grounded dialogue localization model includes:
[0124] A multi-scale feature extraction module, configured to extract text features of the text and multi-granularity visual features of the 2D map images;
[0125] A cross-modal feature fusion module, configured to perform early fusion of the text features and the multi-granularity visual features in the middle layer of the encoder and output the fused multi-modal features;
[0126] A gating network, configured to perform filtering processing on the multi-modal features and output the filtered features;
[0127] A mask query optimizer is used to input the features after filtering processing. By iteratively updating the learnable query, it gradually optimizes the predicted coordinates, and uses the mask mechanism to narrow the attention range to obtain the query result. According to the query result, confidence and heatmap are generated to obtain the final predicted coordinates of the target position.
[0128] In one embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program. The computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used to execute the operations of the embodied dialogue localization method based on iterative optimization and multi-granularity perception.
[0129] In one embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, one or more instructions suitable for being loaded and executed by the processor are stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed Random Access Memory (RAM) memory, or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the embodied dialogue localization method based on iterative optimization and multi-granularity perception in the above embodiments.
[0130] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) that contain computer-usable program code.
[0131] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0132] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0133] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still, the specific implementation manners of the present invention can be modified or equivalently replaced, and any modification or equivalent replacement without departing from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A method for embodied dialogue localization based on iterative optimization and multi-granularity perception, characterized in that: The steps include: Get the text containing multiple rounds of dialogue and the corresponding 2D map image; Based on the acquired text and 2D map image, the trained embodied dialogue localization model is used to predict the target location and obtain the predicted target location coordinates; The embodied dialogue localization model includes: A multi-scale feature extraction module, used to extract text features of the text and multi-granularity visual features of the 2D map image; A cross-modal feature fusion module, used for early fusing the text feature with the multi-granularity visual feature in the middle layer of the encoder, and outputting the fused multi-modal feature; A gating network, used for filtering the multimodal features and outputting the filtered features; The mask query optimizer is used to input the filtered features, learn the query through iterative updates, gradually optimize the predicted coordinates, and use the mask mechanism to narrow the attention range to obtain the query results; generate confidence and heat map according to the query results to obtain the final target position predicted coordinates.
2. The embodied dialogue localization method based on iterative optimization and multi-granularity perception according to claim 1, characterized in that: In the multi-scale feature extraction module, the step of extracting text features of the text includes: Use the pre-trained BERT model to encode text containing multi-turn dialogues into text features.
3. The embodied dialogue localization method based on iterative optimization and multi-granularity perception according to claim 1, characterized in that: In the multi-scale feature extraction module, the step of extracting multi-granularity visual features of the 2D map image includes: Use the backbone network to extract initial multi-granular visual features from 2D map images; The initial multi-granularity visual features are enhanced using multiple parallel convolution branches to obtain enhanced multi-granularity visual features; wherein each convolution branch has a different convolution kernel size.
4. The embodied dialogue localization method based on iterative optimization and multi-granularity perception according to claim 1, characterized in that: In the cross-modal feature fusion module, the step of performing early fusion of the text feature with the multi-granularity visual feature at the middle layer of the encoder comprises: At each layer of the encoder, cross-attention calculations are performed on text features and multi-granularity visual features to obtain aligned multimodal features; Among them, the aligned multimodal features are expressed as: In the formula, F i m represents multimodal features, i represents the number of the layer in the backbone network, and m represents multimodality; Q i is an attention query, K i is the attention key, K i =IW i K ; V i is the attention value, V i =IW i V ; C i is the scaling factor, Q i and K i Dimensions; is a multi-granularity visual feature; I is a text feature; W i Q , W i K and W i V is the projection matrix; ⊙ represents the Hadamard product.
5. The embodied dialogue localization method based on iterative optimization and multi-granularity perception according to claim 1, characterized in that: The filtered feature representation of the gating network output is: In the formula, F i e is the feature after filtering; F i is the feature of the i-th layer in the multi-granularity visual feature; F i m It is a multimodal feature; is the feature weighting coefficient; Where Linear(·) represents linear projection; Tanh(·) and ReLU(·) both represent activation functions.
6. The embodied dialogue localization method based on iterative optimization and multi-granularity perception according to claim 1, characterized in that: In the mask query optimizer, the features after input filtering are executed, the query can be learned through iterative updates, the predicted coordinates are gradually optimized, and the mask mechanism is used to narrow the attention range to obtain the query result; The steps of generating a confidence level and a heat map according to the query result and obtaining the final target position prediction coordinates include: First, the features of different scales in the filtered features are downsampled, then self-attention calculation is performed, and finally the output is upsampled back to the original size to obtain enhanced features. First, a set of learnable queries is initialized, and then the learnable queries are updated through the mask query optimizer loop iteration to finally obtain the query results; in which, in a total of N loop iterations, the decoder is used to decode Each of To refine the query, the output query is processed into a heat map as the attention mask for subsequent cycles; in each iteration, the query is updated by the masked criss-cross attention layer, the standard self-attention layer, and the feed-forward layer with residual connection. The masked criss-cross attention is calculated as follows: Where X is the set of queries output by each layer; represents the binary mask obtained from the last layer; Q is the query of the previous layer after linear projection; K and V are Linear projection of The final heat map M is generated by a multi-layer perceptron and is expressed as:
7. The embodied dialogue localization method based on iterative optimization and multi-granularity perception according to claim 1, characterized in that: The overall loss function used in training the embodied dialogue localization model is expressed as: In the formula, is the overall loss function; is the divergence loss; is the confidence loss; λ1 and λ2 are the weights of the divergence loss and confidence loss respectively.
8. An embodied dialogue localization system based on iterative optimization and multi-granularity perception, characterized in that: include: A data acquisition module, used to acquire text containing multiple rounds of dialogue and corresponding 2D map images; The location prediction module is used to predict the target location based on the acquired text and 2D map image using the trained embodied dialogue localization model to obtain the predicted target location coordinates; The embodied dialogue localization model includes: A multi-scale feature extraction module, used to extract text features of the text and multi-granularity visual features of the 2D map image; A cross-modal feature fusion module, used for early fusing the text feature with the multi-granularity visual feature in the middle layer of the encoder, and outputting the fused multi-modal feature; A gating network, used for filtering the multimodal features and outputting the filtered features; The mask query optimizer is used to input the filtered features, learn the query through iterative updates, gradually optimize the predicted coordinates, and use the mask mechanism to narrow the attention range to obtain the query results; generate confidence and heat map according to the query results to obtain the final target position predicted coordinates.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the embodied dialogue localization method based on iterative optimization and multi-granularity perception is implemented as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the embodied dialogue localization method based on iterative optimization and multi-granularity perception is implemented as described in any one of claims 1 to 7.