Personal dialogue positioning method based on causal-guided diffusion model and related device

Through the causal guided diffusion model, text and visual features in embodied dialogue positioning are extracted and processed, confounding factors are eliminated and denoised through the diffusion network, which solves the shortcomings of the existing embodied dialogue positioning methods in terms of accuracy, generalization ability and anti-interference, and achieves high-precision and robust positioning effects.

CN120146197APending Publication Date: 2025-06-13XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510326971.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing embodied dialogue positioning methods have shortcomings in accuracy, generalization ability and anti-interference, especially in high resolution dependence, insufficient generalization ability and confounding factor interference.

Method used

The embodied dialogue positioning method based on the causal guided diffusion model is adopted, and text and visual features are extracted through the feature extraction module, the causal reasoning module eliminates confounding factors, and the diffusion network performs step-by-step denoising processing to obtain coordinate positioning results.

Benefits of technology

It effectively reduces the dependence on high resolution, improves positioning accuracy and generalization capabilities, enhances the robustness of the model in complex environments, and ensures the accuracy of positioning results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146197A_ABST
    Figure CN120146197A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of a computer vision technology and a personal intelligence technology, and particularly discloses a personal dialogue positioning method based on a causal guide diffusion model and a related device. The personal dialogue positioning method comprises the following steps: acquiring a dialogue text and a top view map of a scene to be positioned, and performing positioning prediction by utilizing a trained causal guide diffusion model to obtain a coordinate positioning result; the causal-guided diffusion model comprises: a feature extraction module; the causal reasoning module is used for taking the text features and the visual features extracted by the feature extraction module as original features, performing confounding factor elimination processing and obtaining confounding-removed features; the de-mixing feature guiding module is used for dynamically adjusting the weights of the original features and the de-mixing features and obtaining screened features; and a diffusion network. According to the technical scheme of the invention, the technical problems of strong resolution dependence, insufficient generalization ability, interference of mixed factors and the like in an owned dialogue positioning method can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision and embodied intelligence, and particularly relates to an embodied dialogue localization method and related device based on a causal-guided diffusion model. Background Art

[0002] Embodied Dialogue Localization (EDL) refers to the task of helping an agent perform precise position localization in an environment through the combination of vision and natural language dialogue. Such tasks have broad application prospects in fields such as robot navigation, emergency rescue, and human-computer interaction. For example, in an emergency rescue scenario, an agent needs to quickly locate the target position through dialogue with humans; in a home service robot, the robot needs to accurately find the target object or position according to the user's voice instructions.

[0003] Currently, the existing embodied dialogue localization methods still face the following challenges:

[0004] (1) Strong resolution dependence: Existing methods usually model the localization problem as an image-to-image conversion problem and use an encoder-decoder architecture to generate a heatmap to predict coordinates. Although these methods perform well in the coarse-grained range, they have obvious deficiencies in precise localization; the heatmap method highly depends on the image resolution, and the improvement of the resolution will bring an exponential increase in computational complexity, limiting its feasibility in practical applications.

[0005] (2) Insufficient generalization ability: The performance of existing methods in unseen environments is poor. Especially when the data distribution is quite different from the training set, the localization accuracy of the model drops significantly. Although data augmentation or using large language models to generate additional dialogue data can improve the generalization ability to a certain extent, these methods are still limited by the inherent bias of the dataset and cannot fundamentally solve the problem.

[0006] (3) Interference from confounding factors: In visual and language inputs, there are a large number of observable and unobservable confounding factors (such as room type, decoration style, lighting conditions, sentence structure, etc.). These factors will cause the model to learn spurious correlations, thereby affecting the accuracy of localization. Existing methods lack an effective mechanism to handle these confounding factors, resulting in unstable performance of the model in complex environments.

[0007] In summary, the existing embodied dialogue localization methods have obvious deficiencies in terms of accuracy, generalization ability, and anti-interference ability. Therefore, designing a localization method that can reduce resolution dependence, improve generalization ability, and effectively handle confounding factors is of great significance for promoting the development of embodied intelligence. Summary of the Invention

[0008] The object of the present invention is to provide an embodied dialogue localization method and related device based on a causal-guided diffusion model, so as to solve one or more of the technical problems such as strong resolution dependence, insufficient generalization ability, and interference of confounding factors existing in the existing embodied dialogue localization methods.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] In the first aspect of the present invention, an embodied dialogue localization method based on a causal-guided diffusion model is provided, including the following steps:

[0011] Obtain the dialogue text and the top-view map of the scene to be located;

[0012] Based on the obtained dialogue text and top-view map, use the trained causal-guided diffusion model to perform localization prediction to obtain the coordinate localization result;

[0013] Among them, the causal-guided diffusion model includes:

[0014] A feature extraction module for extracting the text features of the dialogue text and the visual features of the top-view map;

[0015] A causal reasoning module for performing deconfounding processing on observable and unobservable confounding factors with the text features and visual features extracted by the feature extraction module as the original features to obtain deconfounded features;

[0016] A deconfounded feature guiding module for dynamically adjusting the weights of the original features and the deconfounded features to obtain the screened features;

[0017] A diffusion network for performing regression processing through a step-by-step denoising process with the screened features as the control condition to obtain the coordinate localization result.

[0018] A further improvement of the present invention lies in that

[0019] In the causal reasoning module, the step of performing deconfounding processing on observable and unobservable confounding factors with the text features and visual features extracted by the feature extraction module as the original features to obtain deconfounded features includes: first performing backdoor adjustment, and then performing frontdoor adjustment; among them, the backdoor adjustment eliminates bias by cutting off the backdoor path between the observable confounding factor and the input to obtain the preliminary deconfounded features; the frontdoor adjustment is based on the preliminary deconfounded features, and reduces the influence of unobservable confounding factors by introducing a mediating variable to transfer knowledge to obtain the final deconfounded features.

[0020] A further improvement of the present invention lies in that

[0021] In the step of the backdoor adjustment, the deconfounded visual features in the preliminary deconfounded features Deconfounded text features Are respectively expressed as:

[0022]

[0023] Wherein, LN represents layer normalization; φ v And φ i Both represent learnable fully connected layers, F V And F I Represent the extracted original visual features and text features; Represents the mathematical expectation of the confounding z;

[0024] Among them, the original visual features and text features are uniformly represented as X, then there is:

[0025]

[0026] Wherein, |z i | represents the number of confounding instances belonging to the i-th category in the confounding factor dictionary, ∑ j |z j | represents the total number of confoundings stored in the confounding dictionary; f(X, z) represents a neural network with parameters X and z.

[0027] A further improvement of the present invention lies in that

[0028] The steps for constructing the confounding dictionary include: separately processing text features and visual features to create a confounding dictionary; wherein, for text features, spatial directions and key landmark words are extracted from the dialogue, and then the average features are calculated according to the probability of each word appearing; for visual features, a pre-trained VQA model is used to ask "What kind of room is this?" to obtain each room type, and then the average features of each room type are calculated.

[0029] A further improvement of the present invention lies in that

[0030] In the steps of the front door adjustment, the mediating variable is designed as a feature selector based on the VQ-VAE model, and the deconfounded visual feature F V ' and the deconfounded text feature F I ' are respectively expressed as:

[0031]

[0032] Wherein, And Are quantization features in the VQ-VAE model;

[0033]

[0034] Wherein, Denote the cross-sampled features randomly sampled from the VQ-VAE codebook; Denote the intra-sampled features obtained by the VQ-VAE acting on the current input; two query sets

[0035] A further improvement of the present invention lies in that

[0036] During the training process of the causal-guided diffusion model, the diffusion network and the causal inference module are jointly trained, and the parameters are updated by minimizing the diffusion loss of the diffusion network and the VQ-VAE loss in the causal inference module to update the parameters, and the overall loss function is expressed as:

[0037]

[0038] In the formula, is the overall loss function; is the diffusion loss; is the VQ-VAE loss; γ 1 and γ 2 are respectively the weight coefficients of.

[0039] A further improvement of the present invention lies in that

[0040] In the inference stage of the diffusion network, based on the filtered features of the input, the initial noise coordinates y are sampled from the unit Gaussian distribution t , and the noise is gradually removed through the reverse denoising process to obtain the final coordinate prediction y 0 ;

[0041] Among them,

[0042]

[0043] In the formula, y t-1 represents the noise coordinates at time step t-1; α t represents a fixed sequence of mean coefficients; represents σ t represents the standard deviation of the noise; do(C) represents using the deconfounded features obtained by the causal inference module as the control condition; represents the noise predicted by the model at time step t and under the control condition C, and the noise is corrected by the deconfounded feature guidance module;

[0044]

[0045] In the formula, ε θ represents the uncorrected predicted noise.

[0046] In a second aspect of the present invention, there is provided an embodied dialogue localization system based on a causal-guided diffusion model, comprising:

[0047] A data acquisition module for acquiring dialogue text and a top-view map of the scene to be localized;

[0048] A localization prediction module for performing localization prediction using the trained causal-guided diffusion model based on the acquired dialogue text and top-view map to obtain a coordinate localization result;

[0049] Wherein, the causal-guided diffusion model comprises:

[0050] A feature extraction module for extracting text features of the dialogue text and visual features of the top-view map;

[0051] A causal reasoning module for performing observable and unobservable confounding factor elimination processing using the text features and visual features extracted by the feature extraction module as original features to obtain de-confounded features;

[0052] A de-confounded feature guiding module for dynamically adjusting the weights of the original features and the de-confounded features to obtain filtered features;

[0053] A diffusion network for obtaining a coordinate localization result through step-by-step denoising process regression processing with the filtered features as control conditions.

[0054] In a third aspect of the present invention, there is provided an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the embodied dialogue localization method based on the causal-guided diffusion model according to any one of the first aspects of the present invention.

[0055] In a fourth aspect of the present invention, there is provided a non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the embodied dialogue localization method based on the causal-guided diffusion model according to any one of the first aspects of the present invention.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] The technical solution disclosed by the present invention can effectively reduce the dependence on high resolution and improve the positioning accuracy by introducing a causal-guided diffusion model. Among them, the causal inference module eliminates the confounding factors in the dataset, enhancing the robustness and generalization ability of the model in unseen environments. Specifically, in view of the problem that existing methods usually rely on the generation of high-resolution heatmaps, resulting in high computational complexity and insufficient precise positioning ability, the present invention introduces a diffusion network to directly model the continuous coordinate distribution, avoiding the resolution limitation in the heatmap generation process. Through a step-by-step denoising process, the diffusion network can accurately regress the coordinates, reducing the dependence on high-resolution inputs. The present invention can still achieve high-precision positioning under low-resolution conditions, significantly reducing the computational complexity while improving the positioning accuracy, especially outperforming existing methods in the fine-grained range.

[0058] In view of the problem that the performance of existing methods is poor in unseen environments, especially when the data distribution is quite different from that of the training set and the positioning accuracy of the model drops significantly, in the preferred solution of the present invention, by introducing a causal inference module (including backdoor adjustment and frontdoor adjustment), the confounding factors in the dataset (such as room type, decoration style, lighting conditions, etc.) are effectively eliminated, reducing the overfitting of the model to the training data. The positioning accuracy of the present invention in unseen environments is significantly improved, showing stronger generalization ability and being able to adapt to diverse application scenarios. In view of the problem that existing methods are easily interfered by confounding factors (such as room type, decoration style, lighting conditions, sentence structure, etc.) in visual and language inputs, resulting in the model learning false correlations and affecting the accuracy of positioning, the present invention uses backdoor adjustment (BDA) to handle observable confounding factors and frontdoor adjustment (FDA) to handle unobservable confounding factors, ensuring that the model learns true causal relationships. The present invention can effectively eliminate the interference of confounding factors, improve the robustness of the model in complex environments, and ensure the accuracy of positioning results. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art; obviously, the following-described drawings are some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0060] Figure 1 It is a schematic flowchart of an embodied dialogue positioning method based on a causal-guided diffusion model provided by an embodiment of the present invention;

[0061] Figure 2 It is a schematic diagram of the processing logic of the causal-guided diffusion model in an embodiment of the present invention;

[0062] Figure 3 This is an exemplary analysis schematic diagram of the causal-guided diffusion model in an embodiment of the present invention; among them, Figure 3 in (a) is a schematic diagram of the noise reduction process, Figure 3 in (b) is a schematic diagram of the language interference indication, Figure 3 in (c) is a schematic diagram of the visual interference indication;

[0063] Figure 4 This is a schematic diagram of the network structure of the causal-guided diffusion model in an embodiment of the present invention;

[0064] Figure 5 This is a schematic diagram of the causal graph of the backdoor adjustment and the frontdoor adjustment in an embodiment of the present invention; among them, Figure 5 in (a) is a schematic diagram of the causal graph of the backdoor adjustment, Figure 5 in (b) is a schematic diagram of the causal graph of the frontdoor adjustment;

[0065] Figure 6 This is a schematic diagram of the cumulative matching curve in an embodiment of the present invention;

[0066] Figure 7 This is a schematic diagram of an embodied dialogue localization system based on the causal-guided diffusion model provided by an embodiment of the present invention. Detailed implementation manners

[0067] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention; obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments.

[0068] Based on the technical solutions disclosed in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0069] Please refer to Figures 1 to 3 , an embodied dialogue localization method based on the causal-guided diffusion model provided by an embodiment of the present invention, includes the following steps:

[0070] Step 1, obtain the dialogue text and the top-view map of the scene to be located;

[0071] Step 2: Based on the dialogue text and the top - view map obtained in Step 1, use the trained Causality Guided Diffusion (CGD) model to perform positioning prediction and obtain accurate coordinate positioning results;

[0072] Among them, the causality guided diffusion model includes:

[0073] A feature extraction module for extracting the text features of the dialogue text and the visual features of the top - view map;

[0074] A causal reasoning module for using the text features and visual features extracted by the feature extraction module as the original features to perform the elimination process of observable and unobservable confounding factors and obtain de - confounded features;

[0075] A de - confounded feature guidance module for dynamically adjusting the weights of the original features and the de - confounded features to obtain filtered features. Explanatorily, in the embodiments of the present invention, the de - confounded features are used as the guidance for the denoising process to enhance the generalization ability of the model;

[0076] A diffusion network, with the filtered features as the control condition, and through the step - by - step denoising process regression, obtain the coordinate positioning results.

[0077] In the embodiments of the present invention, aiming at the deficiencies existing in the existing methods in cross - modal alignment and precise positioning tasks, especially problems such as the dependence of traditional methods on high resolution, poor generalization ability caused by data bias, and the limitations of the heat - map method in fine positioning, etc., a precise embodied dialogue positioning method based on a causality guided diffusion model is proposed. This method can significantly improve the positioning accuracy by introducing a causal reasoning module and combining the denoising process of the diffusion network, and effectively reduce the impact of data bias on the model performance. Explanatorily, the diffusion network is used for step - by - step regression of coordinate prediction. The diffusion network gradually adds noise through the forward process and then gradually denoises through the reverse process to finally obtain accurate coordinate prediction; the causal reasoning module is used to process observable and unobservable confounding factors. Input the top - view map and the dialogue text, and encode them into feature vectors through pre - trained visual and text encoders; then, gradually add noise through the forward process of the diffusion network and gradually denoise through the reverse process to finally obtain accurate coordinate prediction.

[0078] In a specific embodiment of the present invention, in the causal reasoning module, the step of performing deconfounding factor elimination processing on the text features and visual features extracted by the feature extraction module as the original features to obtain deconfounded features includes: back-door adjustment (BDA) and front-door adjustment (FDA). BDA eliminates bias by cutting off the back-door path between the confounding factor and the input, while FDA reduces the influence of unobservable confounding factors by introducing mediating variables to transfer knowledge.

[0079] Please refer to Figure 5 , in the embodiment of the present invention, for convenience, the visual sample V and the text sample I are uniformly represented as X, the coordinates to be predicted are represented as Y, and the confounding factors are uniformly represented as

[0080] 1) Back-door adjustment (BDA) includes: First, the text and visual features are processed separately to create a confounding dictionary; for text features, the present invention extracts spatial directions and key landmark words from the dialogue, and then calculates the average feature according to the probability of each word appearing; for visual features, the pre-trained VQA model BLIP is used to obtain each room type by asking "What kind of room is this?", and then the average feature of each room type is calculated.

[0081] According to Bayes' theorem, the prediction without considering the confounding factor can be expressed as:

[0082]

[0083] Among them, p(z|X) may cause the model to learn spurious correlations; for example, since most sofas are gray and placed in the living room, the model can easily learn the spurious association between "gray sofa" and "living room"; The Do operator provides a method to eliminate observable confounding factors by cutting off the back-door link between and X, and this process can be modeled by the neural network f(X,z), which can be expressed as follows:

[0084]

[0085] According to the additivity of expectation, f(X,z) can be expressed as f x (X)+f z (z), which enables the causal relationship to be expressed as

[0086] Next, statistical techniques can be used to calculate

[0087]

[0088] where |z i | represents the number of confounding instances belonging to the i-th category in the confounding factor dictionary, and ∑ j |z j | represents the total number of confoundings accessed in the confounding dictionary.

[0089] Finally, the deconfounded visual features and the deconfounded text features are obtained as follows:

[0090]

[0091] where φ v and φ i represent learnable fully connected layers, F V and F I represent the original visual and text features extracted by the backbone network, and LN represents layer normalization.

[0092] 2) Front-door adjustment (FDA) includes: Back-door adjustment requires prior identification of confounding factors and construction of a confounding dictionary; however, some unobservable confounding factors cannot be directly modeled, which can also lead to bias; to solve this problem, front-door adjustment introduces an additional mediator between X and Y, which creates a front-door path to transfer knowledge:

[0093]

[0094] where m represents the knowledge selected from the mediator and

[0095] Considering the sensitivity of the embodied dialogue localization task to regions, is designed as a VQ-VAE-based feature selector. Specifically, since both images and texts have been pre-encoded as tokens, VQ-VAE is used to project these tokens into the latent space respectively, effectively performing implicit clustering of features; through the learned VQ-VAE, V and I are represented by corresponding discrete coding sequences; subsequently, key features can be extracted from X using the VQ-VAE model, and then these features are used to predict the coordinates Y.

[0096] Using to represent the inner-sampled features obtained by applying VQ-VAE to the current input, and x to represent the cross-sampled features randomly sampled from the VQ-VAE codebook; based on the linear mapping model, the above equation becomes Two embedding functions are used to transmit the input X to two query sets and Then the front-door adjustment It can be approximated as follows:

[0097]

[0098] Represent the quantization features in VQ-VAE as and The final deconfounded features obtained through backdoor adjustment and frontdoor adjustment can be expressed as:

[0099] In a specific embodiment of the present invention, in the diffusion network, the process of performing denoising on the text features and visual features extracted by the feature extraction module as the original features includes:

[0100] Given the top-down map sample M and the dialogue I, first encode them into tokens through pre-trained visual and text encoders respectively, and then project them onto the same dimension, denoted as F V and F I . Next, the coordinates are gradually regressed through the forward and backward processes of the diffusion network.

[0101] First, uniformly sample the time step t from {0,..., T-1}, where T represents the maximum range of the time step; regard the real two-dimensional coordinates y as y in the diffusion network 0 . At the sampled time step t, add independent Gaussian noise to y 0 to obtain the perturbed noise coordinates y t :

[0102]

[0103] where, is a fixed noise sequence.

[0104] After obtaining the noise coordinates y t , use MLP to encode them into tokens, denoted as Then connect the visual token F V and the text token F I to form the condition C. Input the condition C and the noise token together into a standard Transformer. Then extract the noise token from the output of the Transformer and convert it into the predicted noise ε θ through the regression head. By minimizing the mean square error (MSE) loss between the predicted noise and the real noise, the model is gradually optimized during training:

[0105]

[0106] In a specific embodiment of the present invention, in the deconfounding feature guidance module, the process of using the deconfounding features output by the causal inference module as the guidance of the diffusion network includes:

[0107] To simplify the notation, the backdoor adjustment and the frontdoor adjustment in the causal inference module are uniformly denoted as do(C). The classifier-free guidance uses the implicit classifier gradient to adjust the gradient direction in the diffusion network. Inspired by this method, the implicit causal intervention can be expressed as p i (do(C)|y t ) ∝ p(y t |do(C)) / p(y t ), to mitigate the potential adversarial gradient behavior in the causal inference process, where i represents implicit. Using the guidance of this implicit intervention, the diffusion network score estimation is updated as:

[0108]

[0109] where, represents the corrected noise.

[0110] In a specific embodiment of the present invention, the training and inference processes include:

[0111] 1) Training: By jointly training the diffusion network and the causal inference module, minimize the diffusion loss and the VQ-VAE loss

[0112]

[0113] where, γ 1 and γ 2 are weight coefficients.

[0114] 2) Inference: In the inference stage, by sampling the initial noise coordinates y t from the unit Gaussian distribution, and gradually removing the noise through the reverse denoising process, the final coordinate prediction y 0 is obtained:

[0115]

[0116] where

[0117] Figure 4It is a schematic diagram of the main structural framework and working principle of the method in the embodiments of the present invention. Looking from the upper part of the figure, the framework of the present invention mainly consists of a diffusion network and a causal reasoning module. The diffusion network gradually regresses the coordinate prediction through a process of gradually adding noise and denoising, while the causal reasoning module processes observable and unobservable confounding factors through backdoor adjustment (BDA) and frontdoor adjustment (FDA) respectively, ensuring that the model can avoid the interference of irrelevant information during the denoising process. Specifically, the left part of the figure shows the working process of the diffusion network: First, the input top-view map and dialogue text are encoded into feature vectors through pre-trained visual and text encoders; then, noise is gradually added through the forward process of the diffusion network and gradually removed through the backward process, and finally an accurate coordinate prediction is obtained. The causal reasoning module processes observable confounding factors (such as room type, keywords, etc.) and unobservable confounding factors (such as decoration style, lighting conditions, etc.) through backdoor adjustment and frontdoor adjustment respectively, generating deconfounded features as guidance for the diffusion network. The lower part of the figure shows the specific implementation details of the causal reasoning module. Through backdoor adjustment, the model can cut off the backdoor path between the confounding factor and the input, eliminating the influence of observable confounding factors; while through frontdoor adjustment, the model introduces a mediating variable to transmit the causal relationship between the input and the output, further reducing the influence of unobservable confounding factors. Finally, the deconfounded features are used to guide the denoising process of the diffusion network, ensuring that the model is more robust during the denoising process and avoiding the interference of irrelevant information. It can be seen from the figure that whether dealing with observable or unobservable confounding factors, the causal reasoning module can effectively reduce the influence of data bias, thereby improving the generalization ability of the model. Especially in unseen environments, the method of the present invention significantly improves the positioning accuracy through a causally guided denoising process, demonstrating its powerful capabilities in complex and unknown environments. Summarily, the present invention proposes an accurate embodied dialogue localization method based on a causally guided diffusion network, aiming to solve the deficiencies in cross-modal alignment and precise positioning in the prior art. By introducing a causal reasoning module and combining it with a diffusion network, this method significantly improves the positioning accuracy and reduces the impact of data bias on the model performance.

[0118] Please refer to Figure 6 , in the specific embodiments of the present invention, experiments on the WAY dataset compared the performance under the valSeen (seen environment) and valUnseen (unseen environment) datasets respectively. The present invention selected two evaluation metrics, namely Acc0 (the accuracy rate when the distance between the predicted coordinate and the true coordinate is 0 meters) and Acc5 (the accuracy rate when the distance between the predicted coordinate and the true coordinate is less than or equal to 5 meters). Among these two metrics, the larger the value, the better the positioning effect.

[0119] Table 1 shows the experimental results of the method of the present invention on the WAY dataset. As can be seen from Table 1, the method of the present invention has achieved significant performance improvements on both the valSeen and valUnseen datasets. Especially on the valUnseen dataset, the Acc0 and Acc5 of the method of the present invention in the single-turn conversation setting are respectively improved by 9.7% and 29.15% compared to the existing state-of-the-art method (DiaLoc), and are respectively improved by 3.32% and 15.73% in the multi-turn conversation setting. These results indicate that the method of the present invention not only performs well in the seen environment, but also demonstrates strong generalization ability in the unseen environment. Summarily, the present invention reduces the dependence on high resolution by directly modeling the continuous coordinate distribution, effectively improves the positioning accuracy, and eliminates the confounding factors in the dataset through the causal reasoning module, enhancing the robustness and generalization ability of the model in the unseen environment; the experimental results show that the present invention is superior to the prior art in multiple evaluation metrics, especially outstanding in the unseen environment.

[0120] Table 1. Experimental results of the method on the WAY dataset

[0121]

[0122] In summary, the embodiment of the present invention discloses an accurate embodied conversation localization method based on a causal-guided diffusion model. By directly modeling the continuous coordinate distribution through a diffusion network in the local view, it reduces the dependence on high resolution and can achieve precise positioning within a fine range; in the global view, it processes observable and unobservable confounding factors through a causal reasoning module to ensure that the model can avoid interference from irrelevant information during the denoising process, thereby improving the generalization ability of the model; through the causal-guided denoising process, the present invention demonstrates strong positioning ability in complex and unknown environments, significantly improving the positioning accuracy and the generalization performance of the model. The present invention not only provides a new solution for cross-modal alignment and precise positioning tasks, but also lays a solid foundation for future applications in more complex scenarios.

[0123] The following is the device embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For details not disclosed in the device embodiment, please refer to the method embodiment of the present invention.

[0124] Please refer to Figure 7 , in the embodiment of the present invention, there is provided an embodied conversation localization system based on a causal-guided diffusion model, including:

[0125] A data acquisition module, configured to acquire conversation texts and a top-view map of the scene to be located;

[0126] A positioning prediction module, configured to perform positioning prediction using a trained causal-guided diffusion model based on the obtained dialogue text and top-view map, and obtain a coordinate positioning result;

[0127] Wherein, the causal-guided diffusion model includes:

[0128] A feature extraction module, configured to extract text features of the dialogue text and visual features of the top-view map;

[0129] A causal reasoning module, configured to perform observable and unobservable confounding factor elimination processing using the text features and visual features extracted by the feature extraction module as original features, and obtain de-confounded features;

[0130] A de-confounded feature guiding module, configured to dynamically adjust the weights of the original features and the de-confounded features to obtain filtered features;

[0131] A diffusion network, configured to use the filtered features as control conditions and perform regression processing through a step-by-step denoising process to obtain a coordinate positioning result.

[0132] In an embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program. The computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used to execute the operations of the embodied dialogue positioning method based on the causal-guided diffusion model.

[0133] In one embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. Moreover, one or more instructions suitable for being loaded and executed by the processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM (Random Access Memory) or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the method for embodied dialogue localization based on the causal-guided diffusion model in the above embodiment.

[0134] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, optical memories, etc.) containing computer-usable program codes.

[0135] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0136] These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions in the flow Figure 1One or more processes and / or boxes Figure 1 The functions specified in one box or more boxes.

[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one box or more boxes.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. An embodied dialogue localization method based on a causal guided diffusion model, characterized in that: The following steps are involved: Get the dialogue text and the top view map of the scene to be located; Based on the acquired conversation text and bird's-eye view map, the trained causal guided diffusion model is used to perform positioning prediction and obtain coordinate positioning results; Wherein, the causal guided diffusion model includes: A feature extraction module, used to extract text features of the conversation text and visual features of the overhead map; A causal reasoning module, used to use the text features and visual features extracted by the feature extraction module as original features, perform observable and unobservable confounding factor elimination processing, and obtain decongested features; The decongested feature guidance module is used to dynamically adjust the weights of the original features and decongested features to obtain the filtered features; The diffusion network is used to obtain the coordinate positioning result by stepwise denoising process regression processing with the screened features as control conditions.

2. According to claim 1, the method for embodied dialogue localization based on the causal guided diffusion model is characterized in that: In the causal reasoning module, the text features and visual features extracted by the feature extraction module are used as original features to perform observable and unobservable confounding factor elimination processing, and the steps of obtaining deconfounding features include: first performing backdoor adjustment and then performing frontdoor adjustment; wherein the backdoor adjustment eliminates the bias by cutting off the backdoor path between the observable confounding factor and the input to obtain a preliminary deconfounding feature; the frontdoor adjustment is based on the preliminary deconfounding feature, and the knowledge is transferred by introducing an intermediary variable to reduce the influence of the unobservable confounding factor to obtain the final deconfounding feature.

3. The method for embodied dialogue localization based on a causal guided diffusion model according to claim 2, characterized in that: In the step of backdoor adjustment, the de-congested visual features in the preliminary de-congested features De-cluttering text features Respectively expressed as: Where LN represents layer normalization; φ v and φ i Both represent learnable fully connected layers, F V and F I Represents the extracted original visual features and text features; represents the mathematical expectation of the hybrid z; Among them, the original visual features and text features are uniformly represented as X, then: In the formula, |z i | represents the number of confounding instances belonging to the i-th category in the confounding factor dictionary, ∑ j |z j | represents the total number of confounding elements accessed in the confounding dictionary; f(X,z) represents a neural network with parameters X and z.

4. The method for embodied dialogue localization based on a causal guided diffusion model according to claim 3, characterized in that: The steps of constructing the hybrid dictionary include: processing text features and visual features separately to create a hybrid dictionary; wherein, for text features, extracting spatial directions and key landmark words from the conversation, and then calculating the average feature according to the probability of each word appearing; for visual features, using a pre-trained VQA model to ask "What room is this?" to obtain each room type, and then calculating the average feature of each room type.

5. The method for embodied dialogue localization based on a causal guided diffusion model according to claim 3, characterized in that: In the step of front-gate adjustment, the mediating variable is designed as a feature selector based on the VQ-VAE model, and the decongested visual feature F in the final decongested feature V ′, de-mixed text features F I ′ are respectively expressed as: In the formula, and It is the quantitative feature in the VQ-VAE model; In the formula, represents the cross-sampled features randomly sampled from the VQ-VAE codebook; Represents the internal sampling features obtained by VQ-VAE acting on the current input; Two query sets 6. The method for embodied dialogue localization based on a causal guided diffusion model according to claim 5, characterized in that: During the training process of the causal guided diffusion model, the diffusion network and the causal reasoning module are jointly trained to minimize the diffusion loss of the diffusion network and the VQ-VAE loss in the causal reasoning module. To update the parameters, the overall loss function is expressed as: In the formula, is the overall loss function; is the diffusion loss; is the VQ-VAE loss; γ 1 and γ 2 They are The weight coefficient of .

7. The method for embodied dialogue localization based on a causal guided diffusion model according to claim 6, characterized in that: The inference phase of the diffusion network is based on the filtered features of the input by sampling the initial noise coordinate y from the unit Gaussian distribution. t , and gradually remove the noise through the reverse denoising process to obtain the final coordinate prediction y0; in, In the formula, y t-1 represents the noise coordinates at time step t-1; α t represents a fixed mean coefficient sequence; express σ t represents the standard deviation of the noise; do(C) means taking the deconfounding features obtained by the causal inference module as a control condition; represents the noise predicted by the model at time step t and control condition C, and the noise is corrected by the de-confounding feature guidance module; In the formula, ε θ represents the uncorrected prediction noise.

8. An embodied dialogue localization system based on a causal guided diffusion model, characterized in that: include: A data acquisition module, used to acquire the conversation text and a top-view map of the scene to be located; The positioning prediction module is used to perform positioning prediction based on the acquired conversation text and the bird's-eye view map using the trained causal guided diffusion model to obtain the coordinate positioning result; Wherein, the causal guided diffusion model includes: A feature extraction module is used to extract text features of the conversation text and visual features of the overhead map; A causal reasoning module, used to use the text features and visual features extracted by the feature extraction module as original features, perform observable and unobservable confounding factor elimination processing, and obtain decongested features; The decongested feature guidance module is used to dynamically adjust the weights of the original features and decongested features to obtain the filtered features; The diffusion network is used to obtain the coordinate positioning result by stepwise denoising process regression processing with the screened features as control conditions.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the embodied dialogue localization method based on the causal guided diffusion model as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the embodied dialogue localization method based on a causal guided diffusion model as described in any one of claims 1 to 7 is implemented.