System and method for performing identification based on scene representation and information related to the scene
The method enhances multimodal AI systems by refining identification results through iterative dialogue turns, addressing the challenge of incomplete information in real-time vision-language tasks.
Patent Information
- Application Number
- GB2024002970
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2025-09-03
AI Technical Summary
Existing multimodal artificial intelligence systems for vision-language tasks often require complete information prior to processing, which is impractical or impossible in real-time applications like real-time dialogue-based identification.
A computer-implemented method that processes scene representations and information in stages, using unimodal and multimodal encoders and decoders to refine identification results through iterative dialogue turns, allowing for accurate and efficient identification even with incomplete information.
Enables accurate and efficient identification by updating results with each dialogue turn, improving accuracy and reducing error propagation, suitable for real-life applications where complete information is not available.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD Embodiments described herein relate to system and method for performing identification based on, at least, a representation of a scene (scene representation) and information related to the scene. BACKGROUND Multimodal artificial intelligence based systems and methods can be used to process and integrate data of different modes to facilitate reasoning or control. For example, multimodal machine-learning-based systems and methods can be used to process and integrate vision-based data and language-based data to facilitate performing of visionlanguage task. Some existing multimodal machine-learning-based systems and methods for performing vision-language task assume availability of complete information related to the visionlanguage task, which may be impractical or impossible for some applications. For example, for vision-language task such as real-time dialogue-based identification, the entire dialogue may not be available prior to processing using the systems and methods. SUMMARY Embodiments described herein attempt to address one or more of the above problems associated with some existing multimodal artificial intelligence based systems and methods, or, more generally, to provide an improved system and method for performing identification such as localisation. In a first aspect, there is provided a computer-implemented method for performing identification. The computer-implemented method comprises receiving a representation of a scene, receiving a first information related to the scene for performing identification on the representation of the scene, processing at least the representation of the scene and the first information to obtain a first output associated with an identification result, receiving a second information related to the scene for performing identification on the representation of the scene, and processing at least the first output and the second information to obtain a second output associated with an updated identification result. The first information may include at least one query (e.g., at least one question) and at least one response to the at least one query (e.g., at least one answer). The second information may include at least one query (e.g., at least one question) and at least one response to the at least one query (e.g., at least one answer). The first output and / or the identification result may be used to determine whether an identification task has been completed (i.e., whether processing of further information related to the scene is necessary) or to affect content of the second information for performing the identification task. The second output and / or the updated identification result may be used to determine whether an identification task has been completed or to affect content of further (third) information for performing the identification task. The first output may be different from the second output and / or the updated identification result may be different from the identification result (previous or initial identification result). In some examples, the updated identification result is or includes a subset of the identification result (previous or initial identification result). The computer-implemented method takes into account at least the first information (embodied in the first output) and the second information to update, e.g., refine, the identification result such that the updated identification result may become more accurate or more useful than the identification result (previous or initial identification result). The computer-implemented method can provide / update identification result based on each respective (subsequent) information to facilitate performing of the identification task. This may be useful for some real-life practical identification applications in which the complete set of information is not prior known. In some embodiments, the processing of at least the representation of the scene and the first information comprises extracting features from the representation of the scene, extracting features from the first information, and fusing at least the features extracted from the representation of the scene and the features extracted from the first information to obtain the first output. The extraction and fusion operations can help to obtain useful features for facilitating performing of the identification task. In some embodiments, the extraction of features from the representation of the scene is performed using a first encoder. The first encoder may be part of a machine-learning-based identification model. The representation of the scene is data of a first modality, and the first encoder may be a unimodal encoder. A unimodal encoder has a relatively simple construction and may be trained relatively easily. In some examples, the representation of the scene comprises a visual representation or a semantic representation of the scene. In one example, the visual representation of the scene comprises one or more images (e.g., one or more maps) and the first encoder comprises an image encoder. In one example, the visual representation of the scene comprises one or more point clouds, and the first encoder comprises a point cloud encoder. In one example, the semantic representation of the scene comprises one or more graphs such as one or more scene graphs, and the first encoder comprises a graph encoder such as a scene graph encoder. In some examples, the first encoder is a transformer-based encoder. In some embodiments, the extraction of features from the first information is performed using a second encoder. The second encoder may be part of a / the machine-learning-based identification model. The first information may be data of a second modality different from the first modality, and the second encoder may be a unimodal encoder. A unimodal encoder has a relatively simple construction and may be trained relatively easily. In some examples, the first information comprises language-based or visionbased information and the second encoder comprises a language-based data encoder or a vision-based data encoder. In some examples, the first information is languagebased and it comprises text data and the second encoder comprises a text encoder. The text data may be generated from speech data using a speech-to-text converter. The text data may be received via an input device (e.g. keyboard, touch-screen). In some examples, the first information is language-based and it comprises speech data and the second encoder comprises a speech encoder. In some examples, the first information is vision-based and it comprises vision-based data (related to, e.g., pictogram, sign language, etc.) and the second encoder comprises a vision-based data encoder. In some examples, the second encoder is a transformer-based encoder. In some embodiments, the fusing of at least the features extracted from the representation of the scene and the features extracted from the first information is performed using a third encoder. The third encoder may be part of a / the machine learning-based identification model. The third encoder may be a multimodal fusion encoder, e.g., a bimodal fusion encoder. The multimodal fusion encoder can produce a useful output based on data of different modalities. In some examples, the third encoder is a transformer-based fusion encoder. In some embodiments, the computer-implemented method further comprises processing the first output to obtain a representation of the identification result. In one example, the representation of the identification result is machine interpretable. In one example, the representation of the identification result is human interpretable. In one example, the representation of the identification result comprises a visual representation that is human interpretable. The processing of the first output may be performed using a first decoder. The first decoder may be part of a / the machine-learning-based identification model. In some embodiments, the first decoder comprises a heatmap decoder, and the visual representation of the identification result comprises a heatmap. A heatmap can be interpreted relatively intuitively. In some examples, the first decoder is a convolutional-neural-network-based decoder. In some embodiments, the computer-implemented method further comprises outputting the visual representation of the identification result for display. In some embodiments, the computer-implemented method further comprises displaying the visual representation of the identification result to affect the content of the second information. For example, the output or displayed visual representation of the identification result may guide the performing of a subsequent dialogue which affects the second information. For example, the output or displayed visual representation of the identification result can be used to facilitate formulation of at least one query (e.g., at least one question) in a subsequent dialogue. Alternatively, the output or displayed visual representation of the identification result can be used to facilitate determination of whether the identification task has been or can be completed or whether the receiving or processing of further information related to the scene is required. In some cases, the identification task is deemed to have been completed when it is determined that the estimated identification result is correct or accurate or useful enough. In some cases, the identification task is deemed to have been completed if no further information related to the scene is received or no further dialogue occurs, e.g., within a period of time. In some embodiments, the processing of at least the first output and the second information comprises processing only the first output and the second information to obtain the second output. The processing of only the first output (which is derived based on the representation of the scene and the first information) and the second information may allow the processing be performed more efficiently. In some embodiments, the processing of at least the first output and the second information comprises processing at least the representation of the scene, the first output, and the second information to obtain the second output. The processing of the representation of the scene along with the first output (derived based on the representation of the scene and the first information) and the second information, i.e., with the representation of the scene explicitly applied in the processing, may improve the accuracy of the updated identification result and / or reduce the chance or effect of propagation of error in determining the updated identification result. In some embodiments, the processing of at least the first output and the second information comprises extracting features from the second information, and fusing at least the first output and the features extracted from the second information to obtain the second output. The extraction and fusion operations can help to obtain useful features for facilitating performing of the identification task. In some embodiments, the fusing of at least the first output and the features extracted from the second information comprises fusing only the first output and the features extracted from the second information to obtain the second output. The fusing of only the first output and the features extracted from the second information may allow the processing be performed more efficiently. In some embodiments, the fusing of at least the first output and the features extracted from the second information comprises fusing the features extracted from the representation of the scene, the first output, and the features extracted from the second information to obtain the second output. The fusing of the features extracted from the representation of the scene along with the first output (derived based on the representation of the scene and the first information) and the second information, i.e., with the representation of the scene explicitly applied in the fusion, may improve the accuracy of the updated identification result and / or reduce the chance or effect of propagation of error in determining the updated identification result. In some embodiments, the extraction of the features of the second information is performed using a fourth encoder. The fourth encoder may be part of a / the machinelearning-based identification model. The second information may be data of the same modality as the first information, and the fourth encoder may be a unimodal encoder. The fourth encoder and the second encoder may be of the same type. A unimodal encoder has a relatively simple construction and may be trained relatively easily. In some examples, the second information comprises language-based or vision-based information and the fourth encoder comprises a language-based data encoder or a vision-based data encoder. In some examples, the second information is languagebased and it comprises text data and the fourth encoder comprises a text encoder. The text data may be generated from speech data using a speech-to-text converter. The text data may be received via an input device (e.g. keyboard, touch-screen). In some examples, the second information is language-based and it comprises speech data and the fourth encoder comprises a speech encoder. In some examples, the second information is vision-based and it comprises vision-based data (related to, e.g., pictogram, sign language, etc.) and the fourth encoder comprises a vision-based data encoder. In some examples, the fourth encoder is a transformer-based encoder. In some embodiments, the fusing of at least the first output and the features extracted from the second information is performed using a fifth encoder. The fifth encoder may be part of a / the machine-learning-based identification model. The fifth encoder may be a multimodal fusion encoder, e.g., a bimodal fusion encoder. The fifth encoder and the third encoder may be of the same type. The multimodal fusion encoder can produce a useful output based on data of different modalities. In some examples, the fifth encoder is a transformer-based fusion encoder. In some embodiments, the computer-implemented method further comprises processing the second output to obtain a representation of the updated identification result. In one example, the representation of the updated identification result is machine interpretable. In one example, the representation of the updated identification result is human interpretable. In one example, the representation of the updated identification result comprises a visual representation that is human interpretable. The processing of the second output may be performed using a second decoder. The second decoder may be part of a / the machine-learning-based identification model. The second decoder and the first decoder may be of the same type. In some embodiments, the second decoder comprises a heatmap decoder, and the visual representation of the updated identification result comprises a heatmap. A heatmap can be interpreted relatively intuitively. In some examples, the second decoder is a convolutional-neural-network-based decoder. In some embodiments, the computer-implemented method further comprises outputting the visual representation of the updated identification result for display. In some embodiments, the computer-implemented method further comprises displaying the visual representation of the updated identification result to affect the content of a third information. For example, the output or displayed visual representation of the updated identification result may guide the performing of a subsequent dialogue which affects the third information. For example, the output or displayed visual representation of the updated identification result can be used to facilitate formulation of at least one query (e.g., at least one question) in the subsequent dialogue. Alternatively, the output or displayed visual representation of the updated identification result can be used to facilitate determination of whether the identification task has been or can be completed or whether the receiving or processing of further information related to the scene is required. In some cases, the identification task is deemed to have been completed when it is determined that the updated identification result is correct or accurate or useful enough. In some cases, the identification task is deemed to have been completed if no further information related to the scene is received or no further dialogue occurs, e.g., within a period of time. In some embodiments, the second information is received after the first output is obtained and the second information is based at least in part on the identification result. The computer-implemented method is performed for at least two information (i.e., the first information and the second information, e.g., the first and second dialogues) to provide (and update) outputs associated with the identification results. The computer-implemented method can be performed for further information (e.g., dialogue) to further update the identification result. For example, in some embodiments, the computer-implemented method further comprises: receiving a third information related to the scene for performing identification on the representation of the scene, and processing at least the second output and the third information to obtain a third output associated with a further updated identification result. The third information includes at least one query (e.g., at least one question) and at least one response to the at least one query (e.g., at least one answer). The third output and / or the further updated identification result may be used to determine whether an identification task has been completed (i.e., whether the receiving or processing of further information related to the scene is necessary) or to affect the content of further information received. The third output may be different from the first or second output. The further updated identification result may be different from the updated identification result or the identification result. In some examples, the further updated identification result is or includes a subset of the updated identification result. In some examples, the further updated identification result is or includes a subset of the identification result. In some embodiments, the processing of at least the second output and the third information comprises processing only the second output and the third information to obtain the third output. In some other embodiments, the processing of at least the second output and the third information comprises processing at least the representation of the scene, the second output, and the third information to obtain the third output. In some embodiments, the processing of at least the second output and the third information comprises extracting features from the third information, and fusing at least the second output and the features extracted from the third information to obtain the third output. In some embodiments, the fusing of at least the second output and the features extracted from the third information comprises fusing only the second output and the features extracted from the third information to obtain the third output. In some other embodiments, the fusing of at least the second output and the features extracted from the third information comprises fusing the features extracted from the representation of the scene, the second output, and the features extracted from the third information to obtain the third output. In some embodiments, the extraction of the features of the third information is performed using a sixth encoder. The sixth encoder may be part of a / the machine-learning-based identification model. The third information may be data of the same modality as the first information and / or the second information, and the sixth encoder may be a unimodal encoder. The sixth encoder, the fourth encoder, and the second encoder may be of the same type. In some examples, the third information comprises language-based or visionbased information and the sixth encoder comprises a language-based data encoder or a vision-based data encoder. In some examples, the third information is language-based and it comprises text data and the sixth encoder comprises a text encoder. The text data may be generated from speech data using a speech-to-text converter. The text data may be received via an input device (e.g. keyboard, touch-screen). In some examples, the third information is language-based and it comprises speech data and the sixth encoder comprises a speech encoder. In some examples, the third information is vision-based and it comprises vision-based data (related to, e.g., pictogram, sign language, etc.) and the sixth encoder comprises a vision-based data encoder. In some examples, the sixth encoder is a transformer-based encoder. In some embodiments, the fusing of at least the second output and the features extracted from the third information is performed using a seventh encoder. The seventh encoder may be part of a / the machine-learning-based identification model. The seventh encoder may be a multimodal fusion encoder, e.g., a bimodal fusion encoder. The seventh encoder, the fifth encoder, and the third encoder may be of the same type. The multimodal fusion encoder can produce a useful output based on data of different modalities. In some examples, the seventh encoder is a transformer-based fusion encoder. In some embodiments, the computer-implemented method further comprises processing the third output to obtain a representation of the further updated identification result. In one example, the representation of the further updated identification result is machine interpretable. In one example, the representation of the further updated identification result is human interpretable. In one example, the representation of the further updated identification result comprises a visual representation that is human interpretable. The processing of the third output may be performed using a third decoder. The third decoder may be part of a / the machine-learning-based identification model. The third decoder, the second decoder and the first decoder may be of the same type. In some embodiments, the third decoder comprises a heatmap decoder, and the visual representation of the further updated identification result comprises a heatmap. In some examples, the third decoder is a convolutional-neural-network-based decoder. In some embodiments, the computer-implemented method further comprises outputting the visual representation of the further updated identification result for display. In some embodiments, the computer-implemented method further comprises displaying the visual representation of the further updated identification result to affect the content of a fourth information. For example, the output or displayed visual representation of the further updated identification result may guide the performing of a subsequent dialogue. For example, the output or displayed visual representation of the further updated identification result can be used to facilitate formulation of at least one query (e.g., at least one question) in the subsequent dialogue. Alternatively, the output or displayed visual representation of the further updated identification result can be used to facilitate determination of whether the identification task has been or can be completed or whether the receiving or processing of further information related to the scene is required. In some cases, the identification task is deemed to have been completed when it is determined that the further updated identification result is correct or accurate or useful enough. In some cases, the identification task is deemed to have been completed if no further information related to the scene is received or no further dialogue occurs, e.g., within a period of time. In some embodiments, the third information is received after the second output is obtained and the third information is performed based at least in part on the identification result and / or the updated identification result. In some embodiments, the representation comprises a visual representation of the scene or a semantic representation of the scene. For example, the visual representation of the scene may include one or more images. The image(s) may include 2D or 3D image(s). The image(s) may include RGB image(s). The image(s) may include floorplan(s) or map(s) such as sematic map(s), in 2D or 3D. For example, the visual representation of the scene may include one or more point clouds. The point cloud(s) may be 2D or 3D point cloud(s). For example, the semantic representation of the scene may include one or more graphs. The graph(s) may be scene graph(s). In some embodiments, the first information related to the scene comprises a first dialogue related to the scene. For example, the first dialogue may include a languagebased dialogue. The language-based dialogue may include text, speech, etc. For example, the first dialogue may include a vision-based dialogue. The vision-based dialogue may include pictogram, sign language, body language, etc. In some embodiments, the second information related to the scene comprises a second dialogue related to the scene. For example, the second dialogue may include a language-based dialogue. The language-based dialogue may include text, speech, etc. For example, the second dialogue may include a vision-based dialogue. The visionbased dialogue may include pictogram, sign language, body language, etc. In some embodiments, the third information related to the scene comprises a third dialogue related to the scene. For example, the third dialogue may include a languagebased dialogue. The language-based dialogue may include text, speech, etc. For example, the third dialogue may include a vision-based dialogue. The vision-based dialogue may include pictogram, sign language, body language, etc. In some embodiments, the fourth information related to the scene comprises a fourth dialogue related to the scene. For example, the fourth dialogue may include a languagebased dialogue. The language-based dialogue may include text, speech, etc. For example, the fourth dialogue may include a vision-based dialogue. The vision-based dialogue may include pictogram, sign language, body language, etc. In some embodiments, performing the identification comprises identifying one or more objects in the scene. For example, identifying one or more objects in the scene includes identifying presence or absence of an object in the scene and / or identifying one or more properties of an object present in the scene. The one or more properties of the object may include, e.g., a pose, a location, or an orientation of the object. In some embodiments, performing the identification comprises determining a location of a subject in the scene. The scene may relate to an indoor environment and / or an outdoor environment. The subject may include a human subject, an animal, a robot, etc. In some embodiments, each of the information (the first information and the second information) is provided at least in part by the subject. In some embodiments, each of the information includes a respective dialogue, and each of the respective dialogue is performed in part by the subject. In some embodiments, each of the information includes a respective dialogue. In some embodiments, each of the dialogues (e.g., the first and second dialogues) is a humanmachine dialogue (i.e., performed by at least one human subject and at least one computer). In some embodiments, each of the dialogues (e.g., the first and second dialogues) is a human-human dialogue (i.e., performed by two or more human subjects). In some embodiments, the computer implemented method is performed using a machine-learning-based identification model, which includes: a first encoder for extracting features from the representation of the scene, a second encoder for extracting features from the first information, and a third encoder for fusing at least the features extracted from the representation of the scene and the features extracted from the first information to obtain the first output associated with the identification result. In some embodiments, the machine-learning-based identification model further includes a first decoder for processing the first output to obtain the representation of the identification result. In some embodiments, the machine-learning-based identification model further includes a fourth encoder for extracting features of the second information, and a fifth encoder for fusing at least the first output and the extracted features of the second information to obtain the second output associated with the updated identification result. In some embodiments, the machine-learning-based identification model further includes a second decoder for processing the second output to obtain the representation of the updated identification result. In some embodiments, the machine-learning-based identification model further includes a sixth encoder for extracting features of the third information, and a seventh encoder for fusing at least the second output and the extracted features of the third information to obtain the third output associated with the further updated identification result, and a third decoder for processing the third output to obtain the representation of the further updated identification result. In some examples, the machine-learning-based identification model may include further encoders and / or decoders. In some embodiments, the machine-learning-based identification model has been trained using a first training dataset comprising a plurality of sets of data, each of the set of data respectively includes: (i) a plurality of dialogues for performing a corresponding identification operation, (ii) corresponding representation of scenes for use in performing the corresponding identification operation, and (iii) identification result of the corresponding identification operation. Each of the dialogues corresponds to a respective information and is applied to a respective encoder of the machine-learning-based identification model. In some embodiments, the machine-learning-based identification model has been trained using a second training dataset comprising a plurality of sets of data, each of the set of data respectively includes: (i) a plurality of augmented dialogues for performing a corresponding identification operation, (ii) corresponding representation of scenes for use in performing the corresponding identification operation, and (iii) identification result of the corresponding identification operation. The plurality of augmented dialogues (in the second training dataset) are obtained by processing plurality of dialogues (in the first training dataset) using a language processing model. The language processing model may include a large language model (LLM), such as Generative pre-trained transformers (GPT). Each corresponding pair of dialogue and augmented dialogue are semantically consistent. Each of the augmented dialogues is applied to a respective encoder of the machine-learning-based identification model. In some embodiments, the machine-learning-based identification model has been trained using at least some of the sets of data in the first training dataset and at least some of the sets of data in the second training dataset. By training the machine-learning-based identification model with the first training dataset and the second training dataset (including the original dialogues and the augmented dialogues), the performance of the machine-learning-based identification model can be further generalised or improved. In a second aspect, there is provided a system comprising one or more processors configured to perform the computer-implemented method of the first aspect. In a third aspect, there is provided a carrier medium carrying computer readable instructions adapted to cause one or more processors to perform the computer-implemented method of the first aspect. In a fourth aspect, there is provided a non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors. The one or more programs include instructions for performing the computer-implemented method of the first aspect. In a fifth aspect, there is provided a method for training a machine-learning-based identification model. The machine-learning-based identification model may be the machine-learning-based identification model in the first aspect. The method includes training the machine-learning-based identification model using the first training dataset and / or the second training dataset in the first aspect. In some embodiments, the method further comprises: processing the dialogues in the first training dataset using a language processing model to obtain the augmented dialogues. The language processing model may include a large language model (LLM), such as Generative pre-trained transformers (GPT). In a sixth aspect, there is provided a system comprising one or more processors configured to perform the method of the fifth aspect. In a seventh aspect, there is provided a carrier medium carrying computer readable instructions adapted to cause one or more processors to perform the method of the fifth aspect. In an eighth aspect, there is provided a non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors, the one or more programs including instructions for performing the method of the fifth aspect. In the above aspects, the first / second / third / etc. information may together define a more complete set of information. For example, the first / second / third / etc. information may include first / second / third / etc. dialogue, which together define a more complete dialogue. Each of the first / second / third / etc. dialogue may be referred to as a turn of a dialogue (i.e., a turn of the more complete dialogue). Other features and aspects of embodiments of the invention will become apparent by consideration of the detailed description and accompanying drawings. Any feature(s) described herein in relation to one aspect or embodiment may be combined with any other feature(s) described herein in relation to another aspect or embodiment as appropriate and applicable. Terms of degree such as “generally”, “about”, “substantially”, or the like, are used, depending on context, to account for one or more of the following: manufacture tolerance, degradation, trend, tendency, imperfect practical condition(s), etc. BRIEF DESCRIPTION OF THE DRAWINGS Embodiments of the invention will now be described, by way of example, with reference to the accompanying drawings in which: Fig. 1 is a schematic diagram illustrating an operation for performing identification in some embodiments of the invention; Fig. 2 is a schematic diagram illustrating an operation for performing identification in some embodiments of the invention; Fig. 3 is a flow diagram illustrating the operation of Fig. 2; Fig. 4 is a schematic diagram illustrating an operation for performing identification in some embodiments of the invention; Fig. 5 is a schematic diagram illustrating an iterative embodied dialogue localisation operation in some embodiments of the invention; Fig. 6A is a schematic diagram illustrating architecture of a machine-learning-based localisation model in one embodiment of the invention; Fig. 6B is a schematic diagram illustrating architecture of the fusion encoder in the machine-learning-based localisation model of Fig. 6A; Fig. 7 is a schematic diagram illustrating architecture of a machine-learning-based localisation model in one embodiment of the invention; Fig. 8A is a diagram showing a ground truth dialogue and a corresponding augmented dialogue augmented using a large language model in one example; Fig. 8B is a diagram showing a ground truth dialogue and a corresponding augmented dialogue augmented using a large language model in one example; Fig. 9 is a schematic diagram illustrating architecture of a machine-learning-based localisation model in one embodiment of the invention; Fig. 10 is a graph showing cumulative match characteristic (CMC) curves for methods in some embodiments and an existing method in the task of localisation from embodied dialogue (LED), for both single-shot and multi-shot modes; Fig. 11A shows a series of diagrams illustrating the qualitative performance of methods in some embodiments and an existing method in the task of localisation from embodied dialogue (LED), for both single-shot and multi-shot modes; Fig. 11B shows a series of diagrams illustrating the qualitative performance of the methods in some embodiments and an existing model in the task of localisation from embodied dialogue (LED), for both single-shot and multi-shot modes; Fig. 12A is a graph showing performance of the LingUNet-ms method in some embodiments evaluated based on mean localisation error and using the valSeen set; Fig. 12B is a graph showing performance of the LingUNet-ms method in some embodiments evaluated based on mean localisation error and using the valllnseen set; Fig. 12C is a graph showing performance of the DiaLoc-e method in some embodiments evaluated based on mean localisation error and using the valSeen set; Fig. 12D is a graph showing performance of the DiaLoc-e method in some embodiments evaluated based on mean localisation error and using the valUnseen set; Fig. 13A is a graph showing performance of the LingUNet-ms method in some embodiments evaluated based on average estimation confidence (at ground truth location) and using the valSeen set; Fig. 13B is a graph showing performance of the LingllNet-ms method in some embodiments evaluated based on average estimation confidence (at ground truth location) and using the valllnseen set; Fig. 13C is a graph showing performance of the DiaLoc-e method in some embodiments evaluated based on average estimation confidence (at ground truth location) and using the valSeen set; Fig. 13D is a graph showing performance of the DiaLoc-e method in some embodiments evaluated based on average estimation confidence (at ground truth location) and using the valllnseen set; Fig. 14A shows a series of diagrams illustrating the qualitative performance of methods in some embodiments and an existing method in the task of localisation from embodied dialogue (LED), for both single-shot and multi-shot modes; Fig. 14B shows a series of diagrams illustrating the qualitative performance of the methods in some embodiments and an existing method in the task of localisation from embodied dialogue (LED), for both single-shot and multi-shot modes; Fig. 14C shows a series of diagrams illustrating the qualitative performance of the methods in some embodiments and an existing model in the task of localisation from embodied dialogue (LED), for both single-shot and multi-shot modes; Fig. 14D shows a series of diagrams illustrating the qualitative performance of the methods in some embodiments and an existing model in the task of localisation from embodied dialogue (LED), for both single-shot and multi-shot modes; and Fig. 15 is a schematic diagram of an information handling system in some embodiments of the invention. DETAILED DESCRIPTION Fig. 1 illustrates an operation 100 for performing identification in some embodiments of the invention. The operation 100 is a computer-implemented operation performed using a machine-learning-based identification model 100M. The machine-learning-based identification model 100M is arranged to: receive representation of a scene and information related to the scene for performing an identification task using the representation of the scene, process the representation of the scene and the received information related to the scene, and provide an estimated identification result for the identification task. The representation of the scene may be a visual representation of the scene or a semantic representation of the scene. A visual representation of the scene may include image data that represents one or more images (e.g., one or more maps) or point cloud data that represents one or more point clouds. A semantic representation of the scene may include one or more graphs (e.g., one or more scene graphs). The information related to the scene may include a turn of dialogue, which is itself a dialogue and includes at least one query and at least one corresponding response relevant for performing the identification task. The dialogue may include a language-based dialogue or a visionbased dialogue. A language-based dialogue may be represented by text, speech, etc. A vision-based dialogue may be represented by pictogram, sign language, body language, etc. The estimated identification result may include one or more identified parts of the representation of the scene or more generally the scene. The estimated identification result may be used for determining whether the identification task has been or can be completed or for facilitating receiving or processing of further information related to the scene. In an embodiment in which the information related to the scene include a dialogue or a dialogue turn, the estimated identification result may be used for facilitating continuation of dialogue (performing a next turn of the dialogue) for performing the identification task. The dialogue for performing the identification task may be performed between a machine (computer) and the system arranged to perform the computer-implemented operation. The machine (computer) may be one that can perform natural language processing without requiring human natural language input. Alternatively, the machine (computer) may be one that processes human natural language input. In some applications, the identification task is to identify one or more objects in the scene, e.g., to identify presence or absence of an object in the scene and / or to identify one or more properties (e.g., e.g., a pose, a location, or an orientation) of an object present in the scene. In some applications, the identification task is a localisation task, e.g., to determine a location of a subject in the scene. In one example, the identification task is a localisation task to determine, based on the information related to the scene (e.g., the turn of the dialogue), the location of a human subject in a physical environment. In this example, the representation of the scene includes image data representing a floor map of a physical environment and the identification task is to determine a location of a human subject in the physical environment. For the information related to the scene (e.g., the turn of the dialogue), the query may be generated by a computer or a human whereas the response to the query may be provided by the human subject in the environment. The estimated identification result provided by the machine-learning-based identification model 100M relates to an estimated location of the human subject in the environment. In another example, the identification task is a visual grounding task. In this example, the representation of the scene includes image data representing a set of images and the identification task is to determine, based on the information related to the scene (e.g., the turn of the dialogue), one or more relevant images or one or more relevant image parts. For the information related to the scene (e.g., the turn of the dialogue), the query may be generated by a computer or a human whereas the response to the query may be provided by a human subject. The estimated identification result provided by the machine-learning-based identification model 100M relates to estimated relevant image(s) or image part(s) of the set of images. Figs. 2 and 3 illustrate an operation 200 for performing identification in some embodiments of the invention. The operation 200 is a computer-implemented operation performed using a machine-learning-based identification model 200M. Referring to Fig. 2, the machine-learning-based identification model 200M is arranged to: receive a scene representation and information 1,2......N related to the scene (e.g., different turns 1, 2, , N of a dialogue) for performing an identification task using the scene representation, process them, and provide multiple estimated identification results for the identification task, each for a respective information (e.g., a respective turn of the dialogue). The scene representation may be a visual representation (e.g., image data that represents one or more images such as one or more maps or point cloud data that represents one or more point clouds) or a semantic representation (e.g., one or more graphs such as one or more scene graphs). Each of the information 1,2.....N may be a respective turn of a dialogue, and each dialogue turn is a dialogue including one or more queries and one or more responses relevant for performing the identification task. The dialogue may include a language-based dialogue or a vision-based dialogue. A language-based dialogue may be represented by text, speech, etc. A vision-based dialogue may be represented by pictogram, sign language, body language, etc. Each of the estimated identification results may include one or more identified parts of the representation of the scene or more generally the scene. Each of the estimated identification results may respectively be used for determining whether the identification task has been or can be completed or for facilitating receiving or processing of further information related to the scene. In an embodiment in which each of the information related to the scene include a dialogue or a dialogue turn, each of the estimated identification results may respectively be used for facilitating continuation of the dialogue (i.e., a next turn of the dialogue) for performing the identification task. In some examples, the machine-learning-based identification model 200M receive the information 1,2, ..., N (e.g., the different turns of the dialogue) at different times, and generate the different estimated identification results at different times. A subsequent estimated identification result (e.g., estimated identification result N, N is an integer) can be considered as an update to the previous estimated identification result (e.g., estimated identification result N-1). Referring now to Fig. 3, the operation 200 includes, in step 202, processing at least the representation of the scene and the information 1 related to the scene using the machinelearning-based identification model 200M to obtain output 1 associated with estimated identification result 1. The operation 200 also includes, in step 204, processing at least the output 1 and the information 2 related to the scene using the machine-learning-based identification model 200M to obtain output 2 associated with estimated identification result 2. The operation 200 also includes, in step 206, processing at least the output 2 and the information 3 related to the scene using the machine-learning-based identification model 200M to obtain output 3 associated with estimated identification result 3. The operation 200 may continue by repeating the processing of the latest output N-1 (N is an integer) and the data representing the information N related to the scene using the machine-learning-based identification model 200M to obtain output N associated with estimated identification result N. In this operation 200, each information related to the scene is respectively received and processed as the processing is performed. For example, information 1 of the scene is received by the machine-learning-based identification model 200M before step 202; information 2 of the scene is received by the machine-learning-based identification model 200M after step 202 and before step 204; information 3 of the scene is received by the machine-learning-based identification model 200M after step 204 and before step 206; and information N of the scene is received by the machine-learning-based identification model 200M after step 206 and before step 208. In some examples of operation 200, each information related to the scene includes a dialogue turn (a dialogue), and the dialogue is performed as the processing is performed. For example, turn 1 of the dialogue is performed before step 202; turn 2 of the dialogue is performed based on the output 1 or the estimated identification result 1 hence is performed after step 202 and before step 204; turn 3 of the dialogue is performed based on the output 2 or the estimated identification result 2 hence is performed after step 204 and before step 206; and turn N of the dialogue is performed based on the output N-1 or the estimated identification result N-1 hence is performed after step 206 and before step 208. In some examples of operation 200, each output or estimated identification result may be used to guide the performing of the next turn of the dialogue hence to update the identification result. As a result, the dialogue can be performed more efficiently to facilitate performing of the identification task, to obtain a more accurate output or estimated identification result. In some cases, an estimated identification result may be or may include a subset of a previous estimated identification result. However, depending on applications, this may not always be the case. Fig. 4 illustrates an operation 400 for performing identification (e.g., localisation) based in some embodiments of the invention. The operation 400 is a computer-implemented operation performed using a machine-learning-based localisation model such as the machine-learning-based localisation model M. The operation 400 may be applied in the operation 200 illustrated in Fig. 3 in some embodiments. For ease of presentation, the following description of the operation 400 refers to the operation 200 illustrated in Fig. 3. Referring to the operation 400 illustrated in Fig. 4, in 402, scene representation features are extracted from the representation of the scene (e.g., visual or semantic) using the machine-learning-based identification model. The extraction of scene representation features can reduce the dimension of the representation of the scene to facilitate subsequent processing. The extraction of scene I representation features may be performed by a unimodal encoder. In 404, features of turn 1 of a dialogue (language based or vision-based) are extracted using the machine-learning-based identification model. The extraction of features of turn 1 of the dialogue can reduce the dimension of the data representing turn 1 of the dialogue to facilitate subsequent processing. The extraction of features may be performed by a unimodal encoder. In 406, at least the extracted scene representation features and the extracted features of turn 1 of the dialogue are fused using the machine-learning-based identification model to obtain output 1 associated with estimated identification result 1. The fusion may be performed by a multimodal fusion encoder. In 408, the output 1 is processed using the machine-learning-based identification model to obtain a representation of the estimated identification result 1. The output may be processed by a decoder. The representation may be a machine interpretable representation or a human interpretable representation (e.g., visual representation). The representation of the estimated identification result 1 may be interpreted to facilitate continuation of the dialogue for performing the identification task, including facilitating performing of the next turn of the dialogue. In this example, 402 to 408 may be considered as an example of step 202 in operation 200. After the representation of the estimated identification result 1 is obtained and turn 2 of the dialogue is performed, in 410, features of turn 2 of the dialogue are extracted using the machine-learning-based identification model. The extraction of features can reduce the dimension of the data representing turn 2 of the dialogue to facilitate subsequent processing. The extraction of features may be performed by a unimodal encoder. In 412, at least output 1 and the extracted features of turn 2 of the dialogue are fused using the machine-learning-based identification model to obtain output 2 associated with estimated identification result 2. The fusion may be performed by a multimodal fusion encoder. In 414, the output 2 is processed using the machine-learning-based identification model to obtain a representation of the estimated identification result 2. The output may be processed by a decoder. The representation may be a machine interpretable representation or a human interpretable representation (e.g., visual representation). The representation of the estimated identification result 2 may be interpreted to facilitate continuation of the dialogue for performing the identification task, including facilitating performing of the next turn of the dialogue. In this example, 410 to 414 may be considered as an example of step 204 in operation 200. After the representation of the estimated identification result 2 is obtained and turn 3 of the dialogue is performed, 416 to 420 may be performed for the next turn (turn 3) of the dialogue to obtain output 3 and representation of the estimated identification result 3. In other words: in 416, features of turn 3 of the dialogue are extracted using the machine-learning-based identification model; in 418, at least output 2 and the extracted features of turn 3 of the dialogue are fused using the machine-learning-based identification model to obtain output 3 associated with estimated identification result 3; in 420, the output 3 is processed using the machine-learning-based identification model to obtain a representation of the estimated identification result 3. In this example, 416 to 420 may be considered as an example of step 206 in operation 200. The operation 400 may continue for one or more further turns of the dialogue, e.g., until the identification task is completed or deemed complete. The identification, or the identification task, in the above embodiments is defined broadly. For example, the identification or the identification task is not limited to localisation. In one example, the identification task is a localisation task to determine, based on the turn of the dialogue, the location of a human subject in a physical environment. For example, the scene representation includes image data representing a floor map of a physical environment and the localisation task is to determine a location of a human subject in the physical environment. For each respective turn of the dialogue (for performing the localisation task), the query may be generated by a computer ora human whereas the response to the query may be provided by the human subject or another subject in the environment. Each respective estimated identification result provided by the machine-learning-based identification model relates to an estimated location of the human subject in the environment. In another example, the identification task is a visual grounding task. For example, the scene representation includes image data representing a set of images and the identification task is to determine one or more relevant images or one or more relevant image parts (e.g., to identify presence / absence of object in the images, or to identify one or more properties of the object present in the images). For each respective turn of the dialogue (for performing the identification task), the query may be generated by a computer or a human whereas the response to the query may be provided by a human subject. The estimated identification result provided by the machine-learning-based identification model relates to estimated relevant image(s) or image part(s) of the set of images. It should be appreciated that in other embodiments the invention can be applied to other types of applications or scenarios. The following disclosure provides some more-specific embodiments of the invention. In some embodiments of the invention, the identification is iterative embodied dialogue localisation, which can be performed based on a machine-learning-based localisation model and by two agents cooperating or interacting with each other: • an observer (e.g., human or machine) located in an environment • a locator (e.g., human or machine) with knowledge of the environment (e.g., a top-down map of the environment) and can communicate with the observer through dialogue to facilitate determination of the location (e.g., the relative location) of the observer in the environment. In operation, through the dialogue, the locator may ask the observer question(s) about the environment (e.g., presence or absence of features in the environment, characteristics of the environment) and the observer may provide answer(s) based on what the observer can see in the environment. The question(s) and answer(s) may be processed by a machine-learning-based localisation model to determine an estimated localisation result (the estimated location(s) or the relative location(s) of the observer in the environment). The locator may ask further question(s) based on the observer’s answer and the observer may provide further answer(s) to the further question(s), to continue the dialogue. The further question(s) and answer(s) may be processed by the machine-learning-based localisation model to determine a further estimated localisation result (the estimated location(s) such as relative location(s) of the observer in the environment). The process may repeat until the locator has successfully determined the location of the observer in the environment. As a skilled person would appreciate, the successful performing of iterative embodied dialogue localisation would require, among other things, targeted inquiries and useful (e.g., unequivocal) linguistic responses. Fig. 5 illustrates an iterative embodied dialogue localisation operation 500 in some embodiments of the invention. The operation 500 is performed by the locator L and the observer O, using a machine-learning-based localisation model. In these embodiments, the locator L is provided with a top-down map M of the environment in which the observer O is located whereas the observer O has an egocentric view V of the environment. The locator L and the observer O are engaged in a cooperative dialogue to allow the machine-learning-based localisation model to determine the location of the observer O in the environment. In this example, in the first turn of the dialogue, the locator L asks the observer O a question about the observer’s surroundings and the observer O provides an answer based on the observer’s view. The machine-learning-based localisation model processes this first turn of the dialogue to provide 3 estimated (possible) locations of the observer O in the environment. In the second turn of the dialogue, the locator L asks the observer O a further question to disambiguate the estimated locations and the observer O provides a further answer based on the observer’s view. The machine-learning-based localisation model processes this second turn of the dialogue to provide an estimated location of the observer O in the environment. As illustrated in Fig. 5, the machine-learning-based localisation model can provide estimated location(s) of the observer O based on new dialogue information that becomes available (e.g., when a new turn of dialogue is completed). Also, the question posed by the locator L in the dialogue can be based on or affected by the preceding turn(s) of the dialogue and the previous estimation(s). In some embodiments, the machine-learning-based localisation model for the embodied dialogue localisation uses a global map e (a top-down view of the environment) and a dialogue D represented by T turns between the utterance of the locator L and the observer 0, where D = (L1; O1;..., LT, OT), Lt represents the utterance of the locator L at turn i, Ot represents the utterance of the observer 0 at turn i, i = 1,2,3.....The objective of the machine-learning-based localisation model is to estimate the observer’s final location pT in the environment represented by the global map e. In these embodiments, an iterative approach is utilized by using the machine-learning-based localisation model such that multiple estimations plr..., pT, one for each turn of the dialogue, are obtained. Each estimation may include one or more estimated locations. Such iterative approach may enhance communication efficacy and may facilitate understanding between the locator (with a top-down view of the environment) and the observer (with an egocentric view of the environment) through a series of information exchanges between them. In some embodiments, the mean localisation error (LE) of the machine-learning-based localisation model is evaluated based on geodesic distance, which, compared with Euclidean distance, is more meaningful for determining error across rooms in indoor environment. In some embodiments, pixel coordinate pT of the location is snapped to the nearest way-point node gT before calculation of mean localisation error (LE) as: LE = \\gT - ^T||2, where gT\s the ground-truth and gT is the estimation. In some embodiments, the WAY dataset, disclosed in Hahn et al. “Where are you? localisation from embodied dialogue” (2020), is used to train and evaluate the machinelearning-based localisation model. In these embodiments, the WAY dataset is split into a training set (containing 9955 dialogues related to 58 environments), a “valSeen” set (containing 305 dialogues related to 58 environments), and a “valllnseen” set (containing 579 dialogues related to 11 environments). The “valSeen” set is different from the “valUnseen” set in that the “valSeen” set contains new or additional dialogues describing locations of the subjects in the environments. The machine-learning-based localisation model can be trained using the training set and then evaluated using the valSeen and valUnseen sets. In this example, it is assumed that the environment is known to the locator during training and inference, and the inference goal is to estimate the observer’s location in the environment. In some embodiments, the iterative approach may offer improved adaptability and accuracy. For example, the iterative approach may allow the agents (locator and observer) progressively refine their spatial understanding through multiple turns of dialogue, hence accommodating dynamic environments representation and nuanced contextual information. For example, the iterative approach may allow for more effective training, may allow for better generalisation to unseen environment, and / or may be more useful for practical embodied applications. The machine-learning-based localisation model in some embodiments of the invention does not use the entire full dialogue (with multiple dialogue turns) as conditional input. The machine-learning-based localisation model in some embodiments performs multiturn estimation (i.e., with estimation occurring at the end of each turn), which more closely resembles the way human normally performs dialogue-based localisation in real-life, and may be particular suitable for applications such as search and rescue. The machine-learning-based localisation model in some embodiments of the invention is based on transformers (e.g., uses transformer based encoders and / or decoders). The machine-leaming-based localisation model in some embodiments of the invention includes an image backbone that interacts with respective dialogue turn input iteratively, to provide location estimations at multiple timesteps corresponding to the dialogue turns. Fig. 6A shows a machine-learning-based localisation model 600 in one embodiment of the invention. In this embodiment, the machine-learning-based localisation model 600 has a multi-shot multimodal architecture for iterative embodied dialogue localisation. As used herein, the machine-learning-based localisation model 600 is referred to as “DiaLoc-i”. The machine-learning-based localisation model 600 includes an image encoder 602 for processing an image of a map of an environment to provide unimodal visual (map) embeddings, a text encoder 604 for processing the first turn of the dialogue to provide unimodal text (linguistic) embeddings, and a fusion encoder 606 for integrating the unimodal visual (map) embeddings and the unimodal text (linguistic) embeddings to provide an output of hidden states, which relates to the estimation result. The machinelearning-based localisation model 600 also includes a heatmap decoder 608 for decoding the output of hidden states, to provide a heatmap containing the estimated localisation result. The machine-learning-based localisation model 600 also includes another text encoder 610 for processing the second turn of the dialogue to provide unimodal text (linguistic) embeddings, and another fusion encoder 612 for integrating the hidden states output by the fusion encoder 606 and the unimodal text (linguistic) embeddings provided by the text encoder 610, to provide another output of hidden states, which relates to another estimation result. The machine-learning-based localisation model 600 also includes another heatmap decoder 614 for decoding the output of hidden states provided by the fusion encoder 612, to provide a heatmap containing the further estimated localisation result. In this embodiment, each heatmap generally corresponds to the map of an environment but with certain areas presented differently from the others (e.g., different colour, different contrast, different intensity, etc.) to indicate the corresponding estimated localisation result (of one or more estimated locations). Fig. 7 shows another machine-learning-based localisation model in one embodiment of the invention. In this embodiment, the machine-learning-based localisation model 700 has a multi-shot multimodal architecture for iterative embodied dialogue localisation. The machine-learning-based localisation model 700 is generally the same as the machine learning-based localisation model 600, except that the unimodal visual (map) embeddings of the map is explicitly applied for each fusion operation (i.e., to each fusion encoder) and each respective hidden states are fused with the unimodal visual (map) embeddings of the map directly at each timestep. As used herein, the machine-learning-based localisation model 700 is referred to as “DiaLoc-e”. The machine-learning-based localisation model 700 includes an image encoder 702 for processing an image of a map of an environment to provide unimodal visual (map) embeddings, a text encoder 704 for processing the first turn of the dialogue to provide unimodal text (linguistic) embeddings, and a fusion encoder 706 for integrating the unimodal visual embeddings, the unimodal text (linguistic) embeddings, and initial hidden states to provide an output of hidden states, which relates to the estimation result. The machine-learning-based localisation model 700 also includes a heatmap decoder 708 for decoding the output of hidden states, to provide a heatmap containing the estimated localisation result. The machine-learning-based localisation model 700 also includes another text encoder 710 for processing the second turn of the dialogue to provide unimodal text (linguistic) embeddings, and another fusion encoder 712 for integrating the unimodal visual (map) embeddings, the hidden states output by the fusion encoder 706, and the unimodal text (linguistic) embeddings provided by the text encoder 710 to provide another output of hidden states, which relates to another estimation result. The machine-learning-based localisation model 700 also includes another heatmap decoder 714 for decoding the output of hidden states provided by the fusion encoder 712, to provide a heatmap containing the further estimated localisation result. In this embodiment, each heatmap generally corresponds to the map of an environment but with certain areas presented differently from the others (e.g., different colour, different contrast, different intensity, etc.) to indicate the corresponding estimated localisation result (of one or more estimated locations). In the machine-learning-based localisation models 600, 700 in Figs. 6A and 7, the image encoders 602, 702 and the text encoders 604, 610, 704, 710 are unimodal encoders constructed based on transformers. Specifically, in these embodiments, vision transformers (ViT) pre-trained on ImageNet are used as the image encoders 602, 702 for the global map e. The ViT is arranged to generate the visual (map) embedding V e Rm’c, where M = 196 is the number of visual tokens. In these embodiments, the global map e is resized to 224 * 224 before being processed by the image encoder 602, 702. Further, in these embodiments, pre-trained bidirectional encoder representations from transformers (BERT) are applied as the text encoders 604, 610, 704, 710 for encoding the respective turn of the dialogue D = LT, OT). The dialogue D is represented as features L e RN’C, where N =100 is the max token length. In these examples, the feature dimension C for both ViT and BERT is 768. In the machine-learning-based localisation models 600, 700 in Figs. 6A and 7, the fusion encoders 606, 612, 706, 712 are multimodal encoders O constructed based on transformers. Specifically, in these embodiments, each transformer-based multimodal encoder 0 includes a stack of transformer blocks for processing the visual (map) embeddings V provided by the image encoder and text embeddings L provided by the corresponding image encoders. Fig. 6B illustrates the construction of the multimodal encoder O in these embodiments. As shown in Fig. 6B, the multimodal encoder O includes 12 modules connected in series, with each module comprising self-attention (SA), cross-attention (CA), and feed-forward (FF) layers. The cross-attention layer is for multimodal fusion. The output of multimodal encoder is hidden states S e RM£ and it has the same dimension as visual embedding V. The multimodal encoder fuses dialogue input and visual map representations to output the current hidden states St at dialogue turn t. In the machine-learning-based localisation models 600, 700 in Figs. 6A to 7, the heatmap decoders 608, 614, 708, 714 are convolutional neural network based decoders. Each of the heatmap decoders 608, 614, 708, 714 includes a prediction head. The prediction head has multiple convolution and deconvolution layers, and is trained to produce location heatmap H e R11^ using S as input. In these embodiments, in training stage, the heatmap H is up-sampled to the same size as target H e r^o for loss caculation. In inference stage, the estimated location in image space can be obtained via p = Argmax(Softmax( / / )). The machine-learning-based localisation models 600, 700 represent two fusion variants for multi-shot localisation. The main difference between the two models 600, 700 is on how the previous hidden states St_x are integrated with new dialogue turn input Lt. The machine-learning-based localisation model 600, “DiaLoc-i”, is the implicit variant as the hidden states are iteratively updated via St = This ensures that the final states ST are conditioned on all dialogue turns (L1; The machine-learning-based localisation model 700, “DiaLoc-e”, is the explicit variant, in which the visual (map) embedding V is accessible and explicitly used at each respective timestep to fuse with the respective hidden states S for processing by the corresponding fusion encoder via: St = ¢( V 0 In this embodiment, the initial hidden states S (the left-most one in Fig. 7) is initialised to 1 and serves as the prior for the visual map. Turning now to the loss function of the machine-learning-based localisation models 600, 700. For single-shot mode, the entire dialogue is treated as a single dialogue turn (i.e., no additional dialogue turn and related processing) and a single identification is performed. In this case, the operation of the machine-learning-based localisation models 600, 700 in the single-shot mode are generally the same. For single-shot mode, given the estimated heatmap H and the target (ground-truth) heatmap H, the model 600, 700 can be trained to minimise the Kullback-Leibler (KL) divergence between the estimated location and the ground-truth location. In one example, Gaussian smoothing with a standard deviation of 3m is applied to the ground-truth heatmap H. The single-shot loss function Lss can be defined as: LSS{H, H) = log( / / ) (log(H) - log(Softmax(H))) (1) For multi-shot mode of the models 600, 700, the above single-shot loss is applied to each of the estimations H = H[=1 made at each of the turns t e [1, T]. The multi-shot loss function Lms is the weighted sum of these single-shot losses: = ^t=rLss(Ht^aT-^ (2) where a e [0,1] is the decay factor. The use of a non-zero decay factor a results in less penalty being applied to the earlier estimation(s) (e.g., first or first few estimations), which may be based on less complete dialogue context than the later estimations. In some cases, even if the earlier estimation(s) Ht where t <T are applied with the decay factor, the risk of over-fitting still exists. For example, it may not be possible to find the true location (accurate estimation) if one or more of the dialogue turns are ambiguous. Yet, the KL-divergence loss with the ground-truth target may encourage peaky estimation and may lead to over-confident and inaccurate estimation(s) at early timesteps. Thus, in these embodiments, to alleviate this issue, an auxiliary loss Laux is applied to promote the diversity of the earlier estimation(s) and avoid improper premature filtering of false positives. The auxiliary loss function Laux is defined as: ^aux(h,h) t=l ^mse (Sigmoid(Ht) O Hmask, H) (3) where Hmask is the binary mask of the target where H >0, the sigmoid function is applied to Ht to produce the probability scores of the estimated location, and mean squared loss Lmse is used to measure the estimation error. Based on the above, the final loss L in these embodiments of the models 600, 700 is based on the multi-shot loss Lms and the auxiliary loss Laux: = Lms(H,H)+pLaux(H,H) (4) where / 3 is a weighting factor. In these embodiments of the models 600, 700 (except for the ablation studies), (3 is set to 1.0. It should be noted that in some examples, in the training of the model 600, 700, the target heatmap for the heatmap decoders 608, 614 may be the same (if the target to be identified across dialogue turns remains unchanged) or different (if the target to be identified across dialogue turns changes); the target heatmap for the heatmap decoders 708, 714 may be the same (if the target to be identified across dialogue turns remains unchanged) or different (if the target to be identified across dialogue turns changes). Turning now to the training of the machine-learning-based localisation models 600, 700. As mentioned, a training set (containing 9955 dialogues) split from the WAY dataset is used to train the machine-learning-based localisation model 600, 700. In the training of these embodiments, both the input top-down maps and the corresponding ground-truth target heatmaps are resized to 224 x 224. Further, colour jittering, random cropping with ratio [0, 9, 1.0] and scale [0, 75, 1.0], and random rotation of 180° are applied to both the input top-down maps and the corresponding ground-truth target heatmaps for data augmentation. In these embodiments, the unimodal image encoders are ViT-based encoders pre-trained on lmageNet-21k whereas the unimodal text encoders are BERT-based-uncased encoders. The BERT-based-uncased encoders are frozen during training (i.e., weights of unimodal text encoders are kept unchanged). The machine-learning-based localisation models 600, 700 are trained for up to 30 epochs, with batch size 16, using AdamW as the optimiser and setting the learning rate to 2e5. In these embodiments, the best checkpoint for the trained model can be selected according to the accuracy at 5 meters (Acc@5m) on the “val Unseen” set. Experiments are performed to study the performance of the machine-learning-based localisation models 600, 700. Specifically, ablation experiments are performed to investigate the effect of the depth of the multimodal encoders, the difference of the two fusion variants (model 600 vs model 700), the effect of the decay factor a, and the impact of the auxiliary loss. Further experiments are performed to investigate the usefulness of dialogue augmentation in training the machine-learning-based localisation model 600, 700. Further, the operations of the machine-learning-based localisation models 600, 700 are compared with existing method and its variant, for both single-shot mode and multishot mode. The multi-shot mode performance has been examined and analysed in detail. In one ablation experiment, the effect of the depth (i.e., number of blocks) of the multimodal encoder is investigated using the DiaLoc-i model 600 in Fig. 6A with decay a = 0. Table 1 shows the results of this experiment (LE denotes mean localisation error). As shown in Table 1, in this example, for the valUnseen set, when the image encoder (ViT) is frozen (i.e., weights of image encoder are kept unchanged), the best performance is obtained when depth = 3 (i.e., number of blocks N in Fig. 6B = 3). When the image encoder is fine-tuned (not frozen), the performance generally decreases as the depth increases. In this example, the best result for the valUnseen set is obtained when depth = 1, with Acc@5m = 47.09. From the experiment results, it can be determined that fine-tuning the ViT can improve the performance as it may allow the ViT to better adapt to the top-down map visual input. Table 1 - Experimental results of ablation on depth of the multimodal encoders Image encoder (ViT) Depth N valSeen set valUnseen set LEi Acc@5m'|' le; Acc@5mf 1 9.29 47.18 10.54 33.27 Frozen 3 9.45 43.43 9.26 38.23 6 9.21 47.50 9.54 37.72 1 7.62 57.37 9.09 47.09 Fine-tuned 3 7.73 58.01 9.35 42.06 6 8.44 55.12 10.02 39.61 In one ablation experiment, the effect of the two fusion variants (i.e., the two models 600 and 700) and the decay factor a are investigated. As described with reference to Figs. 6A and 7, the model 600 and the model 700 represent two variants for performing multimodal fusion. Specifically, the DiaLoc-i model 600 updates the visual hidden states using dialogue input recursively (visual embeddings only explicitly fused once); the DiaLoc-e model 700 fuses the original visual (map) embedding with the hidden states by performing multiple fusion operations (one at each dialogue turn, hence visual embeddings explicitly fused at each turn). In this experiment, for both models 600, 700, the depth of the multimodal encoder is 3 (i.e., number of blocks N in Fig. 6B = 3), the image encoder ViT is fine-tuned (not frozen), without using auxiliary loss Laux {p = 0). Table 2 shows the results of this experiment (LE denotes mean localisation error). As shown in Table 2, for the valSeen set, the implicit fusion performed using the DiaLoc-i model 600 performs better than the explicit fusion performed using the DiaLoc-e model 700; for the valUnseen set, the explicit fusion performed using the DiaLoc-e model 700 performs better than the implicit fusion performed using the DiaLoc-i model 600. It is believed that the implicit fusion performed using the DiaLoc-i model 600, which leverages shared hidden states, may have a higher risk of over-fitting than the DiaLoc-e model 700. On the other hand, the explicit fusion performed using the DiaLoc-e model 700 uses fixed visual (map) embedding and learnable hidden states, may show better generalisation (i.e., is more suitable for unseen environments). In this experiment, it is found that the use of a non-zero decay factor a may lead to lower estimation accuracy due to the relatively early convergence. In other words, in this experiment, if the decay factor a is not used (or a zero decay factor is used) the performance is better. Table 2 - Experimental results of ablation on fusion schemes and the decay factor Method Decay factor a valSeen set valUnseen set LEI Acc@5m'|' LE! Acc@5mf 0.0 7.73 58.01 9.35 42.06 DiaLoc-i 0.5 7.95 54.68 9.74 36.48 1.0 8.20 55.31 9.75 33.61 0.0 8.99 50.32 9.10 41.60 DiaLoc-e 0.5 9.51 54.68 9.76 41.10 1.0 8.91 51.56 9.21 37.95 In one ablation experiment, the impact of the auxiliary loss Laux (value of p defined in equation (3)) is investigated using the DiaLoc-e model 700 in Fig. 7. In this experiment, the depth of the multimodal encoder is 1, and different decay factors a are used to train different versions of the model 700. For each decay factor a, the same model is trained with auxiliary loss (p = 1) and without auxiliary loss (p = 0). Table 3 shows the results of this experiment (LE denotes mean localisation error). As shown in Table 3, the use of the auxiliary loss Laux (in addition to the multi-shot loss Lms) in the loss function can help to improve the performance of the model 700. Table 3 - Experimental results of ablation on auxiliary loss Decay factor a Aux loss factor P valSeen set valUnseen set LEI Acc@5mt le; Acc@5mf 0.0 0.0 7.81 57.18 9.99 37.89 0.5 7.76 58.01 9.56 40.23 0.5 1.0 7.95 54.68 9.74 36.48 0.0 6.76 63.46 9.58 40.01 1.0 0.5 8.20 55.31 9.75 33.61 1.0 7.92 58.65 8.99 42.06 In one experiment, the effect of augmenting the ground-truth (original) dialogue in the training set (taken from the WAY dataset) using a large language model (LLM) and using the augmented dialogue for training the model is investigated. In this experiment, GPT is applied to investigate the effect of dialogue augmentation on the localisation task. For each sample, the GPT API is prompted to paraphrase the ground-truth (original) dialogue. Specifically, ground truth (original) dialogues of the training set is augmented using gpt-3.5-turbo-16k as the large language model, and the argument dialogue are used for training the DiaLoc-e model 700, with depth of the multimodal encoder = 3, a = 1, and ft = 0. In this example, the prompt “paraphrase the dialog” is used. In the API call, the temperature is set to 0.6 and the top-p is set to 0.5. The GPT augmented dialogue and the ground-truth dialogue are chosen randomly for training the model 700. Figs. 8A and 8B show two examples from the training set of the WAY dataset. For each of these examples, the top-down map and the corresponding target on the left are shown on the left whereas the ground truth (GT) dialogue is shown top right and the GPT-augmented or -paraphrased dialogue is shown bottom right. In both examples, the GPT-augmented or -paraphrased dialogues are generally semantically consistent with their respective ground truth (GT) dialogues. In the example of Fig. 8B, the length of the ground truth (GT) dialogue is reduced generally without altering the meaning. Note that the GPT API in the experiment does not use map information and is purely text-based. Table 4 shows the results of this experiment (LE denotes mean localisation error). As shown in Table 4, the use of the augmented dialogue for training the model 700 may improve the performance of the model for handling both seen maps (maps that have been processed by the model before) and unseen maps (maps that have not been processed by the model before). It can thus be determined that, with the text encoders frozen, an increase in dialogue diversity for training could improve the performance of the model 700. Table 4 - Experimental results of ablation on using additional, augmented dialogues Method Dialogue valSeen set valUnseen set LEi Acc@5m'|' le; Acc@5mf DiaLoc-e GT GT +GPT 8.91 7.07 51.56 60.00 9.21 9.09 37.95 40.71 Further experiments are performed to compare the performance of the methods in some embodiments and the performance of some existing methods in the task of localisation from embodied dialogue (LED), for both the single-shot and multi-shot modes. As mentioned, in single-shot mode, an entire dialog is treated as a dialogue turn and processed and a single identification is performed. In this example, for single-shot mode, a language-conditioned pixel-to-pixel LingUNet (which processes the entire dialogue instead of dialogue turns), which is used as a baseline in Hahn et al. “Where are you? localisation from embodied dialogue” (2020), is compared with the method of the embodiment DiaLoc, which can be DiaLoc-e or DiaLoc-i (as mentioned, in single-shot mode, the DiaLoc-e model 700 and the DiaLoc-i model 600 operate in the same way). In this example, for multi-shot mode, three methods are compared. The first method uses a LingllNet-ms method, a method modified based on LingUNet used as a baseline in Hahn et al. ‘Where are you? localisation from embodied dialogue” (2020). The second method uses the DiaLoc-i model 600. The third method uses the DiaLoc-e model 700. In respect of the first method for multi-shot mode, as LingUNet is designed for singleshot dialog localisation, LingUNet is modified based on the explicit DiaLoc-e model 700 to provide multi-shot estimations. This modified LingUNet model (multi-shot adaptation of LingUNet) is referred to as LingUNet-ms. Basically, the modification is based on leveraging the hidden states of previous iteration for future estimations. Fig. 9 shows an example of the LingUNet-ms model 900. As shown in Fig. 9, the LingUNet-ms model 900 includes a ResNet-18 convolutional neural network 902, as image encoder, for processing an image of a map of an environment to provide unimodal visual (map) embeddings F1, a bidirectional long short-term memory (Bi-LSTM) network 904, as text encoder, for processing the first turn of the dialogue to provide unimodal text (linguistic) embeddings, and a LingUNet model 906, as fusion encoder, for integrating the unimodal visual (map) embeddings F1, the unimodal text (linguistic) embeddings, and initial hidden states H1 to provide an output of hidden states H1 which relates to a first estimation result. The LingllNet-ms model 900 also includes a heatmap decoder 908 (e.g., CNN-based) for decoding the output of hidden states H1’, to provide a heatmap containing the localisation result. The LingllNet-ms model 900 also includes another bidirectional long short-term memory (Bi-LSTM) network 910, as text encoder, for processing the second turn of the dialogue to provide unimodal text (linguistic) embeddings, and another LingllNet model 912, as fusion encoder, for integrating the unimodal visual (map) embeddings F1, the hidden states HT output by the LingllNet model 906, and the unimodal text (linguistic) embeddings provided by the bidirectional long short-term memory (Bi-LSTM) network 910, to provide another output of hidden states H1 ”, which relates to another estimation result. The LingUNet-ms model 900 also includes another heatmap decoder 914 (e.g., CNN-based) for decoding the output of hidden states H1” provided by the LingUNet model 912, to provide a heatmap containing the further estimated localisation result. In the model 900, at each timestep (or dialogue turn) t, the hidden states of previous timestep is fused with unimodal visual (map) embeddings F1 to integrate the dynamic prior information. In respect of the second and third methods for multi-shot mode, the DiaLoc-i model 600 and the DiaLoc-e model 700 are both configured to use depth of the multimodal encoder = 3, a = 0, and p = 1. In this experiment, the accuracy at 0 meter and the accuracy at 5 meters (Acc@0m and Acc@5m) are used as evaluation metrics for both the valSeen set and the valUnseen set. Table 5 - Experimental results of performance of the methods in some embodiments and an existing method in the task of localisation from embodied dialogue Modes Method valSeen set valUnseen set Acc@0mf Acc@5m$ Acc@0m^ Acc@5mf Single- LingUNet 19.87 59.29 6.16 33.33 shot DiaLoc 25.64 66.02 7.02 40.41 LingUNet-ms 14.47 46.15 5.31 36.30 Multishot DiaLoc-i 18.43 57.18 6.42 37.89 DiaLoc-e 18.36 60.00 8.44 47.15 Table 5 shows the results of this experiment. As shown in Table 5, for single-shot mode (i.e., the entire dialogue is used in both training and inference), the DiaLoc method substantially outperforms LingllNet. The improvement in Acc@5m on the valllnseen set is 7.08. This performance enhancement can be attributed to the learning capabilities and versatility of the transformer-based unimodal and multimodal encoders in the DiaLoc method. On the other hand, in the context of the multi-shot localisation task, the evaluation focuses on the final estimation accuracy. First, for the valSeen set, the performance of all three methods in the multi-shot mode is lower than the methods in the single-shot mode. For the valSeen set, the methods based on the DiaLoc-i model 600 and the DiaLoc-e model 700 show better performance than the method based on LingllNet-ms model 900. For the valllnseen set, the method based on the DiaLoc-e model 700 performs best among the three methods. This demonstrates the robust generalisation and reasoning capabilities of the third method based on the DiaLoc-e model 700. For the valllnseen set, the performance of first method based on LingllNet-ms model 900 also appears to be acceptable. This suggests that the iterative multi-shot mode can effectively reduce the issue of over-fitting. In terms of computational complexity, the multi-shot mode is less complex than the single-shot mode, even though the multi-shot mode requires multiple forward passes. In respect of computational costs during the dialogue embedding stage performed using the BERT-based text encoders, for a dialogue with T turns, each turn encoded by N tokens, the single-shot mode involves 0(TN2) multiplications for self-attention computation whereas the multi-shot mode involves only O(N2) multiplications. Fig. 10 shows cumulative match characteristic (CMC) curves for both DiaLoc-e and LingUNet, for single-shot and multi-shot modes, on the WAY dataset. In this example, for single-shot mode, DiaLoc-e is effectively the same as DiaLoc-i whereas LingUNet is denoted as LingUNet-ss, the baseline in Hahn et al. “Where are you? localisation from embodied dialogue” (2020), and in multi-shot mode, DiaLoc-e (different from DiaLoc-i) and LingUNet-ms discussed above are used. In Fig. 10, the x-axis denotes the error threshold for the localisation error and the y-axis denotes the success rate. From Fig. 10, it can be seen that DiaLoc-e consistently outperforms the LingUNet-based (LingUNet-ss and LingUNet-ms) baseline in both single-shot and multi-shot configurations. For the valUnseen set, the multi-shot performance of DiaLoc-e surpasses that of the single-shot baseline (LingUNet) and single-shot DiaLoc-e, hence can obtain improved performance in the processing of new environments. Figs. 11A and 11B show two example estimations performed using the method based on LingllNet (LingUNet-ss for single-shot and LingUNet-ms for multi-shot) and the method based on DiaLoc-e, under single-shot and multi-shot modes, for one case in the valSeen set (valSeen case 67) and one case in the valllnseen set (valllnseen case 245). In each of Figs. 11A and 11B, the first column shows the visual map (Map) and its corresponding ground truth (GT) location, the second column shows the single-shot estimation (LingUNet-ss and DiaLoc-e), and the last three columns show the multi-shot estimations for one, two, and three dialogue turns respectively (LingUNet-ms and DiaLoc-e). As shown in Figs. 11A and 11B, generally, under single-shot mode, the DiaLoc-e method yields more precise and concentrated estimation than the LingUNet-ss method whereas under multi-shot mode, the DiaLoc-e method shows the capability to rectify its previous estimation by leveraging the most recent dialogue information as guidance. For the valSeen case 67 in Fig. 11 A, the DiaLoc-e method effectively corrects the estimation after the second turn (2 / T) whereas the LingUNet-ms method produces relatively noisy distributions. For the valUnseen case 245 in Fig. 11B, in the single-shot mode, the DiaLoc-e method and the LingUNet method both fail to obtain an accurate estimation whereas in the multi-shot mode, the DiaLoc-e method succeeds in refining the estimation hence provides an accurate estimation whereas the LingUNet-ms method converges towards an incorrect area hence produces an inaccurate estimation. The application of multi-shot localisation in some embodiments of the invention can allow early termination of dialogue in real-world applications such as search and rescue applications. In one experiment, the performance of methods LingUNet-ms and DiaLoc-e in multi-shot mode for a dialogue up to timestep (or turn) t is compared using the valSeen set and the valUnseen set. To determine the estimated locations for the localisation task, hard argmax or soft threshholding can be used. For example, argmax can be used to identify the top-1 estimated location based on the estimated heatmap; soft thresholding can be used to identify a set of top-K estimated locations with probabilities surpassing a threshold. In both cases, evaluation metrics such as Recall and Precision can be employed for measuring performance. In this experiment, for simplicity, argmax is used to evaluate the interim estimations. To investigate performance across different dialogue length (i.e., total number of dialogue turns T, not the number of words in the dialogue), in this experiment samples from the valSeen and valllnseen sets are grouped based on their length, and for each group, the mean localisation error (LE) is determined (within each sub-plot, the LE at t where 1 <t <T is shown). It is noted that this approach may be unfair if multiple peaks (estimations) show up and one of them is true. Thus, in some cases, the estimation probability of ground truth (GT) pixel should be considered. The results are shown in Figs. 12A to 12D. As shown in Figs. 12A to 12D, for the DiaLoc-e method, as more dialogue turns are used the localisation error is generally reduced whereas for the LingUNet-ms method lacks such a trend. Figs. 12A to 12D also show that both the LingUNet-ms method and the DiaLoc-e method have reduced performance for dialogues with more number of turns (e.g., dialogues with T = 6 turns). It is believed that this reduced performance may be caused by the unbalanced training data (hence performing additional dialogue turns does not improve the performance). In one experiment, the performance of the methods LingUNet-ms and DiaLoc-e in multishot mode, for a dialogue up to timestep (or turn) t is compared using the valSeen set and the valUnseen set. Unlike the previous experiment which employs localisation error based on top-1 estimation, in this experiment, the estimation confidence is employed. As mentioned, the use of localisation error as a metric may be unfair in cases where multiple peaks (estimations) show up and one of them is true positive. To address this issue, given the heatmap estimation, the pixel-wise probability at ground-truth location is determined in this experiment. To investigate performance across different dialogue length (i.e., total number of dialogue turns T, not the number of words in the dialogue), in this experiment, samples from the valSeen and valUnseen sets are grouped based on their length, and for each group, the average estimation confidence is determined (within each sub-plot, the average estimation confidence at t where 1 <t <T is shown). The results are shown in Figs. 13A to 13D. As shown in Figs. 13A to 13D, for the DiaLoc-e method, as more dialogue turns are used the estimation confidence generally increases whereas for the LingUNet-ms method the estimation confidence generally decreases. In one experiment, the inference runtime and memory usage of the two multi-shot methods, LingUNet-ms and DiaLoc-e, are evaluated. In this experiment, the batch size is set to 1 for both methods and a NVIDIA Titan RTX 24Gb is used for benchmarking. The DiaLoc-e method is evaluated with a depth of 1 in the multi-shot mode on the valSeen set. Table 6 shows the evaluation results. Table 6 - Experimental results of average runtime and memory usage during inference for the methods in some embodiments Method Image Size (height x width) Runtime (second) GPU Usage (MiB) LingUNet-ms 455 x 780 0.768 1273 DiaLoc-e 224 x 224 1.374 3875 Figs. 14A to 14D show four further example estimations performed using the method based on LingUNet (LingllNet-ss for single-shot and LingllNet-ms for multi-shot) and the method based on DiaLoc-e, under single-shot and multi-shot modes, for two cases in the valSeen set (valSeen case 87 and 176) and two cases in the valllnseen set (valllnseen case 224 and 327). In each of Figs. 14A to 14D, the first column shows the visual map (Map) and its corresponding ground truth (GT) location, the second column shows the single-shot estimation (LingUNet-ss and DiaLoc-e), and the last three columns show the multi-shot estimations for one, two, and three dialogue turns respectively (LingUNet-ms and DiaLoc-e). As shown in Fig 14A, for the valSeen case 87: In the single-shot mode, the DiaLoc-e method estimates the correct location whereas the LingUNet-ss method fails, in the multi-shot mode, the LingUNet-ms method produces noisy but correct estimations whereas the DiaLoc-e method is capable of generating concentrated multi-modal estimations (multiple possible locations). As shown in Fig 14B, for the valSeen case 176: In the single-shot mode, the DiaLoc-e method and the LingUNet-ss method both provide reasonable estimations. In multi-shot mode, the DiaLoc-e method can recover from its initial incorrect estimation. As shown in Fig 14C, for the valUnseen case 224: In the single-shot mode, the DiaLoc-e method estimates the correct location whereas the LingUNet-ss method fails. In multishot mode, the LingUNet-ms method generates noisy estimations whereas the DiaLoc-e method can gradually refine its estimation and converge towards the correct location. As shown in Fig 14D, for the valUnseen case 327: In the single-shot mode, the DiaLoc-e method and the LingllNet-ss method both fail. In the multi-shot mode, the LingllNet-ms method converges to a few locations but none of which is correct whereas the DiaLoc-e method incrementally refines its estimation and converges towards the correct location. Fig. 15 shows an information handling system 1500 in some embodiments of the invention. The information handling system 1500 can be used to perform the methods or operations in various embodiments of the invention. For example, the information handling system 1500 may be configured to perform one or more of: the operation 100 in Fig. 1, the operation 200 in Figs. 2 and 3, the operation 400 in Fig. 4, etc. For example, the information handling system 1500 may be configured to train and / or operate one or more of: the model 600 in Fig. 6A, the model in Fig. 6B, or the model 700 in Fig. 7, the model 900 in Fig. 9, etc. The information handling system 1500 generally comprises suitable components necessary to receive, store, and execute appropriate computer instructions, commands, and / or codes. The main components of the information handling system 1500 are a processor 1502 and memory (storage) 1504. The processor 1502 may include one or more: CPU(s), MCU(s), GPll(s), logic circuit(s), Raspberry Pi chip(s), digital signal processor(s) (DSP), application-specific integrated circuit(s) (ASIC), field-programmable gate array(s) (FPGA), or any other digital or analogue circuitry / circuitries configured to interpret and / or to execute program instructions and / or to process signals and / or information and / or data. The memory 1504 may include one or more volatile memory (such as RAM, DRAM, SRAM, etc.), one or more non-volatile memory (such as ROM, PROM, EPROM, EEPROM, FRAM, MRAM, FLASH, SSD, NAND, NVDIMM, etc.), or any of their combinations. Appropriate computer instructions, commands, codes, information and / or data may be stored in the memory 1504. Computer instructions for executing or facilitating execution of the training, operation, and / or method embodiments of the invention may be stored in the memory 1504. One or more machine-learning-based localisation models may be stored in the memory 1504. Data for use in inference and / or data for use in training using the machine-learning-based localisation model may be stored in the memory 1504. The one or more processors 1502 and the memory (storage) 1504 may be integrated or separated (and operably connected). Optionally, the information handling system 1500 further includes one or more input devices 1506. Example of such input device 1506 include: keyboard, mouse, stylus, image scanner, microphone, tactile / touch input device (e.g., touch sensitive screen), image / video input device (e.g., camera), etc. Optionally, the information handling system 1500 further includes one or more output devices 1508. Example of such output device 1508 include: display (e.g., monitor, screen, projector, etc.), speaker, headphone, earphone, printer, additive manufacturing machine (e.g., 3D printer), etc. The display may include a LCD display, a LED / OLED display, or other suitable display, which may or may not be touch sensitive. The information handling system 1500 may further include one or more disk drives 1512 which may include one or more of: solid state drive, hard disk drive, optical drive, flash drive, magnetic tape drive, etc. A suitable operating system may be installed in the information handling system 1500, e.g., on the disk drive 1512 or in the memory 1504. The memory 1504 and the disk drive 1512 may be operated by the one or more processors 1502. Optionally, the information handling system 1500 also includes a communication device 1510 for establishing one or more communication links (not shown) with one or more other computing devices, such as servers, personal computers, terminals, tablets, phones, watches, loT devices, or other wireless computing devices. The communication device 1510 may include one or more of: a modem, a Network Interface Card (NIC), an integrated network interface, a NFC transceiver, a ZigBee transceiver, a Wi-Fi transceiver, a Bluetooth® transceiver, a radio frequency transceiver, a cellular (2G, 3G, 4G, 5G, above 5G, or the like) transceiver, an optical port, an infrared port, a USB connection, or other wired or wireless communication interfaces. Transceiver may be implemented by one or more devices (integrated transmitter(s) and receiver(s), separate transmitter(s) and receiver(s), etc.). The communication link(s) may be wired or wireless for communicating commands, instructions, information and / or data. In one example, the one or more processors 1502, the memory 1504 (optionally the input device(s) 1506, the output device(s) 1508, the communication device(s) 1510 and the disk drive(s) 1512, if present) are connected with each other, directly or indirectly, through a bus, a Peripheral Component Interconnect (PCI), such as PCI Express, a Universal Serial Bus (USB), an optical bus, or other like bus structure. In one embodiment, at least some of these components may be connected wirelessly, e.g., through a network, such as the Internet or a cloud computing network. A person skilled in the art understands that the information handling system 1500 is merely an example and that the information handling systems in other embodiments can have configurations different from the information handling system 1500 (e.g., include additional components, has fewer components, etc.). Although not required, one or more embodiments of the invention can be implemented as an application programming interface (API) or as a series of libraries for use by a developer or can be included within another software application, such as a terminal or computer operating system or a portable computing device operating system. As program modules include routines, programs, objects, components, and data files assisting in the performance of particular functions, the skilled person will understand that the functionality of the software application may be distributed across a number of routines, objects and / or components to achieve the same functionality desired herein. It should be appreciated that where the methods and systems of the invention are either wholly or partly implemented by computing system(s) then any appropriate computing system architecture, including stand-alone computers, network computers, dedicated or non-dedicated hardware devices, may be utilised. Terms such as “computing system”, “computing device”, or the like are intended to include (but not limited to) appropriate arrangement of computer or information processing hardware capable of implementing the function described. The methods and / or operations in embodiments of the invention may be performed at least in part by a system with at least one processor. The methods and / or operations in embodiments of the invention may be arranged as computer readable instructions (e.g., one or more programs) carried by a carrier medium. In some examples, the methods and / or operations in embodiments of the invention may be arranged as computer readable instructions stored in / by a non-transitory computer-readable medium. Some embodiments of the invention can be used to facilitate identification such as iterative embodied dialogue localisation, which requires a locator to possess the capability of finding out an observer’s location following each information (e.g., a turn of the dialogue). The methods and models in some embodiments of the invention can operate with greater efficiency as they do not utilize all information (e.g., the entire dialogue) at the same time for processing during inference and instead only processes one information (e.g., a single turn of the dialogue) at a time. In some practical applications related to localisation, this efficiency improvement can to save time, or even life, as the location of the observer can be estimated or narrowed down relatively quickly. The methods and models in some embodiments of the invention can enhance generalisation to new environment. The risk of over-fitting to training data can be mitigated by not relying on complete dialogues during training of the models. The methods and models in some embodiments of the invention can provide intermediate estimation, which may provide cues for assessing uncertainty and may be particularly useful for identification tasks such as localisation tasks that involve dialogue generation (e.g., embodied visual dialogue and cooperative localisation). The iterative dialoguebased localisation approach in some embodiments of the invention not only aligns with engineering perspectives but also paves the way for improved performance and broader applicability across different types of localisation tasks or more generally identification tasks. The methods and models in some embodiments of the invention can be used to rectify previous estimation with new information and may help a human operator or a machine in question formulation in search and rescue applications. The methods and models in some embodiments of the invention uses the dialogue in an iterative manner based on dialogue turns to progressively refine localisation estimations, with intermediate estimations are made at the end of each turn. Previous estimations could be implicitly or explicitly applied for future estimations. Also, the dialogue can be terminated any time by checking out the intermediate estimations (e.g., if the intermediate estimations can provide useful results). The methods and models in some embodiments of the invention may provide one or more of the following advantages. For example, compared to non-iterative approach, the iterative approach in some embodiments can offer improved adaptability and accuracy. In an example embodied dialogue localisation application, the iterative approach enables agents (locator and observer) to progressively refine their spatial understanding through ongoing dialogue, accommodating dynamic environments representation and nuanced contextual information. For example, the methods in some embodiments, in multi-shot mode when compared with single-shot mode, may allow for more effective training, better generalisation to unseen environments, and ultimately more practical embodied applications. It should be noted that the methods and models of the invention, depend on how and what they are compared with, may provide other advantages not specifically presented herein. The methods and models of the invention can be used in different identification applications such as but not limited to localisation applications. For example, the methods and models in some embodiments can enable interactive multimodal dialoguebased localisation, referring expression comprehension, and visual dialogue. For example, the methods and models in some embodiments can enable interactive image retrieval. Some other example applications of the methods and models in some embodiments include: Al-aided search and rescue, self-localisation assistant, assistant robot in shops and warehouse (visual grounding), maintenance assistant (visual grounding), dialogue-based image retrieval for inspection task, etc. It will be appreciated by a person skilled in the art that variations and / or modifications may be made to the described and / or illustrated embodiments of the invention to provide other embodiments of the invention, without departing from the scope of the invention as defined by the accompanying claims. The described and / or illustrated embodiments of the invention should therefore be considered in all respects as illustrative not restrictive. The steps or operations in some embodiments may be performed in a different order than illustrated, as long as they are practically applicable and logical. While the machine-learning-based localisation model or method embodiments specifically described are related to a specific number of pieces of information or turns of dialogue, the machine-learning-based localisation model or method can be performed for other number of pieces of information or turns of dialogue (one or more turns) to provide other embodiments of the invention. Each dialogue turn can itself be considered as a dialogue. Thus, a more complete dialogue can be defined by multiple dialogues (or dialogue turns). The machine-learning-based localisation model of the invention may have a different architecture than those specifically illustrated. For example, the number of encoders and the number of decoders may be different from those illustrated. For example, the type of encoders and the type of the decoders may be adapted based on the type of data they need to process. In some cases, the observer or subject may move around in the environment during the localisation task and hence the observer or subject may not be at the exact same location during the dialogue. In these cases, the dialogue would evolve as the observer or subject moves around and the dialogue turns can be processed by the localisation model. In some cases, target(s) of the identification task may be changed or updated, e.g., in real time or dynamically.
Claims
1. A computer-implemented method for performing identification, comprising: receiving a representation of a scene;receiving a first information related to the scene for performing identification on the representation of the scene;processing at least the representation of the scene and the first information to obtain a first output associated with an identification result;receiving a second information related to the scene for performing identification on the representation of the scene; andprocessing at least the first output and the second information to obtain a second output associated with an updated identification result.
2. The computer-implemented method of claim 1, wherein the processing of at least the representation of the scene and the first information comprises: extracting features from the representation of the scene;extracting features from the first information; andfusing at least the features extracted from the representation of the scene and the features extracted from the first information to obtain the first output.
3. The computer-implemented method of claim 2, wherein:the extraction of features from the representation of the scene is performed using a first encoder;the extraction of features from the first information is performed using a second encoder; andthe fusing of at least the features extracted from the representation of the scene and the features extracted from the first information is performed using a third encoder.
4. The computer-implemented method of claim 3, wherein:the first encoder is a unimodal encoder, which may be an image encoder;the second encoder is a unimodal encoder, which may be a text encoder or a speech encoder; andthe third encoder is a multimodal fusion encoder.
5. The computer-implemented method of any one of claims 1 to 4, further comprising:processing the first output to obtain a representation of the identification result; and outputting the representation of the identification result.
6. The computer-implemented method of claim 5, wherein the processing of the first output is performed using a first decoder, which may be a heatmap decoder.
7. The computer-implemented method of any one of claims 1 to 6, wherein the processing of at least the first output and the second information comprises: processing at least the representation of the scene, the first output, and the second information to obtain the second output.
8. The computer-implemented method of claim 7, wherein the processing of at least the first output and the second information comprises: extracting features from the second information; andfusing at least the first output and the features extracted from the second information to obtain the second output.
9. The computer-implemented method of claim 8, wherein the fusing of at least the first output and the features extracted from the second information comprises: fusing the features extracted from the representation of the scene, the first output, and the features extracted from the second information to obtain the second output.
10. The computer-implemented method of claim 8 or 9, wherein:the extraction of the features of the second information is performed using a fourth encoder; andthe fusing of at least the first output and the features extracted from the second information is performed using a fifth encoder.
11. The computer-implemented method of claim 10, wherein: the fourth encoder is a unimodal encoder; and the fifth encoder is a multimodal fusion encoder.
12. The computer-implemented method of any one of claims 1 to 11, further comprising: processing the second output to obtain a representation of the updated identification result; andoutputting the representation of the updated identification result.
13. The computer-implemented method of claim 12, wherein the processing of the second output is performed using a second decoder such as a heatmap decoder.
14. The computer-implemented method of any one of claims 1 to 13, wherein the second information is received after the first output is obtained, and the second information is based at least in part on the identification result.
15. The computer-implemented method of any one of claims 1 to 14, wherein performing the identification comprises performing localisation of a subject in the scene, e.g., to determine a location of the subject in the scene.
16. The computer-implemented method of claim 15, wherein:the first information is provided at least in part by the subject; and the second information is provided at least in part by the subject.
17. The computer-implemented method of any one of claims 1 to 14, wherein performing the identification comprises identifying one or more objects in the scene.
18. The computer-implemented method of claim 17, wherein the identification of one or more objects in the scene comprises:identification of presence or absence of an object in the scene; and / or identification of one or more properties of an object present in the scene, such as a pose, a location, or an orientation of the object present in the scene.
19. The computer-implemented method of any one of claims 1 to 18, wherein: the first information related to the scene comprises a first dialogue related to the scene; andthe second information related to the scene comprises a second dialogue related to the scene.
20. The computer-implemented method of claim 19, wherein:the first dialogue comprises a language-based or vision-based dialogue; and the second dialogue comprises a language-based or vision-based dialogue.
21. The computer-implemented method of any one of claims 1 to 20, wherein the representation of the scene comprises: a visual representation of the scene; or a semantic representation of the scene.
22. The computer-implemented method of claim 21, wherein the visual representation of the scene comprises:one or more images such as one or more maps; or one or more point clouds.
23. The computer-implemented method ofclaim 21, wherein the semantic representation of the scene comprises:one or more graphs such as one or more scene graphs.
24. A system comprising one or more processors configured to perform the computer-implemented method of any one of claims 1 to 23.
25. A carrier medium carrying computer readable instructions adapted to cause one or more processors to perform the computer-implemented method of any one of claims 1 to 23.50