Contextualized scene question and answer data generation method oriented to personal interaction
Through the contextualized scenario Q&A data generation method for embodied interaction, the problem that the data generation method in the prior art fails to meet the actual needs of embodied intelligence is solved. The generated data is more in line with the embodied characteristics, and the inference and generalization capabilities of the agent are enhanced.
Patent Information
- Application Number
- CN202411944935.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-16
AI Technical Summary
Existing data generation methods fail to fully consider the actual needs of embodied intelligence, especially in intensive reasoning environments and task generalization capabilities, which fail to achieve the expected results of humans.
A situational scene question and answer data generation method for embodied interaction is proposed. By obtaining target scene data, generating scene context descriptions using the description model, setting multiple scenarios and mapping the description to the first perspective of the agent, constructing an interactive question and answer acquisition system, collecting question and answer data data from real users, and optimizing data quality through keyword splitting and manual correction.
The generated data is more in line with the embodied characteristics, enhances the inference ability and generalization ability of the agent, provides a more comprehensive data generation solution, and promotes the development and application of the embodied intelligence field.
Smart Images

Figure CN120012916A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence embodied interaction technology, and in particular to a method for generating contextualized scenario question and answer data for embodied interaction. Background Art
[0002] If intelligent agents want to truly interact with the physical world, the key is to be in the real physical world and interact with humans, understand and learn the physical relationships between things in the real world and the social relationships between different intelligent agents, so as to act like humans. In recent years, some progress has been made in the research of embodied intelligence, including navigation and object manipulation in three-dimensional scenes. These tasks are inseparable from large-scale data support, and data generation methods for multimodal tasks play a decisive role in the quality and reliability of data. Although the current data has promoted the progress of various embodied perception tasks, the current data generation methods have failed to fully consider the actual needs of embodied intelligence, and the generalization ability in environments with intensive reasoning (such as home) and tasks still does not meet the expected effect of humans. In order to better meet the actual needs of intelligent agents to interact with the world, data generation methods for embodied interaction scenarios need to achieve the following goals: for a given scene context (for example, 3D scanning), understand the situation in the scene (position, direction, etc.) and describe its surrounding environment in a situational manner based on the 3D scene described by language, reasonably answer questions raised by people, and better reflect the interaction with the physical world or people.
[0003] In order to meet the above needs, most of the current data generation methods are oriented towards 3D scene understanding, including single tasks such as 3D scene positioning and description generation, which use large language models to generate task-related annotations or data descriptions. However, these methods are still limited in their application to actual scenarios of embodied intelligence. On the one hand, such methods fail to fully consider situational awareness, and the description of the scene is relatively simple, lacking understanding of the scene and logical reasoning; in addition, the observation of the data is from a third-person perspective rather than a specific self-centered perspective. On the other hand, the current data generation methods consider a single type of task and data, and are limited in data modality, diversity, scale, and task scope. Therefore, there is a lack of contextualized embodied question-and-answer data generation methods for embodied interaction to better meet the actual needs of intelligent agents for embodied understanding and reasoning. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a method for generating contextualized scenario question and answer data for embodied interaction, which can effectively generate contextualized embodied question and answer data for embodied interaction, and provide a better basis for enhancing the reasoning ability and generalization ability of the intelligent body.
[0005] The technical solution adopted by the present invention to solve the technical problem is: to provide a method for generating contextualized scenario question and answer data for embodied interaction, comprising the following steps:
[0006] Obtain target scene data;
[0007] Based on the target scene data, generating a corresponding scene context description using a description model;
[0008] A plurality of different scenarios are set based on the position and orientation of the agent in the scenario, and the scenario context description is mapped to the first perspective of the agent according to the set scenario to obtain a corresponding scenario context description;
[0009] Based on the target scenario data, an interactive question-answering collection system is constructed to collect question-answering data of real users;
[0010] The scenario context description is split into keywords, and the split keywords are used as prior prompts for real users to participate in the interactive question-and-answer process.
[0011] Furthermore, before the step of constructing an interactive question-and-answer collection system based on the target scenario data to collect question-and-answer data of real users, the step further includes:
[0012] Scenario keywords are set based on the set scenario, and the scenario context is updated after the generated scenario context is integrated and screened by the large language model with the scenario keywords.
[0013] Furthermore, the method also includes the step of manually correcting the updated scenario context.
[0014] Furthermore, based on the target scenario data, an interactive question-answering collection system is constructed to collect question-answering data of real users, including:
[0015] Building an interactive UI based on the target scene data so that real users can observe the current scene in real time;
[0016] Guide real users to ask complex reasoning questions and give questions and answers that are consistent with contextual understanding and reasoning;
[0017] Collect the generated question and answer data.
[0018] Furthermore, the complex reasoning problems include navigation problems, common sense problems, and multi-level reasoning problems.
[0019] Furthermore, the step of selecting the question and answer data that meets the contextualization requirements from the collected question and answer data is also included, specifically including:
[0020] Delete question and answer data irrelevant to the current scenario based on comparison of keywords or scenario descriptions;
[0021] Rewriting the templated descriptions and questions that do not need to be combined with scenarios in the candidate question and answer data;
[0022] Reduce the number of set-type questions.
[0023] Furthermore, the method also includes the step of using a grammar checking tool to check and correct errors in the language description.
[0024] Furthermore, the acquiring the target scene includes:
[0025] Get 3D indoor scene dataset;
[0026] According to the richness of the spatial layout and the size of the spatial volume within the scene, the target scene data that meets the set conditions is filtered out.
[0027] Beneficial Effects
[0028] Due to the adoption of the above-mentioned technical scheme, the present invention has the following advantages and positive effects compared with the prior art: the present invention combines the best contextualization to reason about the scene, proposes a contextual description strategy and an interactive question-answering collection and optimization scheme for three-dimensional scenes, covers multi-modal and multi-task embodied understanding tasks, and fully considers the specific perspective centered on the self. At the same time, by introducing a method of interactive human participation in question-answering, the authenticity and rationality of data generation are guaranteed, and data that is more in line with embodied characteristics is obtained; the present invention fully considers the generalization and reasoning ability of the intelligent body, and the task range includes everything from spatial relationship understanding to navigation, common sense reasoning, etc., which provides ideas for proposing a more comprehensive data generation scheme in the subsequent embodied intelligence field, promotes the intelligent body to achieve better embodied understanding and reasoning, and provides support for the future application of embodied intelligence. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic diagram of a flow chart of an embodiment of the present invention;
[0030] Figure 2 This is an example diagram of question and answer data according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall within the scope limited by the appended claims of the application equally.
[0032] The embodiments of the present invention relate to a method for generating contextualized scenario question-answering data for embodied interaction, such as Figure 1As shown in the figure, it includes three steps: scene description, generation of question-answer pairs, and data optimization. 3D indoor scenes are mainly selected from existing public datasets (such as ScanNet), but some scenes may be too crowded / sparse, or small as a whole, making question collection infeasible. Therefore, these scenes are first manually classified according to the richness of object / space layout and spatial volume. After deleting those scenes that do not meet the requirements, multiple scenes with a certain data scale are finally retained. Subsequently, a web-based interactive UI is built to collect question requirements from real users, describe the scene, ask some questions, and give answers that are in line with the situation. The whole process requires people to interact with the three-dimensional scene, which on the one hand ensures the diversity of data and the situational understanding of the scene, and on the other hand ensures that the generation method meets the real needs of real humans. In order to ensure the quality of the generated data, the whole process is divided into three stages of subtasks.
[0033] (1) Generation of contextual description. Different from the previous single description of the three-dimensional scene, which only simply describes the position or direction, the description of the language or scene lacks contextual understanding and awareness, and is often ambiguous. This implementation method obtains a contextual description of the scene observed from the first perspective by assuming the position and orientation of the intelligent agent in the three-dimensional scene. The specific method includes:
[0034] First, the targets and objects in the scene are identified with the help of existing scene understanding methods. For the description of the context, assuming that the input scene is a video, a video description model (ClipBERT or VQA system MCAN) is used to generate a context description of the current scene.
[0035] By assuming the position and orientation of the intelligent agent in the three-dimensional scene, different scenarios are set, and then the context description of the scene is converted to the first-person perspective of the intelligent agent through spatial mapping according to the set scenario to obtain the corresponding context description of the scenario.
[0036] In addition, in order to make the obtained description context more in line with situational needs, the LLM model or GPT is used to integrate and screen the existing information through the perception of the current context, target location and other environments, and generate a variety of description contents that meet the situation. In order to ensure the effectiveness of the results, in this process, keyword information related to the scene and situation is provided, usually keywords related to the environment or possible interactive behaviors (such as kitchen, microwave oven, food) to assist the model in making reasonable judgments, thereby making corrections and updates. This is also one of the key points that distinguishes it from previous data generation methods.
[0037] All scenario descriptions are manually processed and annotated to ensure diversity and reduce ambiguity. Manual processing is mainly due to the fact that it is difficult for models or machines to accurately give rational descriptions based on human thinking logic, so some manual corrections are made to the results that do not conform to the current situation. If necessary, multiple descriptions of the scene can be introduced to enhance data diversity.
[0038] (2) Question collection. A series of question sets are collected for scene and scenario descriptions. In order to make the collected questions involve some substantial reasoning ability, the present invention will evaluate the questions and answers raised by different people. In order to guide them to give some questions and answers that are in line with contextual understanding and reasoning, the present invention guides different users by designing a keyword prompt method. Specifically, the scenario description obtained in the above steps is first split into keywords, and these keywords are used as prompt information as a priori prompts in the process of users participating in interactive Q&A. On the premise, users are guaranteed to collect Q&A data that is more in line with the current situation as much as possible to ensure the rationality of questions and answers. In order to ensure the quality of annotation, questions or answers that are not related to the current situation will be judged and deleted by comparing with keywords or scene descriptions to avoid the collection of invalid questions. For example, "how many chairs are there in the room?" is not included because this question does not need to consider the current situation, while "how many chairs are there behind me?" is more in line with contextual needs. At the same time, humans can ask questions that require complex reasoning, such as navigation, common sense, multi-level reasoning, etc. These tasks can better reflect embodied perception and understanding capabilities. These can greatly ensure the richness of the questions in the data set, the diversity of the language, and the comprehensiveness of the data.
[0039] (3) Interactive answer collection & data optimization. In addition to the answers collected together with the questions, unlike the previous generation method of question and answer descriptions that relies on large language models, in order to ensure the authenticity and rationality of the data, the present invention introduces an interactive UI on the network to send questions to more employees and users, allowing more real humans and users to participate and record their real answers. In order to improve the collection efficiency, the present invention will use the sending function of the network UI to directly send the collection requirements to multiple users, and only allow different users to participate in a limited number of question and answer tasks, so that human performance avoids obvious bias for different question types. In addition, in order to ensure that users can provide real and reasonable answers, users can also observe the current scene in real time through interactive UI visualization, fully understand the contextual information of the scene, and then answer the same questions in the second step to ensure the correspondence between the collected questions and answers and the authenticity and diversity of the question and answer data. In the end, a large-scale, multi-modal, multi-scenario, and multi-situation intelligent question and answer data benchmark that is closer to actual needs is obtained.
[0040] Finally, in order to ensure data quality, the present invention also designs an additional method for post-processing the generated data description to improve the quality and reliability of the results. Specifically, the present invention first uses the grammar checking tool grammar to check and correct errors in the language description; in addition, it uses existing QA question-answering models (such as OK-VQA) to eliminate some meaningless descriptions and questions. For example, those descriptions that are similar to templates (for example, repeating the same sentence pattern) and those questions that are too simple or do not need to be combined with scenarios can be rewritten. Since certain types of questions may tend to favor certain answers, in order to avoid uneven data distribution, some questions are removed to ensure a more even data distribution.
[0041] The question-answer data generated in this embodiment is as follows Figure 2As shown in the figure, it mainly includes: ① 3D scene, including multiple representation methods of indoor scenes, including 3D point cloud, first-person perspective video and bird's-eye view. Dataset users can freely choose from them; ② Situational description, combining the current position (position) and direction (orientation) of a certain agent in the scene (indicated by an arrow), as well as the surrounding objects, to describe the current situation. In order to make the relevant description more in line with the situation, the agent should use daily activities that are in line with the context to describe, such as "I am waiting for my meal to be heated in the microwave" instead of "I am in front of the microwave" (I am in front of the microwave) This is a more mechanical description; ③ Question and answer collection, assuming that the agent is in each specific situation and gives reasonable and correct answers to different people's questions based on the current situation and environment.
[0042] Compared with most previous template-based text generation technologies, the three-dimensional contextualized scenario question and answer data generation method based on interactive human question and answer pairs adopted in the present invention has better versatility and application value. Through the participation of real users or different people, the generated questions and answers are more in line with human contextualized scenarios and the application requirements of embodied intelligence. The obtained data is also more diverse and balanced in comparison, avoiding excessive data bias.
[0043] Building a general question-answering system has always been one of the goals of artificial intelligence. With the advancement of multimodal machine learning, intelligent question-answering is committed to promoting the development of multimodal question-answering systems that are more in line with human contexts. The data generation method proposed in the present invention is also committed to providing data benchmarks for intelligent question-answering systems that assist embodied interaction, and promoting the development and progress of intelligent question-answering systems with rich scene embodied understanding capabilities. The method can select inputs from 3D scans, egocentric videos, or BEV pictures, which makes the proposed data generation method and the subsequently established datasets compatible with various existing intelligent question-answering systems.
Claims
1. A method for generating contextualized scenario question-answering data for embodied interaction, characterized in that: The following steps are involved: Obtain target scene data; Based on the target scene data, generating a corresponding scene context description using a description model; A plurality of different scenarios are set based on the position and orientation of the agent in the scenario, and the scenario context description is mapped to the first perspective of the agent according to the set scenario to obtain a corresponding scenario context description; Based on the target scenario data, an interactive question-answering collection system is constructed to collect question-answering data of real users; The scenario context description is split into keywords, and the split keywords are used as prior prompts for real users to participate in the interactive question-and-answer process.
2. The method according to claim 1, characterized in that Before the step of constructing an interactive question-answer collection system based on the target scenario data to collect question-answer data of real users, the method further includes: Scenario keywords are set based on the set scenario, and the scenario context is updated after the generated scenario context is integrated and screened by the large language model with the scenario keywords.
3. The method according to claim 2, characterized in that It also includes a step of manually correcting the updated situation context.
4. The method according to claim 1, characterized in that: The step of constructing an interactive question-and-answer collection system based on the target scenario data to collect question-and-answer data of real users includes: Building an interactive UI based on the target scene data so that real users can observe the current scene in real time; Guide real users to ask complex reasoning questions and give questions and answers that are consistent with contextual understanding and reasoning, while limiting the number of question-answering tasks they participate in; Collect the generated question and answer data.
5. The method according to claim 4, characterized in that The complex reasoning problems include navigation problems, common sense problems, and multi-level reasoning problems.
6. The method according to claim 1, characterized in that It also includes the steps of filtering out the question and answer data that meets the contextualization needs from the collected question and answer data, including: Delete question and answer data irrelevant to the current scenario based on comparison of keywords or scenario descriptions; Rewriting the templated descriptions and questions that do not need to be combined with scenarios in the candidate question and answer data; Reduce the number of set-type questions.
7. The method according to claim 6, characterized in that The method also includes the step of using a grammar checking tool to check and correct errors in the language description.
8. The method according to claim 1, characterized in that The obtaining of the target scene includes: Get 3D indoor scene dataset; According to the richness of the spatial layout and the size of the spatial volume within the scene, the target scene data that meets the set conditions is filtered out.