Method and device for generating a scenario question and answer model of an agent, equipment and medium
By building a scenario question-answering model, the intelligent agent can accurately and promptly answer scenario questions in real or virtual environments, solving the problem of insufficient scenario question-answering capabilities of the intelligent agent and achieving efficient intelligence improvement.
Patent Information
- Application Number
- CN202510246876.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The intelligent agents in existing technologies lack the ability to answer scenario questions, resulting in a low level of intelligence and an inability to accurately and promptly answer scenario-related questions in real or virtual environments.
By generating a scenario question-answering model, building a scenario question set based on the full scenario information, obtaining scenario recall information, generating matching scenario answers, and training the preset question-answering model, a scenario question-answering model of the intelligent agent is generated.
The intelligent agent can accurately and promptly answer questions about the scene content in real or virtual environments, improving its intelligence level, demonstrating wide applicability and flexibility, approaching the performance of top language models, and achieving efficient scene question-answering capabilities at low-cost training.
Smart Images

Figure CN119739880B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of image processing technology, and more specifically to a method, device, equipment, medium and product for generating a scene question-answering model of an intelligent agent. Background Art
[0002] Artificial Intelligence (AI) is a key driving force behind the new scientific and technological revolution and industrial transformation. It is a new, critical technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. As a key component of intelligence science, AI aims to understand the essence of intelligence and produce new intelligent machines (i.e., agents) that can respond in a manner similar to human intelligence.
[0003] Artificial General Intelligence (AGI) refers to a general artificial intelligence (AI) with efficient learning and generalization capabilities, capable of autonomously generating and completing tasks in complex and dynamic environments. It possesses autonomous perception, cognition, decision-making, learning, execution, and social collaboration capabilities, and is consistent with human emotions, ethics, and moral values. Therefore, one of the most important capabilities of AGI is the ability to provide context-specific question-and-answer services.
[0004] However, existing technologies do not yet have the ability to realize scenario-based question-answering capabilities of intelligent agents. They are usually aimed at question-answering scenarios such as chatting or knowledge graphs. Moreover, the question-answering interaction capabilities of intelligent agents in these question-answering scenarios are also greatly limited, resulting in a low level of intelligence. Summary of the Invention
[0005] In view of at least one of the above problems, the embodiments of the present invention aim to provide methods, devices, equipment, media and products for generating scenario question-answering models of intelligent agents that can realize the generation of scenario question-answering models, thereby significantly improving the timely and accurate interaction capabilities of intelligent agents for scenario questions and answers, enabling the intelligent agents to give more fluent and natural answers to scenario questions, ensuring that they are more in line with real human reactions, thereby achieving a higher level of intelligence.
[0006] One aspect of an embodiment of the present invention provides a method for generating a scenario question and answer model of an intelligent agent, which includes: generating a scenario question set related to the scenario in which the intelligent agent is located based on full scenario information; obtaining scenario recall information related to each scenario question in the scenario question set; based on the scenario recall information, generating a matching scenario answer for each scenario question in the scenario question set to form a scenario answer set; and performing training on a preset question and answer model based on the scenario question set and the matching scenario answer set to generate a scenario question and answer model of the intelligent agent.
[0007] According to one embodiment of the present invention, before generating a scenario problem set related to the scenario in which the agent is located based on the full scenario information, it also includes: obtaining the full scenario information related to the scenario in which the agent is located in response to the received model generation task instructions.
[0008] According to one embodiment of the present invention, in generating a scene problem set related to the scene in which the intelligent agent is located based on the full scene information, it includes: generating local scene information of each object with each object in the full scene information as the center; generating a local scene problem set of the local scene information of each object based on preset scene problem samples and preset problem generation rules, which is used to constitute the scene problem set.
[0009] According to one embodiment of the present invention, obtaining scene recall information related to each scene problem in the scene problem set includes: obtaining scene visual information related to each scene problem in the scene problem set; performing information recall on the full scene information based on the scene visual information to obtain scene recall information.
[0010] According to one embodiment of the present invention, based on the scenario recall information, a matching scenario answer is generated for each scenario question in the scenario question set to form a scenario answer set, including: extracting the problem scenario information in the scenario recall information that matches each scenario question in the scenario question set; and reorganizing the problem scenario information according to preset answer generation rules to generate corresponding scenario answers.
[0011] According to one embodiment of the present invention, in performing training on a preset question-answering model based on a scenario question set and a matching scenario answer set to generate a scenario question-answering model of an intelligent agent, it includes: generating a training data set that conforms to a preset data format based on the scenario question set and the matching scenario answer set; and inputting the training data set into the preset question-answering model according to preset training rules to generate a scenario question-answering model.
[0012] Another aspect of an embodiment of the present invention provides a device for generating a scenario question-answering model for an intelligent agent, comprising a question generation module, an information acquisition module, an answer generation module, and a model training module. The question generation module is configured to generate a scenario question set related to the scenario in which the intelligent agent is located based on full scenario information; the information acquisition module is configured to obtain scenario recall information related to each scenario question in the scenario question set; the answer generation module is configured to generate a matching scenario answer for each scenario question in the scenario question set based on the scenario recall information, thereby forming a scenario answer set; and the model training module is configured to train a preset question-answering model based on the scenario question set and the matching scenario answer set to generate a scenario question-answering model for the intelligent agent.
[0013] Another aspect of an embodiment of the present invention provides an electronic device comprising one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the method for generating the scenario question-answering model of the above-mentioned intelligent agent.
[0014] Another aspect of an embodiment of the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method for generating the scenario question-answering model of the above-mentioned intelligent agent.
[0015] Another aspect of an embodiment of the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the method for generating the scenario question-answering model of the above-mentioned intelligent agent.
[0016] The method for generating a scenario question-answering model for an intelligent agent provided by an embodiment of the present invention can at least partially solve the problem in related arts that the intelligent agent cannot well implement scenario question-answering, resulting in a low level of intelligence, and thus can achieve at least one of the following technical effects:
[0017] Compared to existing intelligent agents that can only achieve chat and knowledge graph-based question-answering capabilities, the scenario-based question-answering model provided by the embodiments of the present invention enables intelligent agents to accurately and promptly answer questions about the context of the scene, based on the context in which they are located. This allows intelligent agents to achieve the scenario-based question-answering capabilities required of general artificial intelligence. Secondly, this scenario-based question-answering model enables intelligent agents to understand and answer questions about contextual information in real-world or virtual environments, including identifying and describing relevant information about visual objects and their relative relationships. This demonstrates broad applicability and flexibility, facilitating the development of cross-domain applications for intelligent agents. Furthermore, by constructing the aforementioned low-cost training dataset and fine-tuning the preset question-answering model, up to 90% of the capabilities of the GPT-40 model can be achieved with limited resources. This means that without relying on large-scale data and computing resources, the performance approaches that of current state-of-the-art language models. Furthermore, based on the generation of this scenario-based question-answering model, intelligent agents can provide more fluent and natural responses to contextual questions, ensuring a closer fit with real human responses, thereby achieving a higher level of intelligence.
[0018] Therefore, the embodiments of the present invention provide an overall framework based on advanced training data construction and model fine-tuning strategies, and provide an efficient and scalable solution, thereby providing a new solution for improving scenario problem capabilities. Practical verification is sufficient to show that this technical solution has demonstrated extremely high practicality and innovation in promoting the development of general artificial intelligence, and has important commercial application value and scientific research value.
[0019] It should be understood that the foregoing general description and the following detailed description are merely exemplary and illustrative and are not intended to limit the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0021] Figure 1 A diagram schematically illustrates an application scenario of a method, apparatus, device, medium, and program product for generating a scenario question-answering model for an intelligent agent according to an embodiment of the present invention;
[0022] Figure 2 A flowchart schematically illustrates a method for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention;
[0023] Figure 3 A flowchart schematically illustrates an application scenario of training data construction for a method for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention;
[0024] Figure 4 A flowchart of an application scenario of model training generation of a method for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention is schematically shown;
[0025] Figure 5 A block diagram schematically illustrates a structure of a device for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention; and
[0026] Figure 6 A block diagram of an electronic device suitable for implementing a method for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention is schematically shown.
[0027] The above-mentioned drawings are part of the description of the embodiments of the present invention, illustrating exemplary embodiments of the present invention. Together with the description, the drawings are used to illustrate the principles of the embodiments of the present invention. It should be understood that the above general description of the drawings and the following detailed description are merely exemplary and illustrative and are not intended to limit the scope of the present invention. DETAILED DESCRIPTION
[0028] In order to make the objectives, technical solutions and advantages of the embodiments of the present invention more clearly understood, the spirit of the contents disclosed in the present invention will be clearly illustrated with the accompanying drawings and detailed descriptions below. After understanding the embodiments of the contents of the present invention, any technician in the relevant technical field can change and modify the contents of the present invention based on the techniques taught by the contents of the present invention without departing from the spirit and scope of the contents of the present invention.
[0029] The exemplary embodiments of the present invention and their description are used to explain the present invention, but are not intended to limit the present invention. In addition, elements / components with the same or similar reference numerals used in the drawings and embodiments are used to represent the same or similar parts.
[0030] The terms “first,” “second,” etc. used in the present invention do not particularly refer to an order or sequence, nor are they used to limit the present invention. They are only used to distinguish elements or operations described with the same technical terms.
[0031] The directional terms used in the present invention, such as up, down, left, right, front, or back, are only used with reference to the directions in the accompanying drawings. Therefore, the directional terms used are used to illustrate and not to limit the present invention.
[0032] The terms “include,” “including,” “have,” “contain,” etc. used in the present invention are open-ended terms, meaning including but not limited to.
[0033] The term "and / or" used in the present invention includes any or all combinations of the items mentioned.
[0034] Regarding the present invention, "plurality" includes "two" and "more than two"; regarding the present invention, "plurality of groups" includes "two groups" and "more than two groups".
[0035] The terms "substantially" and "approximately" used in this disclosure are intended to modify any quantity or error that may vary slightly, but such variations or errors do not alter the essence of the quantity. Generally speaking, the range of such variations or errors may be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments, or other values. Those skilled in the art will appreciate that the aforementioned values may be adjusted based on actual needs and are not intended to be limiting.
[0036] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0037] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to systems having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, and C, etc.). When expressions such as "at least one of A, B, or C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but is not limited to systems having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, and C, etc.). Those skilled in the art should also understand that any transitional conjunctions and / or phrases indicating two or more optional items, whether in the specification, claims, or drawings, should be understood to provide the possibility of including one, either, or both of these items. For example, the phrase "A or B" should be understood to include the possibilities of "A" or "B", or "A and B".
[0038] There are existing solutions in the art that can obtain dialogue / question-and-answer training datasets based on large language models (LLMs). For example, patent CN117033608B provides a knowledge graph generative question-and-answer method based on a large language model. It first generates a training set based on a large language model with stronger capabilities, then fine-tunes the weaker / open-source base model through various fine-tuning methods. Finally, using data distillation methods, the originally weaker base model can achieve a level similar to that of the stronger large language model in certain aspects.
[0039] However, existing technologies lack training solutions for large language models in contextual question-answering scenarios. These scenarios primarily focus on casual conversation or domain knowledge graph-based question-answering. Furthermore, these scenarios often suffer from a lack of intelligence, resulting in a poor user experience and an inability to achieve timely and accurate question-answering. Furthermore, the required language models often require costly data training and fine-tuning, limiting the application and development of this technology in intelligent agents.
[0040] In view of at least one of the above problems, the embodiments of the present invention are intended to enable the generation of a scenario question-answering model, thereby significantly improving the timely and accurate interaction capabilities of intelligent agents for scenario questions and answers. The method, apparatus, equipment, medium, and product for generating a scenario question-answering model for intelligent agents can enable the intelligent agents to give more fluent and natural answers to scenario questions, ensuring that they are more consistent with real human reactions, thereby achieving a higher level of intelligence. To this end, the embodiments of the present invention, for the first time, propose a solution for how to construct a scenario question-answering model with scenario question-answering capabilities, design a specific prompt template, and demonstrate a comparison of the effects of the final fine-tuning model.
[0041] One aspect of an embodiment of the present invention provides a method for generating a scenario question and answer model of an intelligent agent, which includes: generating a scenario question set related to the scenario in which the intelligent agent is located based on full scenario information; obtaining scenario recall information related to each scenario question in the scenario question set; based on the scenario recall information, generating a matching scenario answer for each scenario question in the scenario question set to form a scenario answer set; and performing training on a preset question and answer model based on the scenario question set and the matching scenario answer set to generate a scenario question and answer model of the intelligent agent.
[0042] Figure 1 The application scenario diagram schematically shows the method, apparatus, device, medium and program product for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention.
[0043] like Figure 1As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0044] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0045] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0046] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.
[0047] It should be noted that the method for generating the scenario question-answering model of the intelligent agent provided in the embodiment of the present invention can generally be executed by the server 105. Accordingly, the device for generating the scenario question-answering model of the intelligent agent provided in the embodiment of the present invention can generally be set in the server 105. The method for generating the scenario question-answering model of the intelligent agent provided in the embodiment of the present invention can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the device for generating the scenario question-answering model of the intelligent agent provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0048] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0049] The following will be based on Figure 1 The scene described by Figures 2 to 4The method for generating the scenario question-answering model of the intelligent agent of the disclosed embodiment is described in detail.
[0050] Figure 2 The flowchart of the method for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention is schematically shown.
[0051] like Figure 2 As shown, one aspect of an embodiment of the present invention provides a method for generating a scenario question-answering model of an intelligent agent, which includes operations S201 to S204.
[0052] In operation S201, a scene problem set related to the scene in which the agent is located is generated based on the full scene information;
[0053] In operation S202, scenario recall information related to each scenario question in the scenario question set is obtained; in operation S203, a matching scenario answer is generated for each scenario question in the scenario question set based on the scenario recall information to form a scenario answer set; and
[0054] In operation S204, the preset question-answering model is trained according to the scenario question set and the matching scenario answer set to generate a scenario question-answering model of the intelligent agent.
[0055] The intelligent agent can be the subject of the execution of the method for generating the scenario question-answering model described above in the embodiments of the present invention, or can be an executor controlled by the method. Specifically, it can be a humanoid intelligent robot or other AI device, which typically has its own actuators capable of completing specific motion tasks. For example, a humanoid robot can use a mechanical manipulator to complete the motion task of picking up an object.
[0056] In embodiments of the present invention, a scene can be the real world in which an agent resides, such as the user's bedroom or living room. Alternatively, the scene can be a virtual world that the agent "imagines" or "exists in," such as the agent embedding itself in a game environment or constructing a virtual world based on received environmental information. Objects and people in these real or virtual worlds can constitute the content of the scene.
[0057] Furthermore, scenario-based question-answering (Q&A) involves users asking questions based on the context of the scenario. The agent can then observe the scene information and respond appropriately to the user's questions. Specifically, users can inquire about the type, number, movement, and position of visual objects in the scene, as well as the relative relationships between multiple visual objects. The agent can then provide targeted responses based on the acquired scene content. Scenario-based Q&A is a must-have capability for general artificial intelligence, requiring the agent to be able to respond to questions about the context of the scene it is in.
[0058] The full scene information is the information about all objects and relationships between objects in the environment obtained by the intelligent agent based on the detection of the real environment or virtual environment, and also includes the intelligent agent's own state information.
[0059] Among them, all object information includes information such as objects, people, number of objects, types of objects, size of objects, functions of objects, spatial positions, etc. in the environment; information on relationships between objects can include spatial relationships between objects in the environment (such as a cup on a table) and interactions between objects and people (such as a person standing in front of a table with a cup and a kettle pouring water), and can also include interactions between people (such as a conversation between person A and task B), etc.; the agent's own state information can be the agent's own action execution state information detected at the current moment, such as the agent's walking action in the motion state, action information of the hand-waving action (such as wrist torque, angle of rotation, expected execution time, and spatial positions of key points of fingers, etc.) or static state information (such as spatial position, current static posture, etc.).
[0060] The scenario question set is a data set of questions related to the scenario content generated by the agent based on the above-mentioned full scenario information. These questions can reflect the agent's cognition and understanding of the current scenario.
[0061] Each scenario question in the scenario question set can reflect partial information about the scenario corresponding to the question. This information related to the scenario question can serve as the aforementioned scenario recall information. The scenario recall information can specifically be the scenario information related to the scenario question obtained by extracting the full scenario information based on the scenario question, or it can be partial scenario information.
[0062] The scenario answer is the response content that matches the scenario question after extracting the scenario recall information based on the scenario question. The scenario answer set is the data set composed of these scenario answers.
[0063] Each scenario question in the scenario question set can have a matching scenario answer in the scenario answer set, and the constituent matching relationship between these scenario questions and scenario answers can form training data based on the scenario question set and the scenario answer set. The preset question-answering model can be an ordinary large language model (such as the Qwen model) provided in an embodiment of the present invention, which is used as a base model for generating the scenario question-answering model and does not need to rely on large-scale data and a large amount of computing resources. Therefore, by inputting the training data of the above-mentioned scenario question set and scenario answer set into the preset question-answering model, the generation of the scenario question-answering model of the embodiment of the present invention can be realized.
[0064] Therefore, compared to the existing technology in which the intelligent agent can only realize the question-answering capabilities of small talk and knowledge graph, the scenario question-answering model provided by the embodiment of the present invention can enable the intelligent agent to accurately and promptly answer questions about the scene content according to the scene in which it is located, so that the intelligent agent can achieve the scenario question-answering capabilities that general artificial intelligence should have. Secondly, based on this scenario question-answering model, the intelligent agent can be given the ability to understand and answer questions about scene information in a real-world environment or a virtual world environment, including identifying and describing relevant information about visual objects and their relative relationship information, showing wide applicability and flexibility, and helping to realize the cross-domain application development of the intelligent agent. In addition, only through the construction of the above-mentioned low-cost training data set, combined with the fine-tuning of the preset question-answering model, without relying on large-scale data and computing resources, it has achieved performance close to that of the current top language model. Moreover, based on the generation of the above-mentioned scenario question-answering model, the intelligent agent can give more fluent and natural answers to scene questions, ensuring that they are more consistent with real human reactions, thereby achieving a higher level of intelligence.
[0065] In order to enable those skilled in the art to have a clearer understanding of the method for generating the scenario question-answering model of the above-mentioned intelligent agent in the embodiment of the present invention, the following is further provided: Figure 3-Figure 4 Description.
[0066] Figure 3 A flowchart of an application scenario of a method for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention is schematically shown.
[0067] In an embodiment of the present invention, the method for generating the above-mentioned scenario question-answering model may include two parts: construction of training data and training of a base model.
[0068] In the training data construction in the embodiment of the present invention, candidate questions (i.e., queries) can be generated for a given scenario first, and then based on the given scenario and the candidate questions, relevant information in the scenario is recalled first, and then a matching answer (i.e., answer) is generated.
[0069] Figure 3 A flowchart of an application scenario of a method for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention is schematically shown.
[0070] like Figure 2-Figure 3 As shown, according to one embodiment of the present invention, before operation S201 generates a scene problem set related to the scene where the agent is located according to the full scene information, it also includes:
[0071] In response to the received model generation task instruction, full scene information related to the scene in which the agent is located is obtained.
[0072] The model generation task instruction may be a task execution instruction parameter generated in response to a user input action requiring model generation, and may include parameter content for controlling the intelligent agent to execute the generation of the aforementioned scenario question-and-answer model, so that after receiving the instruction, the intelligent agent can perform environmental detection on the scene in which it is located to obtain the aforementioned full-scene information. In addition, the model generation task instruction may also be a task execution instruction parameter autonomously generated by the intelligent agent based on the current environment or the interaction scenario between the intelligent agent and the user, etc. After generating the instruction, the intelligent agent can respond to the instruction and perform environmental detection on the scene in which it is located to obtain the aforementioned full-scene information.
[0073] This enables the intelligent agent to detect the environment more intelligently and efficiently and obtain the required full-scene information, such as Figure 3 Operation S301 is shown.
[0074] like Figure 2-Figure 3 As shown, according to one embodiment of the present invention, in operation S201, a scene problem set related to the scene where the agent is located is generated based on the full scene information, including:
[0075] Generate local scene information of each object with each object in the full scene information as the center;
[0076] According to the preset scene problem samples and preset problem generation rules, a local scene problem set of the local scene information of each object is generated to form a scene problem set.
[0077] The local scene information is the local scene data obtained by disassembling and extracting the local scene content of the full scene information, wherein the local scene data is related to the corresponding object. Specifically, each object mentioned in the full scene information can be confirmed, and then the corresponding local scene content can be extracted with each object as the center. For example, with the table as the center, all objects related to the table and the information on the relationship between objects in the front, back, top, bottom, left and right of the space where the table is located are extracted as the local scene content of the table; and then with another object other than the table as the center, the information on the objects and the relationship between objects in the space around the object is obtained and extracted as another corresponding local scene content. These local scene contents can constitute the local scene information corresponding to the object.
[0078] The preset scenario question samples are pre-set scenario description information related to scenario questions and answers and question information designed based on the scenario description information. They can be understood as example data of "scene description (Scene) - generated question (Query)".
[0079] For example, sample 1 of the preset scenario question could be: The input scenario description is "You are lying on the floor, with an open door and a red pen to your right." The corresponding output generation questions could include "What is to your right? What is on your right? Is there a door to your right? Is there a pen to your right? What is to the right of your right? Is the door to your right open or closed? What color are the pens on your right? How many pens are there on your right? Where are the pens? Where is the door? Is the door open? How many pens are there? Is there a red pen in the scene? Is there anything the same color as a tomato? Where is the door?" etc.
[0080] Sample 2 of the preset scenario question might be: The input scenario description is "There is a door on your right, on the floor. To your right are a plant, two blue toys, and a cupboard." The corresponding output questions might include "Where is the door? Where is the door? What is to the right of the door? How many plants are to the right of the door? How many toys are to the right of the door? What color are the toys? Which side of the door are the toys? Where is the cupboard by the door? How many toys are there?" etc.
[0081] Sample 3 of the preset scenario question might be: The input scenario description is "A storage box is in front of you, on the floor. To its right are a baby room and a blue cylindrical coffee table." The corresponding generated questions might include "What color is the coffee table? What shape is the coffee table? What's on the coffee table? Where is the locker? Where is the storage box? Is there anything the same color as a tomato? What's to the right of the storage box? What's to the right of the storage box? Where is the baby room? What's in the baby room?" and so on.
[0082] Therefore, the scene descriptions input in the preset scene question samples can include information such as the name, color, and location of objects in the detected environment, as well as information about other objects in front, behind, and around the corresponding objects. Therefore, by imitating the preset scene question samples, the intelligent agent can generate matching scene questions for each input object's corresponding local scene information based on preset question generation rules, thereby forming a scene question set. The preset scene question samples can be preset using a few-shot learning approach.
[0083] The preset question generation rules may be preset rules for generating matching scenario questions based on the scenario description information. In an embodiment of the present invention, the preset question generation rules may include question generation requirements for the question object, questioning method, question subject, question scope, and rule combination. Among them, these question generation requirements may include the following:
[0084] (1) Question object: The question object can be the name, color, shape, number, size, relative position, absolute position, function, etc. of the objects in the local scene.
[0085] (2) Questioning methods: There are at least two types of questioning methods. One is what / who / how, which is a sentence pattern of what; the other is is / has, which is a sentence pattern of “xx?”
[0086] (3) Question subject 1: The subject of a question can have a multi-hop relationship. For example, “What shape is the water cup?” is a 1-hop question, “What can the object on the right side of the water cup be used for?” is a 2-hop question, and “What is behind the object on the right side of the water cup?” is a 3-hop question. Each question has a mutual association relationship.
[0087] (4) Question subject 2: The subject of a question can be one, two or more. For example, the question “Which is bigger, the water cup or the mouse on the table?” has two subjects, while the question “What shape is the water cup on the table?” has one subject.
[0088] (5) Question scope: Questions need to be asked about all items in the input scene description, and the generated questions must conform to Chinese expressions.
[0089] (6) Rule combination: For the generated questions, you can arbitrarily combine the rule requirements of (1) to (5) above to generate questions with different sentence structures as much as possible and cover all sentence structures as much as possible.
[0090] Therefore, according to the above-mentioned preset question generation rules, the scenario question-answering model can have more complex question-asking capabilities, which is more in line with the interactive habits of human scenario question-answering, making each scenario question in the generated scenario question set more natural and in line with the scenario needs.
[0091] Furthermore, the preset question generation model is fine-tuned according to the preset question generation rules, and the preset scene question sample is input into the preset question generation model as a reference sample for scene question generation, so that a local scene question set corresponding to the local scene information of each object can be generated. These local scene question sets may include multiple questions (queries) generated by the scene description of the local scene information of a single object, such as Figure 3 Operation S302 is shown. The above-mentioned preset question generation model can be implemented using a large language model such as GPT-4o.
[0092] Therefore, the preset question generation model can be used to traverse each local scene in the full scene, thereby obtaining a set of local scene questions corresponding to all local scenes in the full scene. These local scene question sets can serve as the components of the scene question set. Furthermore, the scene question-answering model provided by the embodiments of the present invention can enable the intelligent agent to identify and describe the relevant information of visual objects and their relative relationships based on the scene it is in, and accurately and timely provide questionable questions about the scene content.
[0093] like Figure 2-Figure 3 As shown, according to one embodiment of the present invention, in operation S202, the scene recall information related to each scene problem in the scene problem set is obtained, including:
[0094] Obtain scene visual information related to each scene question in the scene question set;
[0095] Information recall is performed on the full scene information according to the scene visual information to obtain scene recall information.
[0096] Scene visual information can be visual information within the scene that is related to the corresponding scene problem and described in natural language, including object description information related to the scene problem, such as the color, shape, position of the object and the natural language description content of nearby objects.
[0097] The scene recall information is the information related to the corresponding scene problem obtained by recalling the information related to the scene visual information in the full scene information based on the above scene visual information. It can cover more scene information related to the objects described by the scene visual information and the relationship between objects, such as Figure 3 Operation S303 is shown.
[0098] For example, if a scene question (Query) in the corresponding scene question set is "Which is greater, the number of sofas or the number of green plants in the scene?", then the corresponding scene visual information may include "the color, shape, quantity, position of the sofa and related visual information of nearby objects", "the color, shape, quantity, position of the green plants and related visual information of nearby objects". Accordingly, the scene recall information obtained by performing recall in the full scene information based on the above scene visual information and scene question may include "the item information in the room now includes 2 sofas. sofa_1, on your left, is located in the living room. Its color is white and its volume is 3373529. In front of it are 1 green plant, 1 green storage box, 1 green closed trash can, 1 hanging painting, 1 paper ball, and 3 yellow building blocks. On its right are 3 green plants, 1 green open door, 1 green closed trash can, and 1 hanging painting. On its left are 2 green plants, and behind it is 1 hanging painting. Its surroundings are divided into three groups: There is 1 pink clothing, 1 white sofa, 1 yellow bookshelf, 1 yellow carpet, sofa_0, on your left, located on the ground, its color is white, its volume is 1768717, there is 1 pink pillow inside it, there are 3 green plants, 1 green open door, 1 green closed trash can on its right, there are 3 green plants, 1 green storage box, 1 paper ball on its left, there is 1 green closed trash can, 2 paintings, 1 yellow painting, 1 mirror in front of it, there are 1 living room, 1 white sofa, 1 dining room around it, there is 1 painting behind it, etc.
[0099] Therefore, by summarizing the scene visual information related to the scene question with the help of the above-mentioned scene recall information, the finally generated scene answer can be guaranteed to match the scene question, while also providing a more complex scene answer that conforms to human answering habits.
[0100] like Figure 2-Figure 3 As shown, according to an embodiment of the present invention, in operation S203, based on the scenario recall information, a matching scenario answer is generated for each scenario question in the scenario question set to form a scenario answer set, including:
[0101] Extracting the problem scenario information matching each scenario question in the scenario question set from the scenario recall information;
[0102] The question scenario information is reorganized according to the preset answer generation rules to generate the corresponding scenario answer.
[0103] like Figure 3Operation S304 is shown. In an embodiment of the present invention, for each generated scenario question, a corresponding scenario answer (Answer) is matched based on the scenario recall information corresponding to the scenario question.
[0104] Specifically, the question scene information is descriptive information directly related to the question content of a certain scenario question, extracted from the matched scenario recall information. For example, if the question content of the scenario question involves location and objects (such as, "What color is the potted plant on the ground?"), the question scene information may include the ground, the potted plant, and the location and color of the potted plant.
[0105] Specifically, the process of extracting scene information for this problem can be understood as the agent's thinking and statistical process. Based on the scene question to be answered and the natural language description of the scene recall information corresponding to the scene question, the required information is extracted relying on some preset extraction rules. The preset extraction rules can limit the requirements or scope of information extraction, such as "If an object is in front of, behind, left, or right of a table, then the object will not be on the table" and "The concept of shape is not a type; it should be a cuboid, cylinder, etc." This can reflect the "thinking" process (i.e., Prompt_Thought) in the process of generating the answer to the scene.
[0106] The preset answer generation rules may be preset rules for generating scenario answers that match the scenario questions based on the above-mentioned question scenario information. In an embodiment of the present invention, the preset answer generation rules may include answer generation requirements such as answering person, answering method, answer restrictions, etc. These answer generation requirements may include the following:
[0107] (1) Answering person: Questions should be answered in the first person, using a child's tone.
[0108] (2) Answering method: The answer should be brief, fluent, and natural, and should conform to the logic of normal human speech and expression.
[0109] (3) Answer restriction 1: The answer should be generated based on the thought process of the input, without adding other information or making additional explanations.
[0110] (4) Answer restriction 2: Do not include the specific item number in the answer, but you can include the item name. For example, do not include toy_01, but you can include "toy car".
[0111] Therefore, according to the above-mentioned preset answer generation rules, the ability to think about answers generated based on understanding can be demonstrated, so that the scenario question and answer model has a more complex answer effect, and at the same time is more in line with the interactive habits of human scenario question and answer, making each scenario answer in the generated scenario answer set more natural and in line with the scenario needs.
[0112] Furthermore, the preset answer generation model is fine-tuned according to the preset answer generation rules, and the problem scenario information is used as input data of the preset answer generation model, so that the preset answer generation model can reorganize the problem scenario information according to the preset answer generation rules to generate a scenario answer (Answer) that matches the corresponding scenario question.
[0113] Among them, the above-mentioned preset answer generation model can be the aforementioned preset question generation model, such as a large language model like GPT-4o, which can provide a two-stage answer generation method based on COT (Chain of Thought), first extracting problem scenario information related to the scenario question from the relevant scenario recall information, and then generating the final matching scenario answer based on the "thinking" process for the above-mentioned problem scenario information, namely Prompt_Final.
[0114] Specifically, the process of generating answers to the above scenarios can be further illustrated by the following example.
[0115] First, set up a scene judgment message System_Msg, such as "SYS_MSG = "You are an AI visual assistant, and you are in a home scene. Your surroundings are described by sentences. You will see multiple objects, and you can see each object's location, color, size, and what objects are in front, behind, to the left, right, above, below, and inside the object. When answering questions, assume you are in the real environment."
[0116] Therefore, the corresponding scenario questions (queries) can be generated as "What color is the potted plant on the ground? Where is the sofa?"
[0117] Further, based on the above scenario question, the corresponding scene recall information (Scene) can be obtained as "The current item information in the room includes 2 sofas. Sofa_1 is in front of you, located in the living room. Its volume is 1995324. Behind it are 1 pink plant, 6 green plants, 1 pink rag, 1 blue rag, and 1 pink plate. It is surrounded by 1 sofa and 1 paper ball. In front of it are 1 green photo, 1 green painting, 1 pink painting, 1 green building block, 3 yellow paintings, 1 blue building block, and 1 silver flower. "Sprinkle, 1 green chair, 1 pink toy, sofa_0, in front of you, located on the ground, its volume is 1344120, inside it are 1 toy, 1 yellow cylindrical coffee table, behind it are 6 green plants, 1 pink dishcloth, 1 pink plate, to its right are 1 pink plant, 1 pink painting, around it are 1 sofa, 1 yellow carpet, 1 toy, in front of it are 1 green photo, 1 green painting, 1 green building block, 1 blue building block, 1 silver shower head, 1 green chair."
[0118] Furthermore, the following problem scenario information can be extracted (i.e., thought) from the above scenario recall information:
[0119] 1. **Sofa Position**:
[0120] - `sofa_0` On the ground, there are green photos, green paintings, green building blocks, blue building blocks, silver shower heads, and green chairs in front.
[0121] - `sofa_1` In the living room, there are green photos, green paintings, pink paintings, green building blocks, yellow paintings, blue building blocks, silver shower heads, green chairs, and pink toys in front.
[0122] 2. **Color of potted plants**:
[0123] - There is a pink plant to the right of `sofa_0`.
[0124] - Behind `sofa_1` there is 1 pink plant and 6 green plants.
[0125] Based on the above question scenario information, the preset answer generation model can be used to fine-tune the preset answer generation rules to reorganize the information and obtain the scenario answers corresponding to the above scenario questions as follows:
[0126] - Color of potted plants on the ground: pink.
[0127] - Position of the sofas: `sofa_0` is on the floor, `sofa_1` is in the living room.
[0128] Therefore, the intelligent agent can accurately and promptly answer questions about the scene content based on the scene it is in, so that the intelligent agent can achieve the scene question-and-answer capability that general artificial intelligence should have. These scene answers constitute the scene answer set, which can enable the intelligent agent to give more fluent and natural answers to scene questions, ensuring that they are more consistent with real human reactions, thereby achieving a higher level of intelligence.
[0129] Figure 4 A flowchart of another application scenario of the method for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention is schematically shown.
[0130] like Figure 2-Figure 4 As shown, according to one embodiment of the present invention, in operation S204, a preset question-answering model is trained according to a scenario question set and a matching scenario answer set to generate a scenario question-answering model of an agent, including:
[0131] Generate a training data set that conforms to a preset data format based on a scenario question set and a matching scenario answer set;
[0132] The training dataset is input into the preset question-answering model according to the preset training rules to generate a scenario question-answering model.
[0133] like Figure 4 As shown, the training data 401 may include training data sets corresponding to the scenario question set and the scenario answer set. The preset data format is a data training format for organizing the scenario question set and the scenario answer set. It can be designed based on the training framework of the preset question-answering model. For example, the input data of the large language model Qwen model can be in the alpaca format, and the data definition of the scenario question set and the scenario answer set can be performed in the instruct-output structure to generate the corresponding training data set.
[0134] The training dataset can be a scenario question-answering dataset defined in the above-mentioned preset data format, which can be used as the designated training data for the preset question-answering model to implement data training of the preset question-answering model to generate the required scenario question-answering model. Specifically, the training dataset can be represented by the following data example 1 and data example 2:
[0135] Data Example 1:
[0136] {
[0137] "system": "Receives input from others, combines [common sense], [scenario information], and [memory] to generate a fluent response to the user's question.",
[0138] "instruction": "[Common sense knowledge]: \tThe foreign name of computer is computer\tThe Chinese name of computer is computer\tThe alias of computer is computer\tThe components of computer belong to computer\n[Scene information]: The current information of items in the room includes: There is no corresponding object in the scene.\n[Memory]:\nQuestion: What is a computer?\n",
[0139] "input": "",
[0140] "output": "A computer is an electronic device that receives information in digital form and performs a series of operations according to predetermined instructions to produce output."
[0141] }
[0142] Data Example 2:
[0143] {
[0144] "system": "Receives input from others, combines [common sense], [scenario information], and [memory] to generate a fluent response to the user's question.",
[0145] "instruction": "[Common Knowledge]: The explanation of a bedside cupboard is: Bedside cupboard. In modern furniture, it is a small cabinet placed on either side of the bed, used for storing miscellaneous items. Its shape is similar to the common bedside cabinets of today. Foreign name: Bedside cupboard, pronunciation: chuang'tou'gui. A bedside book is suitable for the bedside. The book size of "There is a white-faced old woman at the head of the bed" is 16mo. [Scene Information]: The items in the room currently include: 10 books: book_3, on your left, on the desk, has a volume of 1343, with 2 bedside tables, 4 books, and 1 cabinet behind it, and a cupboard to its right. book_1, behind you, on the bookshelf, has a volume of 1445, with 2 books around it, 1 book behind it, 2 bedside tables to its left, and 3 books in front of it. book_4, behind you, on the bookshelf, is blue, has a volume of 731, with 3 books around it, and 1 book behind it. There are 2 books on the left and 2 bedside tables on the right. \tbook_5, behind you, in the bookshelf, its color is blue, its volume is 724, there are 3 books around it, there are 2 books behind it, and there are 2 bedside tables on the left. \tbook_6, behind you, in the bookshelf, its color is blue, its volume is 737, there are 3 books around it, there are 2 books behind it, and there are 2 bedside tables on the left. \tbook_7, behind you, in the bookshelf, its color is blue, its volume is 7 36, it has 3 books around it, 2 books behind it, and 2 bedside tables on its left. \tbook_8, behind you, on the bookshelf, its color is blue, its volume is 716, it has 3 books around it, 2 books behind it, and 2 bedside tables on its left. \tbook_9, behind you, on the bookshelf, its volume is 2857, it has 2 books behind it, 2 bedside tables on its left, and 2 books in front of it. \tbook_0, behind you, on the desk, it The volume of book_2 is 875. There are 3 books behind it and 2 nightstands to its left. Book_2 is behind you, on the bookshelf. Its volume is 1447. There are 2 nightstands to its left and 4 books in front of it. [Memory]: 4 days, 7 hours, and 17 minutes ago, the book was no longer on the nightstand in the bedroom. It was in the yellow TV stand. 4 days, 7 hours, and 17 minutes ago, the book was no longer on the nightstand in the bedroom. It was in the yellow TV stand. 4 days, 7 hours, and 17 minutes ago, the book was no longer on the nightstand in the bedroom.The book is in the yellow TV cabinet.\t4 days, 7 hours, and 17 minutes ago, the book was no longer on the bedside table in the bedroom. The book is in the yellow TV cabinet.\t4 days, 7 hours, and 17 minutes ago, the book was no longer on the bedside table in the bedroom. The book is in the yellow TV cabinet.\t4 days, 7 hours, and 17 minutes ago, the book was no longer on the bedside table in the bedroom. The book is in the yellow TV cabinet.\t4 days, 7 hours, and 17 minutes ago, the book was no longer on the bedside table in the bedroom. The book is in the yellow TV cabinet. t4 days, 7 hours, and 17 minutes ago, the book was no longer on the bedside table in the bedroom. The book was in the yellow TV cabinet.\t4 days, 6 hours, and 42 minutes ago, the book was no longer on the bedside table in the bedroom. The book was in the yellow TV cabinet.\t4 days, 6 hours, and 42 minutes ago, the book was no longer on the bedside table in the bedroom. The book was in the yellow TV cabinet.\t4 days, 6 hours, and 42 minutes ago, the book was no longer on the bedside table in the bedroom. The book was in the yellow TV cabinet.\t4 days, 6 hours, and 42 minutes ago, The book is no longer on the bedside table in the bedroom. The book is in the yellow TV cabinet.\t2 days, 6 hours, and 30 minutes ago, the book is no longer on the bedside table in the bedroom. The book is in the yellow TV cabinet.\t2 days, 6 hours, and 30 minutes ago, the book is no longer on the bedside table in the bedroom. The book is in the yellow TV cabinet.\t2 days, 6 hours, and 30 minutes ago, the book is no longer on the bedside table in the bedroom. The book is in the yellow TV cabinet.\t2 days, 6 hours, and 30 minutes ago, the book is no longer on the bedside table in the bedroom. The book is in the yellow TV cabinet.\t2 days, 26 minutes ago, the book is no longer on the bedside table in the bedroom. The book is in the yellow TV cabinet.\t2 days and 26 minutes ago, the book was not in the bedroom, on the bedside table. The book is in the bedroom, on the yellow TV cabinet.\t2 days and 26 minutes ago, the book was not in the bedroom, on the bedside table. The book is in the bedroom, on the yellow TV cabinet.\t2 days and 26 minutes ago, the book was not in the bedroom, on the bedside table. The book is in the bedroom, on the yellow TV cabinet.\t1 hour and 52 minutes ago, the book was not in the bedroom, on the bedside table.\nQuestion: How many books are there on the right side of the bedside table?\n"
[0146] "input": "",
[0147] "output": "There are 9 books on the right side of the bedside table."
[0148] }
[0149] The preset training rules may be model fine-tuning rules that constrain the data training of the preset question-answering model, such as Lora training rules, wherein the preset training rules may also be set according to the training framework of the preset question-answering model. Figure 4The model 402 shown in FIG. 402, the preset question-answering model can be a common large language model (such as the Qwen model) provided in an embodiment of the present invention, which is used as the base model for generating the scenario question-answering model and does not require large-scale data and a large amount of computing resources. For example, the Qwen2.5 14b model can be trained with the above training data set based on the preset training rules of the Lora training rules to generate the Lora-qwen2.5-14b scenario question-answering model (such as Figure 4 Model 403 is shown), which involves training hyperparameters lr=5e-4 and epoch=3.
[0150] Specifically, a training dataset that conforms to a preset data format can be generated based on the aforementioned scenario question set and scenario answer set. The training dataset can include more scenario questions and data samples of matching scenario answers. Taking the 4700 data samples of the above training dataset as an example, on the same evaluation set, using the same evaluation criteria, after evaluating GPT4o-few-shot model 1, GPT4o-cot model 2, and Lora-qwen2.5-14b model 3, the reasonable rates of the following model question-answering evaluations can be obtained as shown in Table 1:
[0151]
[0152] Model 1 uses a few-shot approach to generate answers using a single prompt from GPT-4o, Model 2 uses the method of generating training data in the embodiment of the present invention, and Model 3 is the model effect after fine-tuning based on the training data. As shown in Table 1, it can be seen that the performance of the fine-tuned model is about 90% of the capability of Model 2 and is better than that of Model 1.
[0153] In other words, the generation method of the above-mentioned scenario question-answering model in an embodiment of the present invention can achieve more than 90% of the interactive capabilities of the GPT-4o model by only constructing and fine-tuning low-cost training data.
[0154] In summary, the method for generating a scenario question-answering model for an intelligent agent provided by an embodiment of the present invention can at least partially solve the problem in related arts of low intelligence levels caused by the inability of intelligent agents to perform scenario question-answering well, and thus can achieve at least one of the following technical effects:
[0155] (1) Enhanced scene question-answering capabilities: The method for generating a scene question-answering model for an intelligent agent provided by the embodiments of the present invention can enable the intelligent agent to understand and answer questions about scene information in a real or virtual environment. This includes recognizing and describing the type, quantity, action, and position of visual objects, as well as the relative relationships between objects. This capability is an important component of developing general artificial intelligence.
[0156] (2) Efficient training method: By constructing low-cost training data and combining it with fine-tuning of large language models, we achieve up to 90% of the GPT-4o model processing power with limited resources. This means that this technology can achieve performance close to that of current state-of-the-art language models without relying on large-scale data and computing resources.
[0157] (3) Versatility: It is applicable not only to real-world scenarios but also to virtual world environments, showing wide applicability and flexibility, which is especially important for developing intelligent agents for cross-domain applications.
[0158] (4) Innovative Framework Design: The overall framework adopts advanced training data construction and model fine-tuning strategies, combined with the advantages of large language models, to provide an efficient and scalable solution, providing new ideas for improving scenario question answering capabilities. These highlights indicate that this technical solution has important commercial and scientific research value in promoting the development of general artificial intelligence.
[0159] Therefore, the above-mentioned scenario question-answering model trained by the method for generating the scenario question-answering model of the intelligent body provided by the embodiment of the present invention can give the embodied intelligent body the ability to understand and answer questions about scenario information in a real or virtual environment, so that it has a more fluent and natural question-answering interaction level, conforms to the interaction habits of normal human question-answering scenarios, and has a higher level of intelligence that is more in line with human reality.
[0160] Based on the above-mentioned method for generating a scene question-answering model of an intelligent agent, the present invention also provides a device for generating a scene question-answering model of an intelligent agent. Figure 5 The device is described in detail.
[0161] Figure 5 The structural block diagram of the device for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention is schematically shown.
[0162] like Figure 5 As shown, the generation device 500 of the scene question-answering model of the intelligent agent in this embodiment includes a question generation module 510, an information acquisition module 520, an answer generation module 530 and a model training module 540.
[0163] The question generation module 510 is used to generate a set of scenario questions related to the scenario the agent is in based on the full scenario information. In one embodiment, the question generation module 510 can be used to perform the operation S201 described above, which will not be repeated here.
[0164] The information acquisition module 520 is used to acquire the scenario recall information related to each scenario question in the scenario question set. In one embodiment, the information acquisition module 520 can be used to perform the operation S202 described above, which will not be repeated here.
[0165] The answer generation module 530 is used to generate a matching scenario answer for each scenario question in the scenario question set based on the scenario recall information to form a scenario answer set. In one embodiment, the answer generation module 530 can be used to perform the operation S203 described above, which will not be repeated here.
[0166] The model training module 540 is used to train the preset question-answering model based on the scenario question set and the matching scenario answer set to generate a scenario question-answering model of the agent. In one embodiment, the model training module 540 can be used to perform operation S202 described above, which will not be repeated here.
[0167] According to an embodiment of the present invention, any multiple modules among the question generation module 510, information acquisition module 520, answer generation module 530, and model training module 540 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to an embodiment of the present invention, at least one of the question generation module 510, information acquisition module 520, answer generation module 530, and model training module 540 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of these. Alternatively, at least one of the question generation module 510, the information acquisition module 520, the answer generation module 530 and the model training module 540 can be at least partially implemented as a computer program module, which can perform the corresponding function when it is executed.
[0168] Figure 6 A block diagram of an electronic device suitable for implementing a method for generating a scenario question-answering model of an intelligent agent according to an embodiment of the present invention is schematically shown.
[0169] The above-mentioned electronic device provided by an embodiment of the present invention includes one or more processors and a memory, and the memory is used to store one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors execute the method for generating the scenario question-answering model of the above-mentioned intelligent agent.
[0170] like Figure 6 As shown, an electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 602 or programs loaded from a storage unit 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0171] Various programs and data required for the operation of the electronic device 600 are stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 executes the programs in the ROM 602 and / or RAM 603 to perform various operations according to the method flow of the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 may also execute the programs stored in the one or more memories to perform various operations according to the method flow of the embodiment of the present invention.
[0172] According to an embodiment of the present invention, electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to bus 604. Electronic device 600 may also include one or more of the following components connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or modem. Communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 610 as needed, so that computer programs read from the removable media can be installed into storage section 608 as needed.
[0173] The present invention also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method for generating the scenario question-answering model of the above-mentioned intelligent agent.
[0174] The computer-readable storage medium may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiment of the present invention.
[0175] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 602 and / or RAM 603 described above, and / or one or more memories other than ROM 602 and RAM 603.
[0176] An embodiment of the present invention also includes a computer program product, which includes a computer program that, when executed by a processor, implements the method for generating the scenario question-answering model of the above-mentioned intelligent agent.
[0177] The computer program includes program codes for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program codes are used to enable the computer system to implement the method provided by the embodiment of the present invention.
[0178] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when the computer program is executed by the processor 601. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0179] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0180] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609 and / or installed from a removable medium 611. When the computer program is executed by the processor 601, the above-described functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0181] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0182] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0183] Those skilled in the art will appreciate that various combinations and / or combinations of features described in the various embodiments and / or claims of the present invention may be made, even if such combinations and / or combinations are not explicitly described in the present invention. In particular, various combinations and / or combinations of features described in the various embodiments and / or claims of the present invention may be made, without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.
[0184] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.
Claims
1. A method for generating a scenario question-answering model of an intelligent agent, characterized in that: include: Generate local scene information of each object with each object in the global scene information as the center; The preset question generation model is fine-tuned according to the preset question generation rules, and the preset scene question samples are input into the preset question generation model as reference samples for scene question generation, thereby generating a local scene question set corresponding to the local scene information of each object. The local scene question set includes multiple questions generated based on the scene description of the local scene information of a single object, which are used to constitute the scene question set; wherein the preset scene question samples are pre-set scene description information related to scene question answering and question information designed based on the scene description information, and are example data of "scene description-generated questions"; the preset question generation rules are preset rules for generating matching scene questions based on the scene description information, including question generation requirements for question objects, questioning methods, question subjects, question scope, and rule combinations; Obtaining scene visual information related to each scene question in the scene question set; performing information recall on the full scene information based on the scene visual information to obtain scene recall information; the scene visual information is visual information within the scene that is related to the corresponding scene question and described in natural language, including object description information related to the scene question; the scene recall information is information related to the corresponding scene question obtained by recalling information related to the scene visual information in the full scene information based on the above-mentioned scene visual information, and includes more scene information related to the objects and relationships between objects described by the scene visual information; Extracting question scenario information from the scenario recall information that matches each scenario question in the scenario question set; fine-tuning a preset answer generation model according to preset answer generation rules, and using the question scenario information as input data for the preset answer generation model, so that the preset answer generation model reorganizes the question scenario information according to the preset answer generation rules to generate a scenario answer that matches the corresponding scenario question; wherein the preset answer generation rules are preset rules related to generation requirements for generating a scenario answer that matches the scenario question based on the question scenario information, including answer generation requirements for answering person, answering method, and answer restrictions; and Training a preset question-answering model based on the scenario question set and the matching scenario answer set to generate a scenario question-answering model for the agent; Among them, the preset answer generation model and the preset question generation model are both large language models that provide a two-stage answer generation method based on the thought chain.
2. The method according to claim 1, characterized in that Before generating the local scene information of each object with each object in the full scene information as the center, the method further includes: In response to the received model generation task instruction, the full scene information related to the scene in which the agent is located is obtained.
3. The method according to claim 1, characterized in that The step of training a preset question-answering model according to the scenario question set and the matching scenario answer set to generate the scenario question-answering model of the agent includes: Generate a training data set that conforms to a preset data format based on the scenario question set and the matching scenario answer set; The training data set is input into the preset question-answering model according to preset training rules to generate the scenario question-answering model.
4. A device for generating a scenario question-answering model of an intelligent agent, used to implement the method according to any one of claims 1 to 3, characterized in that: include: A question generation module, configured to generate a set of scenario questions related to the scenario in which the agent is located based on the full scenario information; An information acquisition module is configured to acquire scene recall information related to each scene problem in the scene problem set, comprising: acquiring scene visual information related to each scene problem in the scene problem set; performing information recall on the full scene information based on the scene visual information to acquire the scene recall information; an answer generation module, configured to generate a matching scenario answer for each scenario question in the scenario question set based on the scenario recall information to form a scenario answer set, comprising: extracting problem scenario information matching each scenario question in the scenario recall information; reorganizing the problem scenario information according to preset answer generation rules to generate corresponding scenario answers, wherein the preset answer generation rules are preset rules related to generation requirements for generating scenario answers matching the scenario questions based on the problem scenario information; and The model training module is used to perform training on the preset question-answering model according to the scenario question set and the matching scenario answer set to generate the scenario question-answering model of the intelligent agent.
5. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to execute the method according to any one of claims 1 to 3.
6. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 3.
7. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Question and answer matching method and system based on question and answer system
CN116860953A
Body scene question answering method and device based on retrieval enhancement generation and electronic equipment
CN119179770A