Method and device for automatically converting agent information into text

By obtaining agent information and generating text descriptions, problems such as agent actions that cannot be used for retrieval and enhance generation are solved, and real-time accuracy of agent dialogue is achieved.

CN120146002AInactive Publication Date: 2025-06-13BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510624072.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the actions of the agent, the changes in the action state, and the perceived actions cannot be used for retrieval and enhance generation, which affects the accuracy of the agent's dialogue.

Method used

By obtaining agent information, including its own actions, action state changes and perceived actions, and automatically generating text descriptions based on preset rules, it is directly applied to the search for enhanced generation RA modules.

Benefits of technology

Real-time accuracy of the agent's dialogue is realized, and the agent's movements, perceived actions and action state changes are converted into text to adapt to the agent's real-time state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146002A_ABST
    Figure CN120146002A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for automatically converting agent information into a text. The method for automatically converting the intelligent agent information into the text comprises the steps that the intelligent agent information is obtained, and the intelligent agent information comprises the action information of an intelligent agent, the change information, caused by the action of the intelligent agent, of the state of an article in a scene and the action information sensed by the intelligent agent; and automatically generating corresponding text description according to a preset rule by using the acquired agent information. Text description is automatically generated through intelligent agent actions, action state changes, perceived actions and the like according to preset rules, the text description can be directly applied to a retrieval enhancement generation RA module, and the actions, the perceived actions and the action state changes of the intelligent agents can be obtained in real time through retrieval enhancement generation in intelligent agent dialogues; the intelligent agent dialogue is accurately adapted according to the real-time state of the intelligent agent, so that the purpose of improving the accuracy of the intelligent agent dialogue is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and in particular, relates to a method and device for automatically converting agent information into text. Background Art

[0002] An agent is an entity that can perceive the environment and take actions to achieve specific goals. It has autonomy, adaptability, and interaction capabilities. By perceiving changes in the environment, the agent makes judgments and decisions based on the knowledge and algorithms it has learned, and then executes actions to affect the environment or achieve a predetermined goal. Agents are widely used in the field of artificial intelligence, commonly found in automated systems, robots, virtual assistants, and game characters, etc. The core lies in the ability to learn autonomously and evolve continuously to better complete tasks and adapt to complex environments.

[0003] Retrieval Augmented Generation (RAG) is one of the current popular cutting-edge technologies for large models. The retrieval augmented generation model combines a language model and information retrieval technology. Specifically, when the model needs to generate text or answer questions, it first retrieves relevant information from a large collection of documents, and then uses this retrieved information to guide the generation of text, thereby improving the quality and accuracy of predictions.

[0004] In the process of implementing this embodiment, the inventors found that in the prior art, agent actions, changes in action states, and perceived actions cannot be used in retrieval augmented generation, thus affecting the accuracy of agent conversations. Summary of the Invention

[0005] Aiming at the problems existing in the prior art, the present invention provides a method and device for automatically converting agent information into text, which at least partially solves the problem that agent actions, changes in action states, and perceived actions in the prior art cannot be used in retrieval augmented generation.

[0006] In a first aspect, an embodiment of the present disclosure provides a method for automatically converting agent information into text, including: Obtain agent information, where the agent information includes the agent's own action information, the information on the change in the state of items in the scene caused by the agent's own behavior, and the action information perceived by the agent; Automatically generate a corresponding text description based on the obtained agent information according to a preset rule.

[0007] Optionally, the automatically generating a corresponding text description based on the obtained agent information according to a preset rule includes: Construct text configuration information based on knowledge data; Construct rule information based on knowledge data and agent motion ability data; Request the configuration file server to obtain agent-related information; Generate corresponding text based on the text configuration information, rule information, and agent-related information.

[0008] Optionally, the generating corresponding text based on the text configuration information, rule information, and agent-related information includes: allocating rule operations according to the rule information and obtaining the corresponding rule operation return results, and integrating all the rule operation return results to generate text information.

[0009] Optionally, the rule information includes status change information and action information. The status change information represents the conditions that need to be met for the action to be executed, and the parameters in the status change information are obtained from the agent information.

[0010] Optionally, the rule information includes the function name, the parameter information required by the function, and the variable name to which the return value after the function execution is assigned; the return value after the function execution is stored in the memory, and the return value is obtained from the memory and used as the actual parameter of the function.

[0011] Optionally, after the step of automatically generating the corresponding text description according to the preset rules for the obtained agent information, it further includes that when it is detected that the state of the object in the scene changes due to the agent's action, according to the difference in the attribute information of the object before and after the state change, the corresponding rule function is called to generate the text describing the state change.

[0012] In a second aspect, an embodiment of the present disclosure further provides a device for automatically converting agent information into text, including: an agent information acquisition module for acquiring agent information, where the agent information includes the agent's own action information, the change information of the item state in the scene caused by the agent's own behavior, and the action information perceived by the agent; A knowledge base for storing common sense knowledge; An action library for representing the motion capabilities possessed by the agent itself; A configuration module for extracting information from the knowledge base and constructing text configuration information for the text generation service; A configuration file for storing the text configuration information; A rule information construction module for constructing rule information according to the knowledge base and the action library; A rule module for storing the rule information; A rule operation module for performing rule operations and returning the result information of the rule operations; A rule engine that reasonably allocates rule operations according to the rule information, integrates all the rule operation return results, and generates text information.

[0013] In a third aspect, embodiments of the present disclosure further provide an electronic device, which includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method for automatically converting agent information into text according to any one of the first aspects.

[0014] In a fourth aspect, embodiments of the present disclosure further provide a computer-readable storage medium, which stores computer instructions for causing a computer to execute the method for automatically converting agent information into text according to any one of the first aspects.

[0015] In a fifth aspect, embodiments of the present disclosure further provide a computer program product, including computer programs / instructions, which when executed by a processor implement the method for automatically converting agent information into text according to any one of the first aspects.

[0016] The method and device for automatically converting agent information into text provided by the present invention, wherein the method for automatically converting agent information into text automatically generates a text description according to preset rules for agent actions, action state changes, perceived actions, etc., and the text description can be directly applied to the RA module of retrieval-augmented generation. In agent conversations, retrieval-augmented generation can be used to obtain the actions, perceived actions, and action state changes of the agent in real time, so that the agent conversation can be precisely adapted according to the real-time state of the agent, thereby achieving the purpose of improving the accuracy of agent conversations. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. Among them, in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.

[0018] Figure 1 is a flowchart of a method for automatically converting agent information into text provided by an embodiment of the present disclosure; Figure 2 is a schematic block diagram of a device for automatically converting agent information into text provided by an embodiment of the present disclosure; Figure 3 is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The embodiments of the present disclosure will be described in detail below in conjunction with the accompanying drawings.

[0020] It should be clear that the following illustrates the embodiments of the present disclosure through specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0021] It should also be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of the aspects described herein can be used to implement a device and / or practice a method. Additionally, this device and / or method can be implemented using other structures and / or functionality in addition to one or more of the aspects described herein.

[0022] It further should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. The diagrams only show the components related to the present disclosure rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.

[0023] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0024] RAG (Retrieval Augmented Generation) is a technical framework that combines retrieval and generation, and is widely used in the field of natural language processing, especially when generating text related to a large amount of background knowledge. The RAG framework mainly includes two core modules: the Retrieval (abbreviated as RA) module and the Generation (abbreviated as GA) module.

[0025] In the RAG framework, the main function of the RA module is to retrieve the most relevant documents or information fragments from a large collection of background knowledge or documents for the current task or context. Specifically, the role of the RA module is as follows: 1. Provide context-related background knowledge: The RA module retrieves relevant documents or fragments from a large number of stored documents or knowledge bases based on the user's question or the context of the generation task. This provides background information related to the question for the subsequent generation module, ensuring that the generated text is factually based and accurate.

[0026] 2. Narrow the search scope: Through the retrieval process, the RA module narrows down the scope of potentially relevant documents, enabling the generation module to generate text based on more precise information. This helps improve the quality and relevance of the generated text.

[0027] 3. Enhance the input of the generation module: The RA module takes the retrieved relevant documents as one of the inputs of the generation module. Together with the user's original input, it provides richer information for the generation module, thereby generating more accurate and detailed text.

[0028] 4. Support the processing of complex tasks: When dealing with complex tasks (such as question-and-answer systems, dialogue systems, summary generation, etc.), the RA module can help the generation module better understand and utilize background knowledge, thus generating text that better meets the user's needs.

[0029] The RA module is an information retriever that is responsible for quickly finding the most relevant knowledge for the current task from a vast amount of data in the RAG framework, providing the necessary background support for the generation module, and making the generated text more accurate, rich, and relevant.

[0030] As Figure 1 shown, this embodiment discloses a method for automatically converting agent information into text, including: Step S101: Obtain agent information, where the agent information includes the agent's own action information, the information on the change in the state of items in the scene caused by the agent's own behavior, and the action information perceived by the agent; The agent information includes the action sequence generated by the cognitive module and the actions issued by the agent itself. Due to the actions made by the agent itself, the information on the change in the scene information occurs. For example, an apple is on the table, and the agent picks up the apple and puts it in the refrigerator. The apple was originally on the table and then becomes in the refrigerator. The action sequence generated by the perception module and the actions made by the agent when seeing other agents. For example, actions such as the agent picking up something and shaking its head.

[0031] The intelligent agent obtains its own action information. The cognitive module of the intelligent agent is responsible for generating its own action sequence and recording relevant action information. For example, when the intelligent agent executes the action of "taking the book on the table and putting it into the schoolbag", the cognitive module will record details such as the type of the action (such as "taking"), the object involved in the action (the book) and its attributes (such as color, size, etc.), the starting point (the table) and the ending point (the schoolbag) of the action. These information will be used as U Action data and stored in a structured manner according to a predefined data format for subsequent text generation processing.

[0032] The Status Change information is obtained by monitoring the impact of the intelligent agent's behavior on the state of objects in the scene. Taking the scene where the intelligent agent moves an apple from the table to the refrigerator as an example, the system will record the state change of the apple before and after the movement. Before the movement, the state of the apple was "is on the table", and after the movement, it became "is in the refrigerator". These state change information not only includes the position change of the object, but may also involve other attribute changes, such as the posture of the object, the relative position relationship with other objects, etc. By comparing the scene states before and after the execution of the intelligent agent's action, the system can accurately capture the item state changes caused by the intelligent agent's behavior and store them in a structured manner, providing data support for generating text describing the scene changes.

[0033] The perception module of the intelligent agent is used to identify the actions of other intelligent agents or objects in the surrounding environment and generate corresponding action sequence information (CV Action). For example, when the intelligent agent observes another intelligent agent executing the action of "shaking the head", the perception module will record information such as the type of the action (shaking the head), the identity of the intelligent agent performing the action (if it can be identified), the timestamp of the action, and the scene location where the action occurs. These CV Action data are also stored in a structured manner so that subsequent texts describing the actions of other intelligent agents or objects can be generated according to preset rules.

[0034] Step S102: Automatically generate corresponding text descriptions based on the obtained intelligent agent information according to preset rules.

[0035] Automatically generate corresponding text descriptions based on the obtained intelligent agent information according to preset rules, including: Construct text configuration information based on knowledge data; The construction of text configuration information is specifically as follows: Extract information such as vocabulary, phrases, and sentence structures related to the actions of the agent and the states of scene objects from CommonBase to construct text configuration information. For example, extract words representing actions (such as "pick up", "place", "move", etc.), adjectives representing object attributes (such as "red", "wooden", etc.), and prepositional phrases representing positional relationships (such as "on...", "in...", etc.). These text configuration information will serve as the basic resource library for generating text descriptions, providing support for vocabulary and expression methods to generate natural, accurate, and language-customized texts based on agent information in the follow-up.

[0036] Construct rule information based on knowledge data and agent motion ability data; Combine the common sense knowledge in CommonBase and the agent motion ability defined in U Motion to construct rule information. The rule information includes the following key parts: State change rule: Defines under what conditions the change of the state of an object in the scene needs to be described. For example, when the position of an object changes, according to the specific position information of the object before and after the change and the attributes of the object, generate a text rule describing the state change. The rule will clearly indicate which parameters (such as object ID, original position, new position, etc.) need to be obtained from the agent information as the judgment conditions and the basis for text generation.

[0037] Action description rule: Define corresponding text generation rules for the actions of the agent itself (U Action) and the actions of other agents or objects (CVAction) respectively. These rules cover how to generate accurate text descriptions of actions based on information such as action type, the objects involved in the action and their attributes, and the action executor. For example, for the action of "Agent A picks up the red book on the table", the rule will define how to organize vocabulary and sentence structures to combine information such as the identity of Agent A (if necessary and available), the action "pick up", the object "red book", and the position "table" into a smooth and natural text description.

[0038] The rule information is stored in a structured form, usually including the function name (corresponding to the specific rule processing function), the parameter information required by the function (parameters obtained from the agent information or determined by other means), and the variable name to which the return value of the function is assigned. In this way, when performing rule operations, the corresponding function can be accurately called according to the rule definition, and the correct parameters can be passed, and at the same time, the result returned by the function is stored in the specified variable for subsequent integration into the final text description.

[0039] Request the configuration file server to obtain agent-related information; Generate corresponding text based on text configuration information, rule information, and agent-related information.

[0040] During the text generation process, if the rule operation requires additional information of the agent (such as name, identity, etc.) to supplement the text content, the rule operation unit will send a request to the configuration file server to obtain this relevant information. The configuration file server stores the detailed configuration data of the agent and can return the corresponding information according to the request. For example, when generating an action description text containing the agent's name, the rule operation unit requests the agent's name information from the configuration file server and passes the obtained information as a variable to the corresponding rule processing function to generate a complete text description.

[0041] Text generation, rule operation execution: The rule engine matches the corresponding rules according to the agent information and assigns rule operation tasks to each Rule Operator. The Rule Operator performs operations according to the rule definition, processes the agent information, and returns the corresponding operation results. For example, for the action of "opening the window" performed by the agent, the corresponding Rule Operator will extract information such as the agent's identity (if necessary), action type "opening", object "window", etc. according to the rules and return a text fragment describing the action "Agent A opened the window".

[0042] Result integration and text generation: The rule engine collects all the results returned by the Rule Operator and integrates them according to the predefined text structure and semantic logic. For example, combine text fragments describing the agent's actions, text fragments describing changes in the scene state, and other relevant supplementary information (such as the agent's expression, time of action execution, etc.) into a complete text description according to a certain word order and logical relationship. The finally generated text information will reflect the agent's behavior and its impact on the scene, as well as other relevant information perceived by the agent.

[0043] With the improvement of the agent's cognitive ability and the dynamic change of the environment, the text generation service will automatically update the rules. For example, when the agent learns a new action pattern or has a more detailed distinction of object attributes, the system will adjust the rule definition in the rule construction unit according to the new cognition and knowledge, and optimize the text configuration information. The rule engine will adopt the updated rules and configuration information in the subsequent text generation process, so as to be able to generate more accurate and more in line with the actual scenario text description, and improve the quality and adaptability of text generation.

[0044] Generate corresponding text based on text configuration information, rule information, and agent-related information, allocate rule operations according to the rule information, and obtain the corresponding rule operation return results, and integrate all the rule operation return results to generate text information.

[0045] The rule information includes status change information and action information. The status change information represents the conditions that need to be met for the action to be executed, and the parameters in the status change information are obtained from the agent information.

[0046] The rule information includes the function name, the parameter information required by the function, and the variable name to which the return value after the function execution is assigned; the return value after the function execution is stored in the memory, and the return value is obtained from the memory and used as the actual parameter of the function.

[0047] After the step of automatically generating the corresponding text description according to the obtained agent information according to the preset rules, it further includes that when it is detected that the state of the object in the scene changes due to the agent's action, according to the difference in the attribute information of the object before and after the state change, the corresponding rule function is called to generate the text describing the state change.

[0048] In this embodiment, according to the action sequence sent by the agent cognitive module, the action state changes, and the action sequence recognized by the perception module, the corresponding text description information is automatically generated and directly used in the RA module in RAG (Retrieval Augmented Generation). As the agent's cognition improves, the text generation service will automatically update the rules for subsequent text generation.

[0049] In this embodiment, according to the agent's own action information, the change information of the state of the items in the scene due to the agent's own behavior, and the perceived action information, the text is automatically constructed.

[0050] Customize the structure and content according to the rules and actions, generate text that conforms to the semantics according to the defined rules, and describe in detail the attribute information of the objects involved, such as color, shape, etc.

[0051] This embodiment also discloses a device for automatically converting agent information into text, including: an agent information acquisition module for acquiring agent information, where the agent information includes the agent's own action information, the change information of the state of the items in the scene caused by the agent's own behavior, and the action information perceived by the agent; A knowledge base for storing common sense knowledge; An action library for representing the movement capabilities of the agent itself; A configuration module for extracting information from the knowledge base and constructing the text configuration information of the text generation service; A configuration file for storing the text configuration information; A rule information construction module for constructing rule information based on a knowledge base and an action library; A rule module for storing rule information; A rule operation module for performing rule operations and returning the result information of the rule operations; A rule engine that reasonably allocates rule operations according to rule information, integrates the return results of all rule operations, and generates text information.

[0052] As Figure 2 shown, Figure 2 the specific functions of each English module are as follows: U Action (Agent Action): The action sequence generated by the cognitive module, the actions issued by the agent itself.

[0053] Status Change: The change information of the scene information caused by the actions made by the agent itself. For example, an apple is on the table, and the agent picks up the apple and puts it in the refrigerator. The apple was originally is on the table and then becomes is in the refrigerator.

[0054] CV Action (Perception Action): The action sequence generated by the perception module, the actions the agent sees other agents making. Such as actions like the agent picking up something and shaking its head.

[0055] CommonBase (Common Sense Knowledge): Refers to common sense knowledge, stored in the OWL format.

[0056] U Motion (Agent's Motion Ability): Represents the motion ability the agent itself possesses. Such as the agent pointing at something, picking up something, etc.

[0057] Vocabulary Build (Text Configuration): Extracts information from CommonBase to construct the text configuration information of the text generation service.

[0058] Properties (Configuration File): Used to store configuration files. In text retrieval, the Properties file is often used to store configuration information, such as the path of the index, query parameters, etc. Through the Properties file, these configuration information can be conveniently managed and modified without directly modifying the code.

[0059] Rule Build (Rule Construction): Constructs rule information based on CommonBase and U Motion.

[0060] Rule Operator (Rule Operation): Execute the Rule operation and return the result information of the Rule operation. Among them, the operator that requires agent information needs to request the profile server to obtain information related to the agent configuration, such as the name information of the agent, etc., and then supplement the text.

[0061] Rule Engine: Reasonably allocate Rule operators according to Rule rules, integrate the return results of all Rule operators, generate text information, and return it to the Client side.

[0062] The result generated by the Rule Engine is finally stored in the database, and the RA module recalls the required text information from this database for dialogue text generation.

[0063] The U motion program is as follows: { "id": "go2bed11111111ec843da9e93dc92099", "name": "Go to sleep", "args": { "type": "object", "properties": { "object_id": { "type": "string", "key_in_qg": "instance_id", "range": "http: / / www.semanticweb.org / bigai / ontologies / 2022 / 9 / common_base#bed" , "required": true }}}}.

[0064] Rules: Associate the content passed in by CV action, U action, and Status Change with the functions in the Rule operator. Generate text from the passed-in content. Each Rule contains condition information and action information. The condition information indicates the conditions that need to be met for the action to be executable. Generally, the actual parameters in the condition are obtained from uaction, cv action, and status change. rule_name represents the name of the rule.

[0065] op represents the name of the function to be used.

[0066] kwargs represents the parameter information required by the corresponding function.

[0067] variables: Appear in the condition, indicating the variable names to which the return value after the execution of the op function is assigned. The variables in the condition will all exist in memory, and the action can obtain these variable values from memory and use them as the actual parameters of op.

[0068] The return value of op in the action is directly assigned to the key corresponding to this dict.

[0069] Rule rules are diverse and include operations corresponding to functions map, operators and, or, etc.

[0070] For example, in the rule UAction: sleep111111111ec843da9e93dc92099, the action is replace_text and the condition is get_agent_name.

[0071] The parameters of replace_text are templates and agent. There are no variables, indicating that there is no need to process its return value information.

[0072] The parameters of get_agent_name are agent_id. The agent_id obtains the agent_id information from the information input by the text generation service. "expr": "u.agent_id" indicates that the agent_id is the value of the agent_id key obtained from the u action / cv action / statuschange.

[0073] The variable "agent" means that the value returned after the execution of the "get_agent_name" method is assigned to the variable "agent", and the agent information is stored in the memory.

[0074] Before executing the "replace_text" method, it is necessary to execute the prerequisite function "get_agent_name" to obtain the parameter information "agent" of "replace_text". When executing "replace_text", the value of the "agent" variable in the memory is obtained through the function "get_variable", and this value is passed as the parameter "agent" of "replcace_text" into the function. The value of the parameter "templcate" is "{agent} sleeps". The Rules program is as follows: { "content": { "rule_name": "UActionEntrypoint", "rules": { "condition": { "op": "get_rule_name", "kwargs": { "u_id": { "op": "get_json_path", "kwargs": { "expr": "u.u_id", "return_list": false } }, "status_code": { "op": "get_json_path", "kwargs": { "expr": "u.status_code", "return_list": false } } }, "variables": "rule_name" }, "action": {​ "op": "run", "kwargs": { "rule_name": { "op": "get_variable", "kwargs": { "variable_name": "rule_name" }}}}}]}, { "rule_name": "CVActionEntrypoint", "rules": { "condition": { "op": "get_json_path", "kwargs": { "expr": "instances[?(@.instance_type == 'Action')]", "return_list": true }, "variables": "action_description" }, "action": { "op": "flatten", "kwargs": { "iterable": { "op": "map", "kwargs": { "func": { "op": "get_function", "kwargs": { "function_name": "describe_action_instance" } }, "arg_list": { "op": "get_variable", "kwargs": { ​"variable_name": "action_description" }}}}}}}]}, { "rule_name": "StatusChangeEntrypoint", "rules": { "condition": { "op": "get_json_path", "kwargs": { "expr": "previous.keys()", "return_list": false }, "variables": "instance_id" }, "action": { "op": "flatten", "kwargs": { "iterable": { "op": "map", "kwargs": { "func": { "op": "get_function", "kwargs": { "function_name": "describe_status_change" } }, "arg_list": { "op": "get_variable", "kwargs": { "variable_name": "instance_id" }}}}}}}]}, { "rule_name": "CVPGEntrypoint", "rules": { ​"condition": { "op": "get_json_path", "kwargs": { "expr": "previous.keys()", "return_list": false }, "variables": "instance_id" }, "action": { "op": "flatten", "kwargs": { "iterable": { "op": "map", "kwargs": { "func": { "op": "get_function", "kwargs": { "function_name": "describe_status_change" } }, "arg_list": { "op": "get_variable", "kwargs": { "variable_name": "instance_id" }}}}}}}]}, { "rule_name": "UAction:sleep111111111ec843da9e93dc92099", "rules": { "condition": { "op": "get_agent_name", "kwargs": { "agent_id": { "op": "get_json_path", "kwargs": {​ "expr": "u.agent_id", "return_list": false } } }, "variables": "agent" }, "action": { "op": "replace_text", "kwargs": { "templates": "{agent} sleeps" , "agent": { "op": "get_variable", "kwargs": { "variable_name": "agent" }}}}}]} ]}。

[0075] The U Action data program is as follows: { "belief_pg_id":"69ee3f06-aa6c-4607-8aa8-2ee94b1187ce", "instance_list": { "instance_id":"toyFloor", "common_base:isIn": "ToyRoom" , "rdf:type": "common_base:floorBoard" }, { "common_base:shape": "unknown" , "instance_id":"PaperBall_0", "rdf:type":​​ "common_base:paperBall" , "common_base:hasColour": "unknown" , "common_base:isOn": "toyFloor" }, { "common_base:shape": "cylinder" , "instance_id":"Bucket", "rdf:type": "common_base:wastebasket" , "common_base:hasColour": "yellow" , "common_base:isOn": "toyFloor" } , "u_info":{ "start_time":1700474280469296, "parent_id":"9d10c6d9-2cc9-497c-8348-c52663659392", "args":{ "properties":{ "obj2pick":"PaperBall_0", "place2put":{ "properties":{ "object_id":"Bucket", "direction":"inside" }}}}, "end_time":1700474280469296, ​​"change_time": 1700474280470527, "status_code": "executed", "unique_id": "0d52079a-eb33-4fe2-9fde-b7fc8f4d59fb", "u_id": "UID#relocate111111ec843da9e93dc92099", "agent_id": "agent_1798255144463241216" } } Status Change data program is as follows: { "previous": { "PaperBall_0": { "rdf:type": "common_base:paperBall" , "common_base:hasColour": "unknown" , "common_base:shape": "unknown" , "common_base:CVReliability": 0.2378875534604251 , "common_base:isOn": "toyFloor" }}, "current": { "PaperBall_0": { "rdf:type": "common_base:paperBall" , "common_base:isIn": "Bucket" , "common_base:hasColour": "unknown" , "common_base:shape": "unknown" , "common_base:CVReliability": 0.2378875534604251 }, "Bucket":{ "rdf:type": "common_base:wastebasket" , "common_base:hasColour": "yellow" , "common_base:shape": "cylinder" , "common_base:CVReliability": 1 , "common_base:isOn": "toyFloor" }, "toyFloor": { "rdf:type": "common_base:floorBoard" , "common_base:CVReliability": 0.9883184432983398 , "common_base:hasColour": "unknown" , "common_base:shape": "unknown" , "common_base:isIn": "ToyRoom"​​}}, "current_time": 1111111111111114, "previous_time": 1111111111111111 } CV action { "Action_42_S_grabWithSingleHand": { "common_base:coordinateX": 0.0 , "common_base:coordinateY": 0.0 , "common_base:coordinateZ": 0.0 , "common_base:endTime": 0 , "common_base:hasAgent": "Robotbaby" , "common_base:hasColour": "unknown" , "common_base:hasPatient": "plate_3" , "common_base:inferTime": 1743580997119603 , "common_base:isPoweredOn": false , "common_base:isTurnedOn": false , "common_base:maxVertexX": 0.0 , "common_base:maxVertexY": 0.0 , "common_base:maxVertexZ": 0.0 , "common_base:minVertexX": 0.0 , "common_base:minVertexY": 0.0 , "common_base:minVertexZ": 0.0 , "common_base:quaternionOmega": 1.0 , "common_base:quaternionX": 0.0 , "common_base:quaternionY": 0.0 , "common_base:quaternionZ": 0.0 , "common_base:shape": "unknown" , "common_base:startTime": 1743580997119603 , "instance_type": "Action" , "isUpright": true , "localMaxVertexX": 0.0 , "localMaxVertexY": 0.0 , "localMaxVertexZ": 0.0 , "localMinVertexX": 0.0 , "localMinVertexY": 0.0 , "localMinVertexZ": 0.0 , "lock": false, "rdf:type": "common_base:grabWithSingleHand" , "scale": 1.0, 1.0, 1.0 , "visible": true , "common_base:exposeRate": 1.0 , "__beliefReliability": 1 }}。

[0076] The electronic device disclosed in this embodiment includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0077] The processor can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions. In one embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the electronic device executes all or part of the steps of the method for automatically converting agent information into text in the various embodiments of the present disclosure described above.

[0078] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, this embodiment may also include well-known structures such as communication buses, interfaces, etc., and these well-known structures should also be included in the protection scope of the present disclosure.

[0079] Such as Figure 3 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. It shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present disclosure. Figure 3 The shown electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0080] Such as Figure 3 As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) or the program loaded from the storage device into the random access memory (RAM). In the RAM, various programs and data required for the operation of the electronic device are also stored. The processing device, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.

[0081] Generally, the following devices can be connected to the I / O interface: input devices including, for example, sensors or visual information acquisition devices; output devices including, for example, display screens; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. The communication device can allow the electronic device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 3 the shown electronic device has various devices, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0082] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, all or part of the steps of the method for automatically converting agent information into text according to the embodiments of the present disclosure are performed.

[0083] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details are not repeated herein.

[0084] The computer-readable storage medium disclosed in this embodiment stores non-temporary computer-readable instructions. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the method for automatically converting agent information into text according to the foregoing embodiments of the present disclosure are performed.

[0085] The above-mentioned computer-readable storage medium includes but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).

[0086] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details are not repeated herein.

[0087] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for illustrative and facilitating understanding purposes, and not for limitation. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0088] In this disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.

[0089] In addition, as used herein, "or" in a list of items starting with "at least one" indicates a disjunctive list, so that for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Further, the term "exemplary" does not mean that the examples described are preferred or better than other examples.

[0090] It should also be noted that in the systems and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.

[0091] Various changes, substitutions, and alterations to the technologies described herein can be made without departing from the teachings defined by the appended claims. In addition, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.

[0092] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0093] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those of skill in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A method for automatically converting agent information into text, characterized in that: include: Obtaining agent information, including agent action information, information about changes in the state of objects in the scene caused by the agent's own behavior, and action information perceived by the agent; The acquired agent information is automatically generated into corresponding text descriptions according to preset rules.

2. The method for automatically converting agent information into text according to claim 1, characterized in that: The acquired agent information is automatically generated into a corresponding text description according to a preset rule, including: Build text configuration information based on knowledge data; Construct rule information based on knowledge data and agent motion ability data; Request the configuration file server to obtain agent related information; Generate corresponding text based on text configuration information, rule information and agent-related information.

3. The method for automatically converting agent information into text according to claim 2, characterized in that: The generating of corresponding text based on text configuration information, rule information and agent related information includes: Assign rule operations according to rule information and obtain corresponding rule operation return results, integrate all rule operation return results, and generate text information.

4. The method for automatically converting agent information into text according to claim 2, characterized in that: Rule information includes state change information and action information. The state change information indicates the conditions that need to be met for the action to be executed, and the parameters in the state change information are obtained from the agent information.

5. The method for automatically converting agent information into text according to claim 4, characterized in that: The rule information includes the function name, the parameter information required by the function, and the variable name to which the return value after the function is executed is assigned; the return value after the function is executed is stored in the memory, and the return value is obtained from the memory and used as the actual parameter of the function.

6. The method for automatically converting agent information into text according to claim 1, characterized in that: After the step of automatically generating a corresponding text description based on preset rules for the acquired intelligent agent information, the method also includes calling a corresponding rule function to generate text describing the state change based on the difference in attribute information of the object before and after the state change when it is detected that the state of an object in the scene has changed due to the action of the intelligent agent.

7. A device for automatically converting agent information into text, characterized in that: include: An agent information acquisition module is used to acquire agent information, including agent action information, information about changes in the state of objects in the scene caused by the agent's own behavior, and action information perceived by the agent; Knowledge base, used to store common sense knowledge; Action library, used to represent the movement capabilities of the agent itself; Configuration module, used to extract information from the knowledge base and construct text configuration information for text generation services; Configuration file, used to store text configuration information; A rule information building module is used to build rule information based on the knowledge base and action base; A rule module, used to store rule information; The rule operation module is used to execute the rule operation and return the result information of the rule operation; The rule engine reasonably allocates rule operations according to rule information, integrates the return results of all rule operations, and generates text information.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for automatically converting intelligent agent information into text as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the method for automatically converting intelligent agent information into text as described in any one of claims 1-6.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the method for automatically converting agent information into text as described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Robot trajectory language description generation method and device and readable storage medium

    CN117506940A

  • Intelligent agent behavior determination method, computer equipment and storage medium

    CN117828039A

  • Multi-agent text information intelligent extraction system, method and equipment based on large language model and medium

    CN119312900A

  • Human-computer interaction method and device based on body-equipped intelligent agent and body-equipped intelligent agent

    CN119376549A

  • Robot control method and device, equipment and medium

    CN119407765A