Intelligent agent control instruction generation method and device, equipment, medium and product
By splitting natural language into split units and generating corresponding functions, building capability logical expressions, the problem of tight coupling of semantic analysis and action execution logic in the existing technology is solved, and flexible reuse and efficient analysis of agent control instructions are realized.
Patent Information
- Application Number
- CN202510624073.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, semantic analysis and action execution logic are tightly coupled, making it difficult to reuse and transfer flexibly; natural language expression is highly context-dependent, and the representation method does not conform to the Chomsky paradigm, which is easy to cause ambiguity; the lack of a unified intermediate representation structure hinders the standardization of cross-module collaboration and planning interfaces.
By splitting natural language into split units, corresponding functions are generated based on these units and trained function models, capability logical expressions are constructed, and agent control instructions are finally generated, semantic analysis and internal logic of agent are decoupled.
This enables the generation method of agent control instruction to be flexibly reused and easy to migrate, improves the analytical accuracy and execution efficiency of natural language instructions, and promotes the standardization of cross-module collaboration and planning interfaces.
Smart Images

Figure CN120146030A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and in particular, relates to a method, device, equipment, medium and product for generating intelligent agent control instructions. Background Art
[0002] With the development of artificial intelligence, especially natural language processing (NLP), language models, and autonomous intelligent agent (Agent) systems, how the system understands, parses, and executes the action requirements expressed in human natural language has become an important issue for intelligent agent system decision-making and interaction. Although currently, with the support of language models, the analysis, understanding, and decision-making capabilities of autonomous intelligent agents have been significantly improved, in complex scenarios such as multi-modal perception, multi-task planning, and multi-agent collaboration, existing methods face many challenges. In the process of implementing this embodiment, the inventors found the following problems in the prior art: 1. The semantic parsing and action execution logic are tightly coupled, making the upstream action planning highly dependent on the downstream action design and difficult to be flexibly reused and migrated; 2. Natural language expressions are highly context-dependent and their representation does not fully conform to the Chomsky normal form, which is prone to ambiguity; 3. There is a lack of a unified intermediate representation structure between natural language and autonomous intelligent agents, which hinders the standardization of cross-module collaboration and planning interfaces. Summary of the Invention
[0003] In view of the problems existing in the prior art, the present invention provides a method, device, equipment, medium and product for generating intelligent agent control instructions, which at least partially solves the problem that the semantic parsing and action execution logic in the prior art are tightly coupled and difficult to be flexibly reused and migrated.
[0004] In a first aspect, an embodiment of the present disclosure provides a method for generating intelligent agent control instructions, including: Splitting the obtained natural language according to a set rule to obtain splitting units; Generating a function corresponding to each splitting unit based on the splitting unit and a trained function model; Constructing an ability logic expression corresponding to the obtained natural language based on the function corresponding to each splitting unit; Generating intelligent agent control instructions based on the ability logic expression.
[0005] Optionally, the function includes an action function, a calculation function, and a special processing function.
[0006] Optionally, the action function includes various actions that the intelligent agent can execute; The calculation function includes numerical operations and set operations.
[0007] Optionally, the function includes a function representation and a parameter introduction. The function representation cannot be empty, and the parameter introduction can be empty.
[0008] Optionally, the ability logic expression corresponding to the natural language obtained by constructing based on the function corresponding to each split unit includes: Adding a corresponding identifier to the function, and rewriting the function based on the identifier to obtain the ability logic expression.
[0009] Optionally, adding a corresponding identifier to the function includes: Adding symbols to the function name in the function representation, and adding symbols to the parameters in the function representation.
[0010] In a second aspect, an embodiment of the present disclosure further provides an intelligent agent control instruction generation device, including: a splitting unit configured to split the obtained natural language according to a set rule to obtain split units; a function generation unit configured to generate a function corresponding to each split unit based on the split unit and a trained function model; a construction unit configured to construct an ability logic expression corresponding to the obtained natural language based on the function corresponding to each split unit; an instruction generation unit configured to generate an intelligent agent control instruction based on the ability logic expression.
[0011] In a third aspect, an embodiment of the present disclosure further provides an electronic device, which includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the intelligent agent control instruction generation method according to any one of the first aspects.
[0012] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the intelligent agent control instruction generation method according to any one of the first aspects.
[0013] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, including computer programs / instructions, which, when executed by a processor, implement the intelligent agent control instruction generation method according to any one of the first aspects.
[0014] The intelligent agent control instruction generation method, device, equipment, medium and product provided by the present invention. Among them, for the intelligent agent control instruction generation method, by converting natural language into functions, and then constructing intelligent agent control instructions based on the functions. By parsing the natural language, the corresponding splitting units are converted into corresponding functions. The natural language data is standardized through the functions, and then the intelligent agent control instructions are constructed based on the functions, decoupling semantic parsing from the internal logic of the intelligent agent, so as to achieve the purpose that the instruction generation method can be flexibly reused and easily migrated. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features and advantages of the present disclosure will become more obvious. Among them, in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0016] Figure 1 It is a flowchart of an intelligent agent control instruction generation method provided by an embodiment of the present disclosure; Figure 2 It is a schematic block diagram of an intelligent agent control instruction generation device provided by an embodiment of the present disclosure; Figure 3 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following will describe the embodiments of the present disclosure in detail in conjunction with the accompanying drawings.
[0018] It should be clear that the following illustrates the implementation manners of the present disclosure through specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.
[0019] Note that the following description relates to various aspects of embodiments within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is for illustrative purposes only. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. Additionally, this apparatus and / or method can be implemented using other structures and / or functionality in addition to one or more of the aspects described herein.
[0020] It should also be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. Only the components related to the present disclosure are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0021] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the aspects described can be practiced without these specific details.
[0022] Morpheme: It is the smallest meaningful unit in a language. It can be a single character or a combination of multiple characters. For example, in the word "trees", both "tree" and "wood" are morphemes, with the meanings of "tree" and "wood" respectively; in "butterfly", "butterfly" is an integral morpheme and cannot be split into individual parts with independent meanings. In this embodiment, natural language can be split into morphemes, or natural language can be split based on tokens.
[0023] For ease of understanding, as Figure 1 shown, this embodiment discloses an agent control instruction generation method, including: Step S101: Split the obtained natural language according to set rules to obtain split units; First, split the obtained natural language according to set rules to obtain split units. The set rules here can be formulated according to various rules, such as splitting according to grammatical structure, semantic units, keywords, etc.
[0024] For a simple example, assume that the obtained natural language is "Let the robot walk forward 3 meters first, then pick up the ball on the ground, and put it into the box beside". According to the set semantic and sequential splitting rules, it can be split into the following split units: 1. "Let the robot walk forward 3 meters first" 2. "Then pick up the ball on the ground" 3. "And put it into the box beside" This splitting method can decompose a complex natural language instruction into multiple relatively simple and semantically clear sub-units, which is convenient for subsequent function matching and processing for each split unit respectively. The reasonable formulation of the splitting rules needs to comprehensively consider factors such as the grammar rules, semantic features of natural language, and the actual functions and application scenarios of the intelligent agent, so as to ensure that the split units can accurately reflect the user's instruction intention and can be effectively mapped to the functions in the function library.
[0025] Step S102: Generate functions corresponding to each split unit based on the split unit and the trained function model; Functions are the key resources for realizing the conversion from natural language to intelligent agent control instructions, and they include action functions, calculation functions, and special processing functions.
[0026] Action functions: cover various actions that the intelligent agent can perform. For example, for a mobile robot, the action functions can include "walk forward (parameter: distance)", "backward (parameter: distance)", "turn left (parameter: angle)", "turn right (parameter: angle)", "grab an object (parameter: object name or position)", "place an object (parameter: placement position)", etc. These action functions directly correspond to the specific operation behaviors of the intelligent agent in the physical world and are the basic action units for the intelligent agent to complete tasks.
[0027] Calculation functions: mainly involve numerical operations and set operations. Numerical operations include basic arithmetic operations such as addition, subtraction, multiplication, division, and modulo, and set operations include operations such as finding the intersection, union, and difference of sets. These calculation functions play an important role in processing natural language instructions involving quantity, position calculation, data set processing, etc. For example, when requirements such as "calculate the sum of the distances between two positions" or "find the set of all objects that meet specific conditions" appear in natural language, the corresponding calculation functions need to be called to implement.
[0028] Special processing functions: are used to handle some special situations that do not conform to conventional action and calculation logics. For example, functions for processing complex semantic phenomena such as fuzzy expressions, metaphors, and context dependencies in natural language. For example, when the user says "Let the robot get to that place quickly", the "quickly" here is a relatively fuzzy concept, and a special processing function may be needed to convert it into specific walking speed parameters, or to determine a reasonable speed value according to the actual situation and preset rules of the intelligent agent.
[0029] Each function consists of two parts: a function representation and a parameter introduction. The function representation cannot be empty. It clarifies the name and basic function of the function and is the key identifier for subsequent function matching and invocation. The parameter introduction can be empty, but when there are parameters, it can elaborate on the meaning, type, value range, etc. of each parameter. This is crucial for accurately mapping the instruction details in natural language to function parameters and ensuring that the generated intelligent agent control instructions more precisely conform to the user's intention.
[0030] Generate functions corresponding to each splitting unit based on the splitting unit and the trained function model. Here, the function model is trained with a large amount of natural language and function call example data, and it can learn the mapping relationship between different expressions of natural language and the corresponding functions.
[0031] Taking the previous splitting unit "Let the robot walk forward 3 meters" as an example, the trained function model can identify the key information in it. For example, "walk forward" corresponds to the "walk forward" function in the action function, and "3 meters" is the distance parameter of this function, thus generating the corresponding function as "walk forward (distance = 3 meters)".
[0032] For the splitting unit "Then pick up the ball on the ground", the function model can determine the "grab object" function in the corresponding action function, and the parameters may be determined as parameters related to the object position or name according to the description "the ball on the ground", such as "grab object (object position = ground, object name = ball)".
[0033] Step S103: Construct the ability logic expression corresponding to the obtained natural language based on the functions corresponding to each splitting unit; After obtaining the functions corresponding to each splitting unit, next, construct the ability logic expression corresponding to the obtained natural language based on these functions. Specifically, add corresponding identifiers to the functions. The ability logic expression constructed in this way not only clearly reflects the structure of each function and its parameters but also provides a unified and standardized expression form for subsequent generation of intelligent agent control instructions, making the instruction generation process more organized and facilitating the parsing and execution by the intelligent agent's control system.
[0034] Step S104: Generate intelligent agent control instructions based on the ability logic expression.
[0035] After constructing the ability logic expression, intelligent agent control instructions can be generated based on this expression. The control system of the intelligent agent can identify and parse each function and its parameters in the ability logic expression and, according to the preset execution logic and the behavior control algorithm of the intelligent agent, convert these functions into specific operation instructions that the intelligent agent can execute sequentially or in parallel.
[0036] Converting the ability logic expression into an executable specific operation instruction includes the following steps: Ability logic expression parsing: Initialize an empty function and its parameter list.
[0037] Scan the ability logic expression character by character. When encountering a function name, record the function name.
[0038] Then continue scanning until encountering the parameter start delimiter (such as a left parenthesis), and start parsing the parameter part.
[0039] For the parameter part, identify the type (such as numeric, string, boolean, etc.) and value of the parameter.
[0040] There may be multiple parameters, separated by a delimiter (such as a comma). Parse each parameter in sequence and add them to the parameter list of the current function.
[0041] When encountering the parameter end delimiter (such as a right parenthesis), complete the parsing of the function and its parameters, and add it to the function and its parameter list.
[0042] If there are multiple function calls in the expression, continue to parse the next function until the entire expression is parsed.
[0043] Task planning: Determine the execution order according to the dependency relationship (if any) between functions. For example, if the output of function A is the input of function B, then function A is executed first, and then function B.
[0044] For functions that can be executed in parallel, mark and group them. For example, if two functions have no data dependency and the hardware resources of the agent support parallel execution, they can be executed in parallel.
[0045] Generate a task execution plan, arranging tasks in sequential or parallel manner.
[0046] Task scheduling: According to the sequence information in the task execution plan, put the tasks into the task execution queue in sequence.
[0047] For the grouped parallel tasks, arrange the parallel execution according to the priority of the group or the availability of the agent resources.
[0048] If resource competition occurs during the task execution (such as multiple tasks needing to use the same sensor simultaneously), arbitration is carried out according to the priority of the tasks, and the high-priority tasks obtain resources first.
[0049] Task execution: Take the task out of the task execution queue, or get the current task to be executed from the task execution group.
[0050] According to the function corresponding to the task, call the corresponding operation module inside the intelligent agent.
[0051] Convert the function parameters into specific parameters of the operation instruction. For example, convert the speed value in the function parameter into the speed control signal of the intelligent motor.
[0052] Execute operation instructions and monitor the results of the operation (such as whether it is successfully executed, execution time, etc.).
[0053] If the operation fails, it is handled according to the preset error handling strategy (such as retry, skip, etc.).
[0054] The ability logic expression constructed through the above technology can be converted into a series of control signals by the intelligent body control system, such as sending a motor drive instruction to make the robot walk forward 3 meters, and then controlling the robotic arm to perform a grasping action to grasp an object named a ball on the ground, thereby realizing the task requirements issued by the user through natural language and completing precise control of the intelligent body.
[0055] In the entire process of the intelligent agent control instruction generation method, each step is closely connected and coordinated with each other, making full use of natural language processing technology, function models and instruction generation, realizing efficient and accurate conversion from natural language to intelligent agent control instructions, providing users with a more convenient and natural human-computer interaction experience, and promoting the widespread application and development of intelligent agent technology in various fields.
[0056] The method for generating intelligent agent control instructions provided in this embodiment effectively solves the complexity and inconvenience of intelligent agent control in the prior art by a series of innovative technical steps, such as splitting natural language, generating corresponding functions based on function models, constructing capability logic expressions, and finally generating intelligent agent control instructions. Its detailed and reasonable function model construction and effective processing of the mapping relationship between natural language and function make this method widely applicable to various intelligent agent application scenarios.
[0057] In this embodiment, ensuring the accuracy of the split unit is a key link in achieving accurate conversion of natural language to intelligent agent control instructions. The following are some methods and strategies for ensuring the accuracy of the split unit: Using advanced natural language processing technology: Selection of word segmentation technology, using mature word segmentation algorithms. For example, for Chinese natural language, a word segmentation method based on dictionary matching (such as forward maximum matching method, reverse maximum matching method) combined with a word segmentation algorithm of statistical language model can be selected. The forward maximum matching method scans the text sequence to be segmented from left to right and takes the longest dictionary word as the word segmentation result; the reverse maximum matching method is the opposite, scanning from right to left. These two methods can effectively process the word segmentation of some simple texts, but for some ambiguous situations (such as the sentence "study the origin of life", the word segmentation result may be "study / origin of life" or "study life / origin"), semantic understanding may be needed to further accurately split.
[0058] Utilize deep learning word segmentation models. Such as a word segmentation model based on the structure of bidirectional long short-term memory network (BiLSTM)-conditional random field (CRF). BiLSTM can capture the context semantic information of the text, and the CRF layer can globally optimize the word segmentation result, thus generating a more accurate word segmentation result. By training on a large-scale natural language corpus, such models can learn the semantic features of vocabulary and context-dependent relationships, and can achieve more accurate splitting for processing complex natural language texts containing polysemous words and variable grammatical structures.
[0059] Syntax analysis assistance, • Use dependency syntax analysis. Dependency syntax analysis constructs a dependency relationship tree between words to clarify the grammatical structure of the sentence. For example, in the sentence "Let the robot take the book from the table to the sofa", dependency relationship analysis can determine that "robot" is the execution subject, "take" is the main action, "book" is the object of the action, and "from the table" and "to the sofa" respectively represent the starting point and ending point of the action. The clarification of this grammatical structure helps to reasonably split the sentence into different functional units, such as the action execution unit (the robot takes the book), the position starting point unit (from the table), and the position ending point unit (to the sofa), thus providing a clear splitting basis for subsequent function matching and instruction generation.
[0060] Adopt constituent syntax analysis. Constituent syntax analysis decomposes a sentence into phrase structure components, such as noun phrases, verb phrases, etc. It can identify various phrase boundaries and structural levels in a sentence. For example, for the sentence "The fast-moving robot stopped in the corner of the room", constituent syntax analysis can divide the noun phrase "The fast-moving robot" as the subject, the prepositional phrase "in the corner of the room" as the locative adverbial, and the verb phrase "stopped" as the action. This way of splitting based on phrase components can better conform to the grammatical norms of natural language, improve the accuracy of splitting units, and avoid splitting mistakes caused by incorrect grammatical structures.
[0061] Use a language model with a large number of parameters. The language model learns the patterns and rules of language through unsupervised learning on a large-scale text corpus. Learn language knowledge through pre-training. Due to the powerful semantic understanding ability of the language model, it can perform word segmentation more accurately, especially when dealing with ambiguous words and new words. The language model adopted in this embodiment is not limited and can be a self-developed language model or a large language model such as GPT (Generative Pre-trained Transformer) for processing.
[0062] Construct high-quality training data and rule base: Diversify and refine the annotation of training data, and collect natural language instruction data in a wide range of fields. Include natural language samples in different scenarios (such as industrial robot control instructions, smart home control instructions, game agent control instructions, etc.) and different levels of complexity (from simple single-action instructions to complex multi-step combined instructions). For example, for the smart home control scenario, collect various types of instructions such as "Turn on the lights in the living room", "Lower the air conditioner temperature in the bedroom by 2 degrees", "Let the sweeping robot clean the living room and kitchen" to ensure that the model can learn various expression forms and instruction structures.
[0063] Fine-annotate the collected data. The annotation content includes the boundary of each split unit of the natural language instruction, the semantic category of each split unit (such as action, object, location, time, etc.), and the corresponding intelligent agent control intention. For example, for the instruction "Remind me to have a meeting at 3 o'clock", the split units are annotated as "at 3 o'clock" (time unit) and "Remind me to have a meeting" (action-object unit), and at the same time, the intention of this instruction is annotated as setting a reminder. Through this kind of fine annotation, the training model can learn the manifestation forms and position characteristics of different semantic components in natural language, so as to more accurately identify and distinguish each split unit in the actual splitting process.
[0064] Formulate reasonable splitting rules and patterns, and formulate rules based on semantic and syntactic features. For example, it is stipulated that when encountering words representing time (such as "today", "tomorrow", "morning", "afternoon", "o'clock", etc.) and their surrounding contexts conform to the syntactic position of adverbial of time, they are regarded as an independent time split unit; when a verb representing an action (such as "open", "close", "move", "grab", etc.) and its related object (such as the object of the action) appear, they are combined into an action-object split unit. At the same time, considering the flexibility of language, some rules can also be formulated to handle special cases such as elliptical sentences and inverted sentences. For example, for the instruction "Get me a pen" with the subject omitted, it can be supplemented according to the context (such as the previously mentioned subject "robot") and split into the corresponding subject (executor) and action-object unit.
[0065] Continuously optimize the priorities of rules and conflict resolution strategies. When multiple rules may apply to a natural language segment simultaneously, determine reasonable rule priorities. For example, for a complex sentence that contains both time expressions and action-object expressions, first split the time units according to the grammar rules of time adverbs, and then split the remaining part according to the relevant rules of action-object. Moreover, for possible rule conflict situations (such as different splitting results of the same text segment by two different rules), establish a conflict resolution mechanism, such as selecting the most appropriate splitting method according to criteria such as semantic consistency, grammatical correctness, and the rationality of the agent control instructions, and continuously optimize these rules and strategies through verification and adjustment of a large number of actual cases to improve the accuracy of splitting units.
[0066] Introduce context understanding and semantic analysis, and use context information for dynamic splitting adjustment, memorizing and referring to the dialogue history. In a multi-round dialogue scenario, the agent needs to remember the previous dialogue content and context information to accurately split the current natural language instruction. For example, if the user first says "Let the robot go to the study", and then says "Put the book on the table", the agent should be able to combine the context of the previous round of dialogue, know that the "book" is the book mentioned before, and that "Put on the table" is a location-related action, so as to correctly split the second instruction into "the book (object unit)" and "Put on the table (action-location unit)". By maintaining the dialogue state (such as recording information about previously mentioned objects, locations, actions, etc.), some omitted or referential content can be supplemented and accurately split when splitting subsequent instructions, avoiding splitting errors caused by the lack of context.
[0067] Consider the scenario background knowledge. Assist in splitting according to the specific application scenario and relevant background knowledge where the agent is located. For example, in an intelligent factory environment, if it is known that there are various types of robots and production equipment in the scenario, when receiving the instruction "Let the assembly robot install the parts on the conveyor belt", the agent can use the scenario background knowledge (such as the functions of the assembly robot, the storage location of the parts, the location of the conveyor belt, etc.) to accurately split into "assembly robot (executor unit)", "install the parts (action-object unit)", and "on the conveyor belt (action-location unit)", and be able to understand the logical relationship between these splitting units, so as to generate the correct control instruction.
[0068] Application of deep semantic understanding technology, using word embedding models. Word embedding models (such as Word2Vec, GloVe, etc.) can map words into a low-dimensional semantic space, making words with similar semantics closer in the space. In this way, the agent can better understand the semantic similarity and relevance of words in natural language. For example, when words like "turn on", "open", and "start" appear in the instruction, the word embedding model can recognize their semantic similarity, all meaning to make a certain device or function start working, thus helping to accurately split the text fragments containing these related words into action units of the same category, avoiding inaccurate splitting caused by different expression forms of words.
[0069] Combined with semantic role labeling. Semantic role labeling (SRL) aims to identify the semantic roles of each component in a sentence, such as the agent, patient, time, location, etc. For the natural language instruction "Please ask the cleaning robot to clean the floor of the living room", semantic role labeling can determine that the "cleaning robot" is the agent, "clean" is the action, and "the floor of the living room" is the patient (the object of the action), and can further clarify the time (if mentioned in the instruction) and location (living room) where the action occurs. Using the results of semantic role labeling, each unit in the instruction can be split more accurately, and the role played by each unit at the semantic level can be clarified, providing an accurate semantic information basis for subsequent function mapping and control instruction generation, thereby improving the accuracy of splitting units.
[0070] Provide a user feedback channel. For example, set a feedback button or an opinion collection box on the user interface of the agent control system. When the user finds that the execution result of the agent for the natural language instruction does not meet expectations, they can conveniently submit feedback, explaining the original meaning of the instruction and the deviation of the agent's execution. At the same time, the system can also automatically record information such as the user's evaluation of the execution result (such as through a simple satisfied / dissatisfied button) and the user's re-corrected input of the instruction (if the user modifies the instruction expression again), in order to better analyze the accuracy problem of splitting units.
[0071] Analyze user feedback in a timely manner and update the system. Classify, count, and deeply analyze the collected user feedback to determine which splitting unit problems have led to execution deviations in the agent control instructions. For example, the user feedback is "I asked the robot to get a cup, but it got the wrong position. Maybe it didn't correctly understand the position description 'the cup on the left'". For such feedback, analyze whether the splitting unit accurately splits out the position-object composite unit 'the cup on the left'. If it is a problem in the splitting link, relevant splitting rules, models, or algorithms need to be adjusted and optimized to improve the splitting accuracy of subsequent similar instructions. By continuously paying attention to and utilizing user feedback, a closed-loop mechanism for continuously optimizing the accuracy of splitting units is formed, enabling the system to better adapt to the actual needs and natural language expression habits of users.
[0072] In summary, through a variety of measures such as comprehensively applying advanced natural language processing technologies, constructing a high-quality data and rule foundation, introducing context and semantic understanding, and establishing an artificial review and user feedback mechanism, this embodiment can effectively ensure the accuracy of natural language splitting units, thereby laying a solid foundation for the accuracy and reliability of the agent control instruction generation method.
[0073] In a specific application scenario, the ability logic expression (logic_form) of this embodiment proposes a mechanism for mapping natural language action instructions to a context-free semantic space. This mechanism constructs a composable, inferable, and plannable intermediate representation layer as the "abstract semantic protocol layer" connecting the natural language interface and the execution control logic. In this way, on the one hand, the ability logic expression realizes the decoupling from the internal logic of the agent, and on the other hand, it also realizes the support for modeling complex actions.
[0074] The ability logic expression (also known as logic form) is a representation form of natural language instructions.
[0075] Specifically, in natural language, although there are various sentence patterns, instructions are usually expressed as an imperative sentence, which necessarily contains a predicate verb, and there may also be an object and other attributives and auxiliary verbs that modify this action; for machine language, the execution logic of actions or operations is usually implemented by a function, and various input information is represented by various parameters of the function. Correspondingly, we can establish a correspondence between machine language instructions and natural language instructions by referring to the above logic. Specifically, refer to Table 1: Table 1 Correspondence Table of Machine Language Instructions and Natural Language Instructions
[0076] According to the above logic, a series of actions and operations that humans can perform are defined as a series of functions, specifically including action functions, calculation functions, and special function functions, and each function has corresponding parameters. Next, a specific introduction is given as follows: 1. Action functions: Include various actions that the machine can perform in the embodied environment, including actions that can be executed without parameters (such as standing up, handstanding, etc.) and actions that require parameters when executed (such as picking up something, walking, etc.), as shown in Table 2: Table 2 Example Table of Action Functions
[0077] 2. Calculation functions: It represents various operations on various numerical values, including numerical operations (such as addition, subtraction, multiplication, division, greater than, less than, etc.) and set operations (such as belongs to, contains, counting, etc.), as shown in Table 3: Table 3 Example Table of Calculation Functions
[0078] 3. Special processing functions: Such functions are usually for supporting some special processing (such as obtaining corresponding entities based on text) and some implicit actions (such as querying information about a certain entity or event), as shown in Table 4: Table 4 Example Table of Special Processing Functions
[0079] According to the above representation method, it can be obtained that: as long as the corresponding functions are defined, all action instructions and operation logics can be represented as a series of function representations. And for the convenience of parsing, add symbols to the function names (such as rewriting the grounding function as " "), and add symbols to the parameters (such as rewriting the parameter text as <arg>text< / arg> ). Correspondingly, the representation of natural language instructions, that is, the ability logic expression, can be obtained. As shown in Table 5: Table 5 Example Table of Ability Logic Expressions
[0080] In actual applications, for human natural language instructions, after converting them into ability logic expressions using the model, the functional execution of the instructions can be achieved through the logic of the functions.
[0081] As Figure 2 shown, this embodiment also provides an intelligent agent control instruction generation device, including: a splitting unit, configured to split the obtained natural language according to the set rules to obtain a splitting unit; A function generation unit, configured to generate a function corresponding to each split unit based on the split unit and the trained function model; A construction unit, configured to construct an ability logic expression corresponding to the obtained natural language based on the function corresponding to each split unit; An instruction generation unit, configured to generate an agent control instruction based on the ability logic expression.
[0082] The electronic device disclosed in this embodiment includes a memory and a processor. The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0083] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the electronic device executes all or part of the steps of the agent control instruction generation method of the foregoing embodiments of the present disclosure.
[0084] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, this embodiment may also include well-known structures such as communication buses, interfaces, etc., and these well-known structures should also be included in the protection scope of the present disclosure.
[0085] Such as Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. It shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present disclosure. Figure 3 The shown electronic device is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.
[0086] Such as Figure 3 As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the electronic device are also stored. The processing device, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.
[0087] Typically, the following devices can be connected to the I / O interface: input devices including, for example, sensors or visual information acquisition devices; output devices including, for example, display screens; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. The communication device can allow the electronic device to communicate wirelessly or wirelesly with other devices (such as edge computing devices) to exchange data. Although Figure 3 an electronic device having various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.
[0088] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by the processing device, all or part of the steps of the intelligent agent control instruction generation method according to the embodiments of the present disclosure are executed.
[0089] For a detailed description of this embodiment, reference can be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.
[0090] The computer-readable storage medium disclosed in this embodiment stores non-temporary computer-readable instructions. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the intelligent agent control instruction generation method according to the foregoing embodiments of the present disclosure are executed.
[0091] The above-mentioned computer-readable storage medium includes but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).
[0092] For a detailed description of this embodiment, reference can be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.
[0093] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-disclosed specific details are only for illustrative and facilitating understanding purposes, rather than limitations, and the above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0094] In this disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.
[0095] In addition, as used herein, the "or" used in the listing of items starting with "at least one" indicates a disjunctive listing. So, for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Further, the term "exemplary" does not mean that the examples described are preferred or better than other examples.
[0096] It should also be noted that in the systems and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.
[0097] Various changes, substitutions, and alterations to the technologies described herein can be made without departing from the teachings defined by the appended claims. In addition, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.
[0098] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0099] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some variations, modifications, alterations, additions, and subcombinations thereof.
Claims
1. A method for generating control instructions for an intelligent agent, characterized in that: include: The acquired natural language is split according to the set regulations to obtain split units; Generate a function corresponding to each split unit based on the split unit and the trained function model; The capability logic expression corresponding to the natural language obtained based on the function corresponding to each split unit is constructed; Generate agent control instructions based on capability logic expressions.
2. The method for generating control instructions for an intelligent agent according to claim 1, characterized in that: The functions include action functions, calculation functions and special processing functions.
3. The method for generating intelligent agent control instructions according to claim 2, characterized in that: The action function includes various actions that the agent can perform; The calculation function includes numerical calculation and set calculation.
4. The method for generating control instructions for an intelligent agent according to claim 1, characterized in that: The function includes a function representation and a parameter introduction. The function representation cannot be empty, but the parameter introduction can be empty.
5. The method for generating intelligent agent control instructions according to claim 4, characterized in that: The capability logic expression corresponding to the natural language obtained by constructing the function corresponding to each split unit includes: Add corresponding identifiers to the functions, and rewrite the functions based on the identifiers to obtain the capability logic expression.
6. The method for generating intelligent agent control instructions according to claim 5, characterized in that: The adding of corresponding identifiers to the functions includes: Add the function name in the function representation Symbols, added to the parameters in the function representation symbol.
7. An intelligent agent control instruction generating device, characterized in that: include: A splitting unit is used to split the acquired natural language according to a set rule to obtain splitting units; A function generating unit, used for generating a function corresponding to each split unit based on the split unit and the trained function model; A construction unit, used to construct a capability logic expression corresponding to the acquired natural language based on the function corresponding to each split unit; The instruction generation unit is used to generate agent control instructions based on capability logic expressions.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the intelligent body control instruction generation method described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the intelligent agent control instruction generation method described in any one of claims 1-6.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the method for generating intelligent agent control instructions as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Automatic function generation method and system based on truth table
CN115329949A
Function calling method based on natural language
CN118642787A
Cited By
Multi-agent-based industrial process control system and method, agents and medium
CN121209434A