Value-driven agent capability subset dynamic recommendation method and device
By obtaining the agent's current U instance function and V value function, and combining the Infostate object, U instance functions related to the current V value function are recommended, which solves the problem of too large space for agent planning and reasoning, and significantly improves the efficiency and accuracy of planning and reasoning.
Patent Information
- Application Number
- CN202510624071.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
In the prior art, the planning and inference space of agents are too large, resulting in a surge in computing complexity and too large search space, and lack of a general and user-insensitive filtering mechanism, resulting in inefficient planning.
By obtaining the current U instance function and V value function of the agent, and combining the Infostate object, U instance functions related to the current V value function are recommended. Based on the Python version of TongPL, the agent's ability search space is reduced.
It significantly improves the efficiency and accuracy of planning reasoning, ensures that the recommended capability examples are highly consistent with the current needs and goals of the agent, and enhances the agent's adaptability.
Smart Images

Figure CN120146201A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method, device, electronic device, computer-readable storage medium, and computer program product for dynamically recommending a subset of intelligent agent capabilities based on value driving. Background Art
[0002] In the field of artificial intelligence, the planning and reasoning capabilities of an intelligent agent are important indicators to measure its intelligence level. However, with the complexity of application scenarios, the value function, capabilities, or action space of an intelligent agent often becomes extremely large, which poses a huge challenge to planning and reasoning. To solve this problem, the original TongPL (Lua version) proposed a method for recommending a specific capability space (U) based on a value function (V), and by reducing the U space, the speed and accuracy of planning and reasoning are improved.
[0003] In the original TongPL (Lua version) solution, the UV library is defined based on the SPO (subject-predicate-object triple) of OWL (Web Ontology Language). The V-to-U algorithm calculates the similarity between UV by comparing SPO strings, thereby achieving TopN U recommendations. However, this solution has a strong dependence on the definition of functions, especially the OWL-based SPO structure, which is difficult to use on a large scale in practical applications.
[0004] At the same time, some other solutions such as watchpoints, although they can track the situation where variables change, ignore the situation of variable usage (that is, there is no change, just reference). This results in these solutions being unable to provide sufficient information for relevance calculation, thereby limiting their application effects in intelligent planning systems.
[0005] To solve the problem of the overly large planning and reasoning space of complex intelligent agents, in fact, in most scenarios, a considerable amount of unnecessary information can be filtered out before planning and reasoning. For example, a filtering mechanism based on common sense and context can significantly improve the planning efficiency and accuracy.
[0006] Therefore, how to design a general, open, and user-invisible filtering mechanism to solve the problem of the overly large planning and reasoning space of complex intelligent agents has become an urgent technical problem to be solved. Summary of the Invention
[0007] In view of this, the present invention provides a method, device, electronic device, computer-readable storage medium, and computer program product for dynamically recommending a subset of intelligent agent capabilities based on value driving, so as to solve the problem of the overly large planning and reasoning space of complex intelligent agents in the prior art.
[0008] To solve the above technical problems, an embodiment of the present invention provides a method for dynamically recommending a subset of agent capabilities based on value drive, and the method includes: Obtain all current U instance functions of the agent and the current V value function of the agent; wherein, the current U instance function and the current V value function are the agent capabilities U and the value function V written based on the Python version of TongPL; Obtain the current Infostate object; wherein, the Infostate object is a data representation of the information state, and is used to store and transmit the information required by the agent during the decision-making process; Based on the above current U instance function, current V value function, and current Infostate object, recommend U instance functions related to the current V value function.
[0009] Optionally, the method further includes: The general name of different value dimensions of the agent is denoted as the value function V, where the value function V includes n dimensions, and each dimension of the value function V has a corresponding calculation function denoted as f_i and a weight denoted as w_i. The total calculation function V is f_1 * w_1 + f_2 * w_2 +... + f_n * w_n, and the calculation result is used as the value function of the overall V. The improvement of the value function of the overall V determines the driving force of the agent's actions.
[0010] Optionally, recommending U instance functions related to the current V value function based on the above current U instance function, current V value function, and current Infostate object includes: Based on trace+ast, run the value function of the overall V and each U instance function to obtain the usage attribute counter and change attribute counter of the value function of the overall V and each U instance function with respect to the Infostate object; Construct a UV graph based on the coincidence degree of the usage attribute names of the usage attribute counter and the change attribute names of the change attribute counter; Calculate the total in-edge weight sum of the U instance function nodes connected to the value function node of the overall V according to the UV graph to obtain a correlation calculation result; Recommend the relevant U instance functions according to the correlation technical result.
[0011] Optionally, running the value function of the overall V and each U instance function based on trace+ast to obtain the usage attribute counter and change attribute counter of the value function of the overall V and each U instance function with respect to the Infostate object includes: During the function call process, establish a node class StackNode for each call stack; among them, each node class StackNode of the call stack includes at least the function definition node of the ast, all ast nodes indexed by line number, and the parsed Infostate variable name; and establish a data structure StackInfo for each call stack; among them, each data structure StackInfo of the call stack includes at least the mapping relationship between the Infostate variable id and the Infostate variable name, the mapping relationship between the local Infostate variable and the Infostate variable name, the flag indicating that the Infostate variable has changed, the usage attribute counter, and the change attribute counter. Use the ast and inspect modules of Python to parse the source code of the UV function to obtain an ast-based tree structure. Use the tree structure to recursively count the usage and change of Infostate variables in a single call stack. Use the trace module to trace the usage and change of Infostate variables in multiple call stacks to obtain the total usage attribute counter and change attribute counter.
[0012] Optionally, constructing a UV graph based on the coincidence degree of the usage attribute name of the usage attribute counter and the change attribute name of the change attribute counter includes: The value function of the overall V and each U instance function have a total usage attribute counter and a change attribute counter for the Infostate object. The value function of the overall V and each U instance function are used as nodes in the graph; When there is usage or change for the same Infostate object attribute between two nodes, the two nodes are related, and a directed edge will be established between the two nodes, and the edge weight is the sum of the counts of all coincident Infostate object attributes, thereby constructing the UV graph.
[0013] Optionally, calculating the sum of the in-edge weights of the U instance function nodes connected to the value function node of the overall V according to the UV graph to obtain a correlation calculation result, and recommending the relevant U instance functions according to the correlation technology result includes: Calculate the sum of the in-edge weights of each U instance function node connected to the value function node of the overall V according to the UV graph as the correlation score of the U instance function node; Sort the correlation scores of all U instance function nodes and recommend the top N relevant U instance functions.
[0014] Optionally, the method further includes: Maintain one or more U instance functions; wherein, each of the U instance functions corresponds to a specific ability of the agent; When a U instance function is configured to modify one or more state parameters in the Infostate object during execution.
[0015] On the other hand, an optional embodiment of the present invention provides a device for dynamically recommending a subset of agent capabilities based on value drive, the device comprising: A first acquisition module, configured to acquire all current U instance functions of the agent and the current V value function of the agent; wherein, the current U instance functions and the current V value function are the agent capabilities U and value function V written based on the Python version of TongPL; A second acquisition module, configured to acquire the current Infostate object; wherein the Infostate object is a data representation of the information state, and is used to store and transmit information required by the agent during the decision-making process; A recommendation module, configured to recommend U instance functions related to the current V value function based on the above-mentioned current U instance functions, current V value function, and current Infostate object.
[0016] Optionally, the device further comprises: A calculation module, configured to collectively denote different value dimensions of the agent as the value function V, wherein the value function V includes n dimensions, each dimension of the value function V has a corresponding calculation function denoted as f_i and a weight denoted as w_i, and the overall calculation function V is f_1 * w_1 + f_2 * w_2 +... + f_n * w_n, and the calculation result is used as the value function of the overall V. The improvement of the value function of the overall V determines the driving force of the agent's actions.
[0017] Optionally, the recommendation module comprises: A first acquisition unit, configured to run the overall V value function and each U instance function based on trace+ast, and obtain the usage attribute counter and change attribute counter of the overall V value function and each U instance function with respect to the Infostate object; A construction unit, configured to construct a UV graph based on the coincidence degree of the usage attribute names of the usage attribute counter and the change attribute names of the change attribute counter; A second acquisition unit, configured to calculate the sum of the in-edge weights of the U instance function nodes connected to the overall V value function node according to the UV graph, and obtain a correlation calculation result; A recommendation unit, configured to recommend the relevant U instance functions according to the correlation technology result.
[0018] Optionally, the first acquisition unit includes: A creation subunit, configured to create a node class StackNode for each call stack during a function call process; wherein, the node class StackNode of each call stack includes at least a function definition node of the ast, all ast nodes indexed by line number of the code, and a parsed Infostate variable name; and create a data structure StackInfo for each call stack; wherein, the data structure StackInfo of each call stack includes at least a mapping relationship between an Infostate variable id and an Infostate variable name, a mapping relationship between a local Infostate variable and an Infostate variable name, a flag indicating that the Infostate variable has changed, a usage attribute counter, and a change attribute counter; A first acquisition subunit, configured to parse the source code of the UV function using the ast and inspect modules of Python to obtain an ast-based tree structure; A statistics subunit, configured to recursively count the usage and change situations of Infostate variables in a single call stack using the tree structure; A second acquisition subunit, configured to trace the usage and change situations of Infostate variables in multiple call stacks using the trace module to obtain a total usage attribute counter and a change attribute counter.
[0019] Optionally, the construction unit includes: A third acquisition subunit, configured to have a total usage attribute counter and a change attribute counter for the value function of the overall V and each U instance function with respect to an Infostate object, and use the value function of the overall V and each U instance function as nodes in the graph; A construction subunit, configured to when there is usage or change for the same Infostate object attribute between two nodes, the two nodes are related, and a directed edge is established between the two nodes, and the edge weight is the sum of the counts of all overlapping Infostate object attributes, thereby constructing the UV graph.
[0020] Optionally, the recommendation unit is further configured to calculate the sum of the edge weights of the incoming edges of each U instance function node connected to the value function node of the overall V according to the UV graph as the relevance score of the U instance function node; sort the relevance scores of all U instance function nodes, and recommend the top N relevant U instance functions.
[0021] Optionally, the apparatus further includes: A maintenance module, configured to maintain one or more U instance functions; wherein, each U instance function corresponds to a specific ability of the intelligent agent; A modification module for modifying one or more state parameters in an Infostate object when a U instance function is configured to do so during execution.
[0022] In yet another aspect of an alternative embodiment of the present invention, there is provided an electronic device, comprising: a processor; and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of the above embodiments.
[0023] In yet another aspect of an alternative embodiment of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to any one of the above embodiments.
[0024] In yet another aspect of an alternative embodiment of the present invention, there is provided a computer program product, the computer program product including instructions that, when executed, cause a computer to execute the method according to any one of the above embodiments.
[0025] Through an alternative embodiment of the present invention, all current U instance functions of the agent and the current V value function of the agent are obtained; wherein, the current U instance functions and the current V value function are the agent capabilities U and value function V written based on the Python version of TongPL; the current Infostate object is obtained; wherein, the Infostate object is a data representation of the information state, used to store and transmit the information required by the agent during the decision-making process; based on the above current U instance functions, current V value function, and current Infostate object, a U instance function related to the current V value function is recommended. It solves the problems in the prior art that as the complexity of the agent increases, the explosion of the value function V and the explosive growth of the ability set U lead to a sharp increase in computational complexity and an overly large search space, and the existing screening mechanism based on string similarity is inefficient when facing high-complexity V and U, lacking a general and user-invisible filtering mechanism, resulting in low planning efficiency and difficulty in meeting the actual application requirements. The embodiment of the present invention can accurately recommend a U instance function related to the current value function of the agent by obtaining the current U instance functions and V value function of the agent and combining the current Infostate object. This recommendation mechanism is based on the Python version of TongPL, which can effectively narrow the ability search space of the agent, thereby significantly improving the efficiency of planning and reasoning. At the same time, since the recommendation result is closely related to the current value function of the agent, it can ensure that the recommended ability instances are highly consistent with the current needs and goals of the agent, thereby improving the accuracy and effectiveness of planning decisions. In addition, by using the Infostate object to store and transmit information during the decision-making process, the alternative instance of the present invention can also enhance the adaptive ability of the agent in different scenarios, enabling it to more flexibly handle various complex situations.
[0026] It should be understood that the above general description and the following specific embodiments are only exemplary and explanatory, and they do not limit the scope of what the present invention intends to claim. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The following attached drawings are a part of the specification of the present invention, which illustrate exemplary embodiments of the present invention. The attached drawings and the description of the specification are used together to explain the principles of the present invention.
[0028] Figure 1 It is a flowchart of a value-driven intelligent agent ability subset dynamic recommendation method according to an embodiment of the present invention.
[0029] Figure 2 It is a schematic diagram of a value-driven intelligent agent ability subset dynamic recommendation method according to an embodiment of the present invention.
[0030] Figure 3It is another flowchart of the value-driven intelligent agent capability subset dynamic recommendation method according to an embodiment of the present invention.
[0031] Figure 4 It is a structural block diagram of a value-driven intelligent agent capability subset dynamic recommendation device according to an embodiment of the present invention.
[0032] Figure 5 It is another structural block diagram of a value-driven intelligent agent capability subset dynamic recommendation device according to an embodiment of the present invention.
[0033] Figure 6 It is a structural block diagram of a recommendation module according to an embodiment of the present invention.
[0034] Figure 7 It is a structural block diagram of a first acquisition unit according to an embodiment of the present invention.
[0035] Figure 8 It is a structural block diagram of a construction unit according to an embodiment of the present invention.
[0036] Figure 9 It is another structural block diagram of a value-driven intelligent agent capability subset dynamic recommendation device according to an embodiment of the present invention.
[0037] Figure 10 It shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. Detailed implementation manners
[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the spirit of the content disclosed by the present invention will be clearly described below with reference to the drawings and in detail. After any person skilled in the art understands the embodiments of the content of the present invention, they can make changes and modifications based on the technologies taught by the content of the present invention, which do not deviate from the spirit and scope of the content of the present invention.
[0039] The illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention. In addition, elements / components using the same or similar reference numerals in the drawings and embodiments are used to represent the same or similar parts.
[0040] Regarding the "first", "second",... used in this article, they do not particularly refer to the order or sequence, nor are they used to limit the present invention. They are only used to distinguish elements or operations described with the same technical terms.
[0041] Regarding the directional terms used in this article, such as: up, down, left, right, front, or back, etc., they are only references to the directions in the drawings. Therefore, the directional terms used are for explanation and not for limiting this creation.
[0042] As used herein, terms such as "comprising", "including", "having", "containing", etc. are all open-ended terms, meaning including but not limited to.
[0043] As used herein, "and / or" includes any one or all combinations of the recited things.
[0044] As used herein, "a plurality of" includes "two" and "more than two"; "a plurality of groups" includes "two groups" and "more than two groups".
[0045] As used herein, terms such as "substantially", "about", etc. are used to modify any quantity or error that can vary slightly, but these slight variations or errors do not change its essence. Generally, the range of such slight variations or errors modified by such terms can be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments, or other values. Those skilled in the art should understand that the aforementioned values can be adjusted according to actual needs and are not limited thereto.
[0046] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted to have a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0047] In cases where expressions such as "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In cases where expressions such as "at least one of A, B, or C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but not be limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). Those skilled in the art should also understand that substantially any disjunctive conjunction and / or phrase representing two or more alternative items, whether in the specification, claims, or drawings, should be understood to give the possibility of including one of these items, either of these items, or both items. For example, the phrase "A or B" should be understood to include the possibility of "A" or "B", or "A and B".
[0048] To solve the problem of excessive planning and reasoning space for complex agents (e.g., the different value dimensions and large ability sets of agents), in fact, in most scenarios, a considerable amount of unnecessary information can be filtered before planning and reasoning (e.g., the action of eating when hungry, cleaning when the room is messy). Among them, the general term for the different value dimensions of the agent is denoted as V, including those defined manually or learned from data and experience. The ability set of the agent is denoted as U, including all the things the agent can do, which can be relatively simple actions or more complex behaviors. For example, walking, talking, eating, sleeping, picking up objects, putting down objects, cleaning the room, etc. Abilities can be innate knowledge given to the agent manually by writing code, or new knowledge generated by the agent through interaction with the environment or communication with other agents. As the agent's abilities become stronger, this set will also become larger. When making planning decisions, the agent needs to input its current value V and ability set U, and consider which abilities (i.e., what things to do) can maximize its own value score. When the dimension of V and the number of U are large, the computational cost required for planning decisions will become very large, and a method is needed to improve the planning efficiency. To improve the planning efficiency and accuracy, combined with the design principles of TongPL (Python version), and in the principle of being as general and open as possible, there needs to be a user-invisible way to filter the overly large ability set.
[0049] Based on the value V, implement dynamic U recommendation at runtime to help the agent's planner perform more effective and rapid planning, while reducing the limitations or additional workload when users write the UV library. For this purpose, in this embodiment, a method for dynamically recommending a subset of agent capabilities based on value-driven is provided. Figure 1 It is a flowchart of the method for dynamically recommending a subset of agent capabilities based on value-driven according to an embodiment of the present invention, as Figure 1 shown, and this process includes the following steps: Step S101: Obtain all current U instance functions of the agent and the current V value function of the agent; wherein, the current U instance function and the current V value function are the agent capabilities U and value function V written based on the Python version of TongPL. In an alternative embodiment, the different value dimensions of the agent are collectively denoted as the value function V, including those defined manually or learned from data and experience, such as hunger level, thirst level, cleanliness, tidiness, curiosity, interestingness, fatigue, etc. Among them, the value function V includes n dimensions, and each dimension of the value function V has a corresponding calculation function denoted as f_i and a weight denoted as w_i. The overall calculation function V is f_1 * w_1 + f_2 * w_2 +... + f_n * w_n, and the calculation result is used as the value function of the overall V. The improvement of the value function of the overall V determines the driving force of the agent's actions. The set of agent capabilities is denoted as U, including all the things the agent can do, which can be relatively simple actions or relatively complex behaviors. For example, walking, speaking, eating, sleeping, picking up an item, putting down an item, cleaning a room, etc. The capabilities can be innate knowledge given to the agent manually by writing code or new knowledge generated by the agent through interaction with the environment or communication with other agents. As the agent's capabilities become stronger and stronger, this set will also become larger and larger.
[0050] Step S102: Obtain the current Infostate object; wherein, the Infostate object is a data representation of the information state, used to store and transfer the information required by the agent during the decision-making process.
[0051] Step S103: Recommend U instance functions related to the current V value function based on the above current U instance functions, current V value functions, and current Infostate objects.
[0052] Through the above steps, all current U instance functions of the agent and the current V value function of the agent are obtained; wherein, the current U instance function and the current V value function are the agent capabilities U and value function V written based on the Python version of TongPL; the current Infostate object is obtained; wherein, the Infostate object is a data representation of the information state, used to store and transmit the information required by the agent during the decision-making process; based on the above current U instance function, current V value function, and current Infostate object, a U instance function related to the current V value function is recommended. The problem in the prior art is solved that with the increase in the complexity of the agent, the explosion of the value function V and the explosive growth of the ability set U lead to a sharp increase in computational complexity and an overly large search space, and the existing string similarity-based screening mechanism is inefficient in the face of high-complexity V and U, lacking a general and user-invisible filtering mechanism, resulting in low planning efficiency and difficulty in meeting the actual application requirements. In the embodiments of the present invention, by obtaining the current U instance function and V value function of the agent and combining the current Infostate object, a U instance function related to the current value function of the agent can be accurately recommended. This recommendation mechanism is based on the Python version of TongPL, which can effectively narrow the ability search space of the agent, thereby significantly improving the efficiency of planning and reasoning. At the same time, since the recommendation result is closely related to the current value function of the agent, it can ensure that the recommended ability instances are highly consistent with the current needs and goals of the agent, thereby improving the accuracy and effectiveness of planning decisions. In addition, by using the Infostate object to store and transmit information during the decision-making process, the above steps of the optional instances of the present invention can also enhance the adaptive ability of the agent in different scenarios, enabling it to more flexibly handle various complex situations. In a specific application scenario, it is applied to the attention mechanism of complex agents. For example, through the correlation calculation from the value (V) of the agent to the ability set (U), the most relevant ability subset can be screened according to a specific value. Among them, the value of the agent can be artificially defined or learned from data or experiences, such as values of hunger, thirst, cleanliness, tidiness, curiosity, etc. The ability set includes what the agent can do, such as eating, drinking, cleaning dirty dishes, cleaning the room, exploring unknown areas, etc. If the agent is not currently hungry but the room is messy, then when the agent plans and decides what to do next, it can only select a part of the ability set (that is, without considering eating, only considering the abilities related to cleaning the room), which can reduce the unnecessary occupation of planning and reasoning resources, improve the accuracy and reliability of the planning and reasoning results, and enhance the efficiency of the planning and reasoning process.Meanwhile, the above embodiments have no restrictions on the data structure of the variables themselves. They can handle both simple Python data structures (such as arrays, tuples, dictionaries) and complex class instances or nested structures, such as classes with many member variables, dictionaries nested with other dictionaries or arrays (such as the Infostate object mentioned below), etc.
[0053] In an alternative embodiment, as Figure 2 shown, the input parameters are: all current U instances of the agent, the V calculation function, and the Infostate object; the return parameter is: all U instances associated with V. Among them, the U instance and the V calculation function are the ability and value functions written by the user based on TongPL (Python version). When running, the Agent instance of the agent will maintain a list of U instances and the total V calculation function object. In addition, the Infostate object refers to all environmental and state information, which is maintained in memory in the form of a dictionary, and the data is generated and provided by the perception module (vision, hearing, etc.). The V calculation function will calculate the corresponding score according to the Infostate object (such as whether the environment is clean, whether oneself is hungry, etc.), and the U instance function will operate on the Infostate object during planning. For example, the U for eating will change the self-hunger degree in the Infostate object to not hungry.
[0054] The above step S103 involves recommending U instance functions related to the current V value function based on the current U instance function, the current V value function, and the current Infostate object. In an alternative embodiment, as Figure 3 shown, it includes the following steps: Step S301: Run the overall V value function and each U instance function based on trace + ast to obtain the usage attribute counter and change attribute counter of the overall V value function and each U instance function with respect to the Infostate object. Specifically, the trace module of Python can be used to trace the usage and change of the Infostate object in the overall V and each U function, and all attributes in the Infostate object will be traced and recorded. For example, if Infostate = {'a': {'b':[1, 2]}}, then for the statement temp = Infostate['a']['b'][0], it is a usage of the Infostate object, and the usage attribute name is Infostate['a']['b'][0]; correspondingly, for the statement Infostate['a']['b'] = [4, 5], it is a change of the Infostate object, and the change attribute name is Infostate['a']['b'].
[0055] Step S302: Construct a UV graph based on the coincidence degree of the usage attribute names using the usage attribute counter and the changed attribute names changing the attribute counter. Specifically, the coincidence degree of the usage attribute counter name (string) and the changed attribute counter name (string) is used. The corresponding usage attribute counter contains multiple usage attribute names and their corresponding quantities (key-value pairs), and the changed attribute counter contains multiple changed attribute names and their corresponding quantities (key-value pairs). In this alternative embodiment, relying on the trace + ast to run the value function of the overall V and each U instance function, the usage attribute counter and the changed attribute counter of the overall V and each U with respect to the Infostate object can be accurately obtained. This process not only reveals the specific operations of the U instance and the V calculation function at the code level, but also quantifies their dependence and influence on each attribute in the Infostate object. In this way, the embodiments of the present invention can accurately identify the internal correlation between the U instance and the V calculation function, laying a solid foundation for the subsequent recommendation process. At the same time, using the coincidence degree of the usage attribute name (string) and the changed attribute name (string), combined with the usage attribute counter and the changed attribute counter, a UV association graph is constructed. This graph uses the attribute name as the node and the usage and change of the attribute as the edge, and quantifies the association strength between the nodes through the coincidence degree. The construction of this graph structure transforms the originally complex code-level association relationship into an intuitive graphical representation, making the association between the U instance and the V calculation function clearer and easier to understand. Moreover, the construction of the graph structure also provides an efficient data basis for the subsequent correlation calculation.
[0056] Step S303: Calculate the total in-edge weight sum of the U instance function nodes connected to the value function node of the overall V according to the UV graph to obtain the correlation calculation result. In addition, calculate the correlation score through the total in-edge weight sum of the U nodes connected to the V node, and recommend the TopN U set according to the correlation score. This process ensures that the U instance set recommended to the agent has the highest correlation with the V calculation function, so that the current needs and goals of the agent can be maximally met during the decision-making process. Based on the recommendation mechanism of the correlation score, it not only improves the accuracy and effectiveness of the recommendation, but also avoids the interference of irrelevant or low-correlation U instances, further enhancing the decision-making efficiency and performance of the agent.
[0057] Step S304: Recommend the relevant U instance function according to the correlation technology result.
[0058] Furthermore, regarding the value function of the overall V and each U instance function based on trace + ast, the value function of the overall V and the usage attribute counter and change attribute counter of each U instance function with respect to the Infostate object are obtained. In an optional embodiment, during the function call process, a node class StackNode of each call stack is established; wherein, the node class StackNode of each call stack includes at least the function definition node of the ast, all ast nodes indexed by line number of the code, and the parsed Infostate variable name; and a data structure StackInfo of each call stack is established; wherein, the data structure StackInfo of each call stack includes at least the mapping relationship between the Infostate variable id and the Infostate variable name, the mapping relationship between the local Infostate variable and the Infostate variable name, the flag indicating that the Infostate variable has changed, the usage attribute counter, and the change attribute counter. Since Python coding is relatively flexible, it is necessary to track the usage and change situations of the Infostate object in various cases. For example, some local temporary variables actually refer to the Infostate object, and there are also reference situations between functions. The source code of the UV function is parsed using the ast and inspect modules of Python to obtain the ast-based tree structure. Specifically, the ast (abstract syntax tree, an module of Python) and inspect modules of Python are used to parse the source code to obtain the ast tree structure and convert it into the above data structure. The usage and change situations of the Infostate variable are recursively counted in a single call stack using the tree structure. Specifically, a NodeVisitor class is constructed to be responsible for updating the data information of each call stack. Among them, a recursive method is used to parse all variables in the data structure, and then the Name, Attribute, and Subscript nodes of the ast are used to count the usage and change situations of the variables. The trace module is used to track the usage and change situations of the Infostate variable in multiple call stacks to obtain the total usage attribute counter and change attribute counter.Through this alternative embodiment, during the function call process, the node class StackNode and data structure StackInfo of each call stack are established, and the source code is parsed and variable tracing is performed using Python's ast, inspect, and trace modules, achieving precise statistics and analysis of the usage and changes of the Infostate object in each call stack, with the following remarkable technical effects: 1) Precise capture of function call context: By establishing the node class StackNode for each call stack, including the function definition node of ast, all ast nodes indexed by line number, and the parsed Infostate variable names, the embodiment of the present invention can precisely capture the context information of function calls. This refined context capture provides accurate basic data for subsequent variable tracing and analysis, ensuring that the analysis of the usage and changes of the Infostate object can be accurate to the specific function call level. 2) Comprehensive recording of the mapping relationship and changes of Infostate variables: By establishing the data structure StackInfo of each call stack, including the mapping relationship between the Infostate variable id and name, the mapping relationship between local Infostate variables and names, the flag indicating that the Infostate variable has changed, the usage attribute counter, and the change attribute counter, the embodiment of the present invention can comprehensively record the mapping relationship and changes of Infostate variables in each call stack. This comprehensive recording mechanism not only covers the mapping of global and local variables, but also quantifies the usage and change degree of variables through the change flag and usage / change counters, providing rich data support for subsequent analysis. 3) Effective tracking of the flexibility of Python coding: In response to the flexibility of Python coding, the embodiment of the present invention ensures accurate tracking of the usage and changes of the Infostate object in various complex coding situations by tracking the usage and changes of the Infostate object in each case, such as the reference of local temporary variables to the Infostate object and the reference between functions. This effective tracking mechanism overcomes the challenges brought by the flexibility of Python coding and ensures the accuracy and integrity of the analysis. 4) Efficient parsing of source code and conversion into data structure: By using Python's ast and inspect modules to parse the source code, obtaining the tree structure based on ast functions and converting it into the above data structure, the embodiment of the present invention realizes the efficient conversion of source code into data structure. This conversion process not only retains the structural information of the source code, but also converts it into a data form convenient for subsequent analysis and processing, improving the efficiency and accuracy of processing.5) Accurately count the usage and changes of Infostate variables: By constructing the NodeVisitor class and recursively counting the usage and changes of Infostate variables in a single call stack, the embodiments of the present invention achieve accurate counting of the usage and changes of Infostate variables. Specifically, by recursively parsing all variables in the data structure and using the Name, Attribute, and Subscript nodes of the ast to count the usage and changes of variables, the comprehensiveness and accuracy of the counting are ensured. 6) Comprehensively track the usage and changes of Infostate variables in multiple call stacks: Using the trace module to track the usage and changes of Infostate variables in multiple call stacks, obtaining the total usage attribute counter and change attribute counter, the embodiments of the present invention achieve comprehensive tracking of the overall usage and changes of Infostate variables in multiple call stacks. This comprehensive tracking mechanism provides a global perspective for subsequent analysis and decision-making, ensuring the integrity and accuracy of the analysis. 7) There is no need to impose any restrictions or changes on the traced variables, and there is no need to modify the code at runtime. Since ast is static code parsing, it is easier to control the processing methods and corner cases in different situations. In summary, the embodiments of the present invention can track the usage and changes of any Python function with respect to any Python variable by establishing a fine-grained call stack data structure and using the ast, inspect, and trace modules of Python for source code parsing and variable tracking, and return them in the form of counters, achieving accurate statistics and analysis of the usage and changes of Infostate objects in each call stack. The realization of these technical effects enables the agent to more accurately and comprehensively understand its own state and the interaction with the environment, thereby making more reasonable and efficient decisions, which has important practical application value and broad application prospects.
[0059] Regarding constructing a UV graph based on the coincidence degree of the used attribute name using an attribute counter and the changed attribute name changing the attribute counter, in an alternative embodiment, the value function of the overall V and each U instance function both have a total used attribute counter and a changed attribute counter for the Infostate object. The value function of the overall V and each U instance function are used as nodes in the graph. When there is a use or change for the same Infostate object attribute between two nodes, the two nodes are related, and a directed edge is established between the two nodes. The edge weight is the sum of the counts of all coincident Infostate object attributes. The direction of the edge is defined as from the use count to the change count. Use represents the consumption node, and change represents the production node. Generally, the overall V is a pure consumption node, and U is a production node. We are more concerned about the relevance from V to U. For example, the overall V and U1 have use and change counts (counts greater than 0) for Infostate['x'] and Infostate['y']. Therefore, an edge needs to be established between the overall V and U1. Additionally, the overall V has a use count greater than 0, and U1 has a change count greater than 0. Therefore, the direction of the edge is from the overall V to U1. Edges will be established between the overall V and U, and between U and U according to the above method. Edges also need to be established between U and U because some Us do not directly affect V but are indirectly associated with V through other Us. Finally, a directed graph between U and V is obtained.
[0060] Regarding calculating the total sum of the in-edge weights of the U instance function nodes connected to the value function node of the overall V according to this UV graph to obtain a relevance calculation result, and recommending the relevant U instance functions according to this relevance technical result. In an alternative embodiment, according to this UV graph, calculate the total sum of the in-edge weights of each U instance function node connected to the value function node of the overall V as the relevance score of this U instance function node. Sort the relevance scores of all U instance function nodes, and recommend the top N relevant U instance functions. Specifically, based on the UV directed graph obtained from the above embodiment, calculate the total sum of the in-edge weights of each U node connected to the overall V (for example, if U1 has two in-edges, then the score of U1 is the sum of the weights of these two edges) as the relevance score of this U node. Sort the relevance scores of all U nodes, and the top N U instances can be recommended for efficient reasoning in the downstream planning decision-making.
[0061] In an alternative embodiment, maintain one or more U instance functions; wherein, each U instance function corresponds to a specific ability of the agent, and a U instance function is configured to modify one or more state parameters in the Infostate object during execution.
[0062] A method for variable tracking and correlation calculation based on Python provided by an optional embodiment of the present invention utilizes the Abstract Syntax Tree (AST) of Python to understand and record the reading and writing situations of any Python variable. Using the above capabilities and the Trace module of Python, it realizes the tracking of the usage of a variable in any Python function, and can further establish a correlation calculation graph based on a variable between different functions.
[0063] In this embodiment, a device for dynamically recommending subsets of agent capabilities based on value-driven is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0064] This embodiment provides a device for dynamically recommending subsets of agent capabilities based on value-driven, as Figure 4 shown, including: A first acquisition module 41, configured to acquire all current U instance functions of the agent and the current V value function of the agent; wherein, the current U instance function and the current V value function are the agent capabilities U and value function V written based on the Python version of TongPL. A second acquisition module 42, configured to acquire the current Infostate object; wherein, the Infostate object is a data representation of the information state, used to store and transmit the information required by the agent during the decision-making process. A recommendation module 43, configured to recommend U instance functions related to the current V value function based on the above current U instance function, current V value function, and current Infostate object.
[0065] Optionally, as Figure 5 shown, the device further includes: A calculation module 45, configured to generally denote different value dimensions of the agent as the value function V, wherein the value function V includes n dimensions, each dimension of the value function V has a corresponding calculation function denoted as f_i and a weight denoted as w_i, and the total calculation function V is f_1 * w_1 + f_2 * w_2 +... + f_n * w_n, and the calculation result is used as the value function of the overall V. The improvement of the value function of the overall V determines the driving force of the agent's actions.
[0066] Optionally, as Figure 6 shown, the recommendation module 43 includes: The first acquisition unit 431 is configured to obtain the value function of the overall V and the usage attribute counter and change attribute counter of each U instance function with respect to the Infostate object based on the value functions of the overall V and each U instance function running with trace + ast; The construction unit 432 is configured to construct a UV graph based on the coincidence degree of the usage attribute name of the usage attribute counter and the change attribute name of the change attribute counter; The second acquisition unit 433 is configured to calculate the total weight sum of the incoming edges of the U instance function nodes connected to the value function node of the overall V according to the UV graph, and obtain a correlation calculation result; The recommendation unit 434 is configured to recommend the relevant U instance functions according to the correlation technology result.
[0067] Optionally, as Figure 7 shown, the first acquisition unit 431 includes: The establishment subunit 4311 is configured to establish a node class StackNode for each call stack during the function call process; wherein, the node class StackNode of each call stack includes at least the function definition node of the ast, all ast nodes indexed by line number of the code, and the parsed Infostate variable name; and establish a data structure StackInfo for each call stack; wherein, the data structure StackInfo of each call stack includes at least the mapping relationship between the Infostate variable id and the Infostate variable name, the mapping relationship between the local Infostate variable and the Infostate variable name, the flag indicating that the Infostate variable has changed, the usage attribute counter, and the change attribute counter; The first acquisition subunit 4312 is configured to parse the source code of the UV function using the ast and inspect modules of Python to obtain a tree structure based on the ast; The statistics subunit 4313 is configured to recursively count the usage and change of the Infostate variable in a single call stack using the tree structure; The second acquisition subunit 4314 is configured to use the trace module to trace the usage and change of the Infostate variable in multiple call stacks to obtain the total usage attribute counter and change attribute counter.
[0068] Optionally, as Figure 8 shown, the construction unit 432 includes: The third acquisition subunit 4321 is configured to have a total usage attribute counter and change attribute counter for the value function of the overall V and each U instance function with respect to the Infostate object, and use the value function of the overall V and each U instance function as nodes in the graph; A construction subunit 4322 is used to establish a directed edge between two nodes with a weight equal to the total count of all overlapping Infostate object attributes when there is usage or change of the same Infostate object attribute between the two nodes, and then construct the UV graph.
[0069] Optionally, the recommendation unit 434 is further configured to calculate the total weight of the incoming edges of each U instance function node connected to the value function node of the overall V as the correlation score of the U instance function node based on the UV graph; sort the correlation scores of all U instance function nodes, and recommend the top N relevant U instance functions.
[0070] Optionally, as Figure 9 shown, the device further includes: A maintenance module 46 for maintaining one or more U instance functions; wherein each U instance function corresponds to a specific ability of the agent. A modification module 47 for modifying one or more state parameters in the Infostate object when a U instance function is configured to execute.
[0071] The further functional descriptions of the above-mentioned modules are the same as those in the corresponding embodiments above, and will not be elaborated here.
[0072] An exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiment of the present invention.
[0073] An exemplary embodiment of the present invention further provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.
[0074] An exemplary embodiment of the present invention further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.
[0075] Reference Figure 10, the structural block diagram of an electronic device that can be used as a server or a client of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0076] As Figure 10 shown, the electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0077] A plurality of components in the electronic device 1000 are connected to the I / O interface 1005, including: an input unit 1006, an output unit 1007, a storage unit 1008, and a communication unit 1009. The input unit 1006 can be any type of device that can input information into the electronic device 1000. The input unit 1006 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 1007 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1008 can include but is not limited to magnetic disks, optical disks. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0078] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above. For example, in some embodiments, the music data processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. In some embodiments, the computing unit 1001 can be configured to execute the method according to the embodiments of the present invention by any other suitable means (e.g., by means of firmware).
[0079] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0080] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0081] As used in this invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0082] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0083] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or in a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0084] A computer system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0085] Obviously, the above embodiments are only examples for clear illustration and not limitations on the implementation. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A value-driven method for dynamically recommending agent capability subsets, characterized in that: The method comprises: Obtain all current U instance functions and current V value functions of the agent; wherein the current U instance function and the current V value function are the agent capability U and value function V written based on the Python version of TongPL; Get the current Infostate object; the Infostate object is a data representation of the information state, which is used to store and transmit the information needed by the agent in the decision-making process; Based on the current U instance function, the current V value function and the current Infostate object, a U instance function related to the current V value function is recommended.
2. The value-driven agent capability subset dynamic recommendation method according to claim 1 is characterized in that: The method further comprises: The different value dimensions of the agent are collectively referred to as the value function V, where the value function V includes n dimensions. Each dimension of the value function V has a corresponding calculation function denoted by fi and a weight denoted by wi. The total calculation function V is f_1 * w_1 +f_2 * w_2 + ... + f_n * w_n. The calculation result is taken as the value function of the overall V. The improvement of the value function of the overall V determines the driving force of the agent's actions.
3. The value-driven agent capability subset dynamic recommendation method according to claim 2 is characterized in that: Recommending a U instance function related to the current V value function based on the current U instance function, the current V value function and the current Infostate object includes: Run the value function of the overall V and each U instance function based on trace+ast to obtain the use attribute counter and change attribute counter of the overall V value function and each U instance function about the Infostate object; Constructing a UV map based on the coincidence of the used attribute name of the used attribute counter and the changed attribute name of the changed attribute counter; Calculate the sum of the incoming edge weights of the U instance function nodes connected to the value function node of the overall V according to the UV graph to obtain a correlation calculation result; The relevant U instance function is recommended based on the correlation technical results.
4. The value-driven agent capability subset dynamic recommendation method according to claim 3 is characterized in that: Based on trace+ast, run the value function of the overall V and each U instance function to obtain the use attribute counter and change attribute counter of the Infostate object of the value function of the overall V and each U instance function, including: In the function calling process, a node class StackNode of each call stack is established; wherein, the node class StackNode of each call stack at least includes the function definition node of ast, all ast nodes indexed by code lines, and the parsed Infostate variable name; and a data structure StackInfo of each call stack is established; wherein, the data structure StackInfo of each call stack at least includes the mapping relationship between the Infostate variable id and the Infostate variable name, the mapping relationship between the local Infostate variable and the Infostate variable name, the flag of the Infostate variable change, the use attribute counter, and the change attribute counter; Use Python's ast and inspect modules to parse the source code of the UV function and obtain an ast-based tree structure; Use the tree structure to recursively count the usage and changes of Infostate variables in a single call stack; The trace module is used to track the usage and changes of Infostate variables in multiple call stacks to obtain the total usage attribute counter and change attribute counter.
5. The value-driven agent capability subset dynamic recommendation method according to claim 3 is characterized in that: Constructing a UV map based on the coincidence of the used attribute name of the used attribute counter and the changed attribute name of the changed attribute counter comprises: The overall V value function and each U instance function have a total usage attribute counter and a change attribute counter about the Infostate object. The overall V value function and each U instance function are taken as nodes in the graph; When two nodes use or change the same Infostate object attribute, the two nodes are related, and a directed edge is established between the two nodes. The edge weight is the sum of the counts of all overlapping Infostate object attributes, thereby constructing the UV graph.
6. The value-driven agent capability subset dynamic recommendation method according to claim 5, characterized in that: According to the UV graph, the sum of the input edge weights of the U instance function nodes connected to the value function node of the overall V is calculated to obtain a correlation calculation result. According to the correlation technical result, the relevant U instance function is recommended, including: Calculate the sum of the edge weights of each incoming edge of the U instance function node connected to the value function node of the overall V according to the UV graph as the relevance score of the U instance function node; Sort the relevance scores of all U instance function nodes and recommend the TopN related U instance functions.
7. The value-driven agent capability subset dynamic recommendation method according to any one of claims 1 to 6, characterized in that: The method further comprises: Maintain one or more U instance functions; wherein each U instance function corresponds to a specific capability of the agent; In a U instance function is configured to modify one or more state parameters in the Infostate object when executed.
8. A value-driven intelligent agent capability subset dynamic recommendation device, characterized in that: The device comprises: The first acquisition module is used to obtain all current U instance functions and the current V value function of the agent; wherein the current U instance function and the current V value function are the agent capability U and value function V written based on the Python version of TongPL; The second acquisition module is used to acquire the current Infostate object; wherein the Infostate object is a data representation of the information state, and is used to store and transmit the information required by the intelligent agent in the decision-making process; The recommendation module is used to recommend a U instance function related to the current V value function based on the current U instance function, the current V value function and the current Infostate object.
9. The value-driven agent capability subset dynamic recommendation device according to claim 8, characterized in that: The device also includes: A calculation module is used to record the different value dimensions of the intelligent agent as a value function V, wherein the value function V includes n dimensions, each dimension of the value function V has a corresponding calculation function recorded as fi and a weight recorded as wi, and the total calculation function V is f_1 * w_1 + f_2 * w_2 + ... + f_n * w_n. The calculation result is used as the value function of the overall V. The improvement of the value function of the overall V determines the driving force of the intelligent agent's actions.
10. The value-driven agent capability subset dynamic recommendation device according to claim 9, characterized in that: The recommendation module includes: A first acquisition unit is used to run the value function of the overall V and each U instance function based on trace+ast to obtain the use attribute counter and the change attribute counter of the overall V value function and each U instance function about the Infostate object; A construction unit, configured to construct a UV map based on the coincidence of the usage attribute name of the usage attribute counter and the change attribute name of the change attribute counter; A second acquisition unit is used to calculate the sum of the incoming edge weights of the U instance function nodes connected to the value function node of the overall V according to the UV graph to obtain a correlation calculation result; A recommendation unit is used to recommend the relevant U instance function according to the correlation technical result.
11. The value-driven agent capability subset dynamic recommendation device according to claim 10, characterized in that: The first acquisition unit includes: Establish a subunit, which is used to establish a node class StackNode of each call stack during a function call process; wherein the node class StackNode of each call stack at least includes a function definition node of ast, all ast nodes indexed by code lines, and a parsed Infostate variable name; and establish a data structure StackInfo of each call stack; wherein the data structure StackInfo of each call stack at least includes a mapping relationship between an Infostate variable id and an Infostate variable name, a mapping relationship between a local Infostate variable and an Infostate variable name, a flag indicating that an Infostate variable has changed, a use attribute counter, and a change attribute counter; The first acquisition subunit is used to parse the source code of the UV function using Python's ast and inspect modules to obtain an ast-based tree structure; The statistics subunit is used to recursively count the usage and changes of Infostate variables in a single call stack using a tree structure; The second acquisition subunit is used to use the trace module to track the usage and change of the Infostate variable in multiple call stacks to obtain a total usage attribute counter and a change attribute counter.
12. The value-driven agent capability subset dynamic recommendation device according to claim 10, characterized in that: The building blocks include: The third acquisition subunit, for the overall V value function and each U instance function, has a total usage attribute counter and a change attribute counter about the Infostate object, and takes the overall V value function and each U instance function as nodes in the graph; The construction subunit is used to establish a directed edge between the two nodes when the same Infostate object attribute is used or changed between the two nodes, and the edge weight is the total count of all overlapping Infostate object attributes, thereby constructing the UV graph.
13. The value-driven agent capability subset dynamic recommendation device according to claim 12, characterized in that: The recommendation unit is also used to calculate the sum of the edge weights of the incoming edges of each U instance function node connected to the value function node of the overall V according to the UV graph as the relevance score of the U instance function node; sort the relevance scores of all U instance function nodes, and recommend the TopN related U instance functions.
14. The value-driven agent capability subset dynamic recommendation device according to any one of claims 8 to 13, characterized in that: The device also includes: A maintenance module, used to maintain one or more U instance functions; wherein each U instance function corresponds to a specific capability of the agent; The modification module is used to modify one or more state parameters in the Infostate object when a U instance function is configured to be executed.
15. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1-7.
17. A computer program product, characterized in that The computer program product comprises instructions which, when executed, cause a computer to perform the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Q-learning-based multi-agent initiative recommendation method for agriculture capital electronic commerce
CN103914560A
MOM service code retrieval and recommendation method based on abstract syntax tree, computer program product and terminal
CN118820396A
System assisted data blending
US20160188843A1