Autonomous intelligent system based on value driving
By introducing a value-driven mechanism into the autonomous intelligent system, including value extraction, reverse sampling generation and planning modules, the problem of insufficient intelligence in the existing agents in the communication system is solved, and smarter communication and autonomous behavior is achieved.
Patent Information
- Application Number
- CN202510714540.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing autonomous agents based on large language models cannot achieve smarter communication in communication intelligent systems, and lack the autonomy of internal drive modules to guide autonomy.
A value-driven autonomous intelligent system is proposed, including a value extraction module, a reverse sampling generation module and a planning module. By extracting value from the scene, reverse sampling generates a target state analysis map, and planning and generating action trajectory, the autonomous behavior of the agent is realized.
By driving the behavior of an agent, more intelligent communication can be achieved and the autonomy and communication capabilities of autonomous intelligent systems can be enhanced.
Smart Images

Figure CN120235249A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology for realizing autonomous intelligence, and particularly relates to a value-driven autonomous intelligent system. Background Art
[0002] The related research work on autonomous intelligence can be divided into two stages. The first stage is an intelligent system centered on cognitive architecture, which believes that as long as several cognitive modules are implemented, the autonomy of the system can be achieved. However, there is no stable internal driving module in these modules to tell the system how to propose new tasks. The second stage is an intelligent agent system centered on large language models. The large language model tells the intelligent agent which tasks should be executed in different contexts, so as to achieve autonomous communication. However, the black box and unexplainable characteristics of the large language model itself also make it unclear where the autonomy comes from. The following briefly analyzes the typical work in these two stages.
[0003] Cognitive architecture stage: Since the birth of artificial intelligence, more than 80 cognitive architecture models have been proposed. Most of these models are at the conceptual level, either verified on very simple examples or remain at the thinking level (please refer to "A Review of 40 Years of Cognitive Architecture Research" for details). Well-known cognitive architectures include: ACT-R (Adaptive Control of Thought-Rational) developed by Carnegie Mellon University, which is used to simulate the thinking and behavior processes of humans, including modules such as perception, memory, and decision-making; SOAR is a general problem-solving architecture developed by the University of Michigan, which emphasizes symbolic operation and is suitable for complex task modeling; LIDA (Learning Intelligent Distribution Agent) is based on the global workspace theory in cognitive science and is suitable for multi-task scenarios; CLARION (Connectionist Learning with Adaptive Rule Induction ON-line) combines symbols and connectionism to simulate the conscious and unconscious processes of humans; SPAUN (Semantic Pointer Architecture Unified Network) developed by the University at Buffalo, which simulates the cognitive functions of the human brain and combines neural network structures and biological constraints.
[0004] Large Language Model Stage: Recently, with the rise of large models, cognitive architectures based on large models that include functional modules such as memory, reasoning, and reflection have been successively proposed. Most of them are re-implementations of traditional cognitive architectures. For specific references, please refer to "Cognitive Architectures for Language Agents". A representative work on autonomous agents based on large language models is AI Village launched by Stanford in 2023. Researchers provided each agent with a short biography, including name, age, job, family, interests, and some habits, and then let them play freely. Then, these agents rely on a large language model to generate their behaviors according to the biographies they defined. The way the agents simulate the behaviors of the characters is similar to that of humans. They wake up, make breakfast, go to work, have lunch, and chat with other agents they encounter. They also remember what happened, reflect, and make plans. For example, when the researcher in charge of this landscape suggested that a character plan a Valentine's Day party, she invited friends and acquaintances, and many of them showed up at the right time and place. A batch of similar large language model-driven autonomous agents have emerged after AI Village.
[0005] Currently, autonomous agents based on large language models adopt prompt techniques, and drive the behaviors of the agents by manually writing specific behavioral steps. When applied to communication-based intelligent systems, more intelligent communication cannot be achieved. Summary of the Invention
[0006] One of the objectives of the present invention is to provide a value-driven autonomous intelligent system, which realizes more intelligent communication by driving the behaviors of agents with values.
[0007] A value-driven autonomous intelligent system provided by an embodiment of the present invention includes: an agent, where the agent includes: a value extraction module, an inverse sampling generation module, and a planning module; Among them, the value extraction module extracts values from the scene to obtain predicted values; the inverse sampling generation module is used to inversely sample and generate a target state parsing graph based on the given expected values and predicted values; the planning module is used to plan and generate an action trajectory based on the initial parsing graph and the target state parsing graph.
[0008] Preferably, the value extraction module extracts values from the scene to obtain predicted values and performs the following operations: Parse the scene to generate a large parsing graph; Take values for the values of each region in the scene based on the large parsing graph; Extract the value of the attention region from the large parsing graph as the predicted value.
[0009] Preferably, the reverse sampling generation module is used to generate a target state analysis graph based on the given expected value and predicted value, including: Determine the difference between the expected value and the predicted value; Based on the difference, establish the probability distribution of the target state analysis graph; Based on the probability distribution of the target state analysis graph, guide the sampling of the sub-analysis graph corresponding to the large analysis graph to obtain the target state analysis graph.
[0010] Preferably, the pre-configured algorithm rules include: genetic algorithm or monte carlo algorithm.
[0011] Preferably, before calculating the probability distribution, expand the probability distribution to obtain a functional form.
[0012] Preferably, the intelligent agent further includes: an iteration module; the iteration module is used to update the initial state and the target state after the actions of the action trajectory, and cyclically control the actions of the value extraction module, the reverse sampling generation module, and the control planning module; when the difference between the expected value and the predicted value is within a preset range, end the iteration.
[0013] Preferably, the value-driven autonomous intelligent system further includes: An input module, an intelligent agent construction module, and an output module; wherein, the intelligent agent construction module includes: a value calibration unit, a value extraction unit, and a construction unit; Among them, the value calibration unit is used to generate a calibrated value; the value extraction unit is used to retrieve relevant values from the value library based on the calibrated value; the construction unit is used to construct an intelligent agent according to the calibrated value and / or the retrieved value.
[0014] Preferably, the steps for constructing the values in the value library are as follows: Group the data samples to obtain multiple sample groups; Each sample group corresponds to training a value.
[0015] Preferably, the steps for training the value are as follows: Extract the fluid state from the data samples; Discover value scalars from the fluid state; Train the pre-configured model according to the value scalars.
[0016] Preferably, the value extraction unit is further used to extract the values in the value library marked with representative labels; the calibrated values generated by the value calibration unit are multiple; when multiple values are generated, output each value and receive weight configuration.
[0017] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structure particularly pointed out in the written description and the drawings.
[0018] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0019] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 It is a schematic diagram of an autonomous intelligent system based on value-driven in an embodiment of the present invention; Figure 2 It is a driving principle diagram starting from scenarios in an embodiment of the present invention; Figure 3 It is a sampling process of an object parsing diagram based on value differences in an embodiment of the present invention; Figure 4 It is a schematic diagram of autonomous behavior generation based on iterative optimization of object value in an embodiment of the present invention; Figure 5 It is a schematic diagram of another autonomous intelligent system based on value-driven in an embodiment of the present invention; Figure 6 It is a training step diagram of value in an embodiment of the present invention. Detailed Embodiments
[0020] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0021] An embodiment of the present invention provides an autonomous intelligent system based on value-driven, as Figure 1 shown, including: an agent, wherein the agent includes: a value extraction module 1, a reverse sampling generation module 2, and a planning module 3; Among them, the value extraction module 1 extracts values from the scenario to obtain predicted values; the reverse sampling generation module 2 is used to generate an object state parsing diagram by reverse sampling based on the given expected value and predicted value; the planning module 3 is used to plan and generate an action trajectory based on the initial parsing diagram and the object state parsing diagram.
[0022] The value-driven autonomous intelligent system of the present invention drives the actions of intelligent agents through value. The principle is to first determine the difference between the value contained in the scenario and the given expected value, and generate a target state parsing graph by reverse sampling based on the difference. Then, the planning module is used to plan and generate an action trajectory based on the initial parsing graph and the target state parsing graph. That is, given the expected value V*, the predicted value V, and the value difference delta_V = V* - V, combined with the geometric personality trait vector h, a target state parsing graph (Goal PG) can be generated by reverse sampling. Given the initial parsing graph LPG and the target parsing graph GoalPG, an action trajectory is planned and generated to enable reaching the target state from the initial state, as shown in Figure 2 ; In addition, as shown in the figure, from the initial parsing graph (LPG), a sub-parsing graph (eLPG) is obtained through encoder processing, and then a top-down or bottom-up attention mechanism is introduced to obtain a sub-parsing graph with an attention mechanism. A learning flow state is obtained from the sub-parsing graph, and the learning flow state and the geometric personality trait vector h are combined and passed through a value function MLP to obtain a predicted value. The difference between the given expected value and the predicted value is ; Then, the selected sub-PG (the smallest unit that makes up the parsing graph) is derived by reverse deduction of the differences, and the target parsing graph Goal PG is obtained by sampling it. The planning module plans and generates an action trajectory based on causal chain search for task planning: when the initial state and the target state are given, to establish the transfer trajectory between these two states, a causal chain needs to be searched in the causal library to establish the connection between the two states. The causal chain is a very stable state transfer relationship, and using the causal chain will also greatly reduce the search space; where the given initial state is represented as the initial parsing graph, and the target state is represented as the target parsing graph; to implement the causal chain, it can be achieved by a pre-configured causal library, that is, when the initial state and the target state are given, to establish the transfer trajectory between these two states, a causal chain needs to be searched in the causal library to establish the connection between the two states. In addition, to reduce the sampling space, sampling can be only performed in the area of interest in terms of structure, and at the same time, using the initial state as the initial sample for sampling will greatly speed up the sampling process. Among them, the scenario can be specifically represented as one image or multiple images; when the autonomous intelligent system is applied in a virtual communication environment, the scenario is obtained through the user input that communicates with the agent or is automatically generated in the virtual communication environment; when the automatic intelligent system is applied in a real communication scenario (installed in a smart robot), the scenario is obtained through the corresponding configured image acquisition module for acquisition. For example, in a virtual environment, the agent can complete a series of tasks. A specific task is that the agent tidies up the items on the desktop. It first learns the function of the value dimension of "tidy", and then samples various forms of tidiness based on this function, including placing the items on a line or a circle, and then plans specific operation steps based on this line or circle, and then executes to complete the task, showing value-driven autonomy.
[0023] Among them, the value extraction module extracts the value from the scenario to obtain the predicted value and performs the following operations: Parse the scenario to generate a large parsing graph; Take the value of each area in the scenario based on the large parsing graph; Extract the value of the area of interest from the large parsing graph as the predicted value.
[0024] Extract the large parsing graph (LPG) of the scenario, which depicts the value assignment of each area in the scenario, thus forming a value map (Value Map). When the value Vi of the area of interest differs from the expected value V*, an expected difference (or called expected gap) delta_V is generated. This expected difference will then drive the agent to generate a target task state, thus generating a new task.
[0025] Among them, the reverse sampling generation module is used to reverse sample and generate the target state parsing graph based on the given expected value and predicted value, including: Determine the difference between the expected value and the predicted value; Establish the probability distribution of the target state parsing graph based on the difference; Guide the sampling of the sub-parsing graph corresponding to the large parsing graph based on the probability distribution of the target state parsing graph to obtain the target state parsing graph.
[0026] The pre-configured algorithm rules include: genetic algorithm or monte carlo algorithm.
[0027] Before calculating the probability distribution, expand the probability distribution to obtain the function form.
[0028] Establish the probability distribution p(pg*) of the target parsing graph pg* based on delta_V, as Figure 3 shown, use MCMC (Markov Chain Monte Carlo method) sampling ; expressed as ; ; where it should be noted that The sampling is guided by the scenario / task grammar to ensure effectiveness ; focus on sub sampling rather than the entire LPG sampling ; perform conditional sampling based on the existing LPG to simplify the subsequent task planning ; The form of this distribution function can be further expanded to obtain the function form with pg* as the variable. Here, the genetic algorithm or monte carlo algorithm can be used to calculate the pg* with the maximum probability. Taking the monte carlo algorithm as an example, three aspects need to be considered during the sampling process. First, the sampling process is carried out under the guidance of a task grammar rule, which can ensure the rationality of the sampling task; second, the entire sampling process only samples the sub-parsing graph of interest and does not need to sample the entire large parsing graph; finally, sampling based on the parsing graph of the initial state will be more efficient.
[0029] In one embodiment, the intelligent agent further includes: an iteration module; the iteration module is used to update the initial state and the target state after the actions of the action trajectory, and cycle to control the actions of the value extraction module, the reverse sampling generation module, and the control planning module; when the difference between the expected value and the predicted value is within the preset range, end the iteration.
[0030] See Figure 4 For searching the effective causal link path from to adopt a forward / backward search strategy, such as ,[[]] , RRT, CEM; during the search, evaluate and calculate , stop if the condition is met; if not, based on the newly generated Return and sample new ; Since the target task state sampled may not fully reflect the target value, and in addition, the target task state cannot be perfectly achieved in task planning, the uncertainties brought by these two factors result in the inability to complete the value-driven task in one go. It is necessary to iteratively update the initial state and the target state, approaching the target value step by step, so as to implement multiple tasks. When the distance between the true value and the target value is within the acceptable range, the task generation ends.
[0031] In one embodiment, a value-driven autonomous intelligent system, such as Figure 5 shown, further includes: an input module 11, an agent construction module 12, and an output module 13; wherein, the agent construction module 12 includes: a value calibration unit 121, a value extraction unit 122, and a construction unit 123; Among them, the value calibration unit 121 is used to generate a calibrated value; the value extraction unit 122 is used to retrieve relevant values from the value library based on the calibrated value; the construction unit 123 is used to construct an agent according to the calibrated value and / or the retrieved value; The value-driven autonomous intelligent system of this embodiment retrieves relevant values from a pre-configured value library through the value extraction unit on the basis of a manually set value, thereby constructing multiple agents, and further realizing more intelligent communication. Through agents guided by multiple different values, insights into the multi-faceted nature of the same problem are achieved, and further more intelligent communication is realized.
[0032] The working principle of the agent can be specifically understood as follows: First, extract the large parse graph (LPG) of the given initial scene, which depicts the value assignments of each region in the scene, thus forming a value map (Value Map). When there is a difference between the value Vi of the region of interest and the expected value V*, an expected difference (or called expected gap) delta_V is generated. This expected difference will drive the agent to generate a target task state, thereby generating a new task, that is, given the expected value V*, predict the value V, and its value difference delta_V = V* - V. The geometric personality trait vector h can be inversely sampled to generate a target state parse graph. Given the initial parse graph LPG and the target parse graph Goal PG, plan to generate an action trajectory so that it can reach the target state from the initial state; among them, based on delta_V, establish the probability distribution p(pg*) of the target parse graph pg*, such as Figure 3As shown, the form of this distribution function can be further expanded to obtain a functional form with pg* as the variable. Here, a genetic algorithm or a Monte Carlo algorithm can be used to calculate the pg* with the maximum probability. Taking the Monte Carlo algorithm as an example, three aspects need to be considered during the sampling process. First, the sampling process is carried out under the guidance of a task grammar rule, which can ensure the rationality of the sampling task. Second, the entire sampling process only samples the concerned sub-parse graph and does not need to sample the entire large parse graph. Finally, sampling based on the parse graph in the initial state will be more efficient. Since the target task state obtained by sampling may not fully reflect the target value, and in task planning, the target task state cannot be perfectly achieved either, the uncertainty brought by these two aspects leads to the inability to complete the value-driven task in one go. It is necessary to iteratively update the initial state and the target state, approaching the target value step by step, thereby implementing multiple tasks. When the distance between the true value and the target value is within the acceptable range, the task generation ends.
[0033] In one embodiment, the steps for constructing the values in the value library are as follows: Group the data samples to obtain multiple sample groups; Each sample group is correspondingly trained for a value.
[0034] The data samples can be grouped in a random manner, that is, it can be set how many groups of samples are needed, and then the samples are selected by a random extraction and filling method. Each sample group can train a value, which is then stored in the value library for later use.
[0035] In one embodiment, as Figure 6 shown, the steps for training the value are as follows: Step 1: Extract the flow regime from the data samples; Step 2: Discover the value scalar from the flow regime; Step 3: Train the pre-configured model according to the value scalar.
[0036] Value, that is, the value function. Its specific generation process is as follows: Sample Data: This is the raw data collected from the environment or task, which may be images, audio, video, text, or structured data, etc. Fluent / Latent Variable: These are the potential dynamic features or hidden patterns behind the data, which can reflect the changes of the data in different states. The latent variable represents a certain potential attribute or state of a data sample. Value Scalar: This is a numerical value representing the value of a certain data sample, usually obtained through the output of the model, representing the pros and cons of the target state or behavior corresponding to the sample. The generation of latent variables and sample data is as follows: For sequential data (such as time series, text, etc.), we use RNN or Transformer, where the latent variable is used as the input of the hidden layer of the network to capture the dependencies and state transitions in the time series. For image data, we use CNN or UNet, where the latent variable is used as an additional feature encoding to help extract the key features in the image and generate data samples. For graph data (such as social networks or traffic networks), we can use GNN (Graph Neural Network) to process the relationships between nodes in the data through the graph structure, and then infer the latent variables and sample data. Through these different architectures, latent variables that conform to different modal data can be generated to help establish an accurate value function.
[0037] In one embodiment, the value calibration unit is used to generate calibrated values and perform the following operations: Output preset value questionnaire items; Receive feedback information for each value questionnaire item; Based on the feedback information and the pre-configured parameter correspondence table, determine the parameter data; Based on the determined parameter data, generate calibrated values.
[0038] The calibrated value is essentially for the user and is constructed based on the information given by the user. How to understand the user's information can be done by configuring value questionnaire items. The user configures different values according to the value questionnaire items. The value questionnaire items are configured as multiple, and each value questionnaire item corresponds to multiple feedback items, and each feedback item corresponds to different parameter data values; then, by integrating all the questionnaire items, the calibrated value can be obtained; furthermore, the user can select a scenario in the early stage and give different value questionnaire items according to different scenarios.
[0039] To extract relevant values from the value library, the value extraction unit is used to retrieve relevant values from the value library based on the calibrated values and perform the following operations: Arrange the parameter data in the calibrated value in the pre-configured order to form a first data set; Arrange the parameter data in each value in the value library in the pre-configured order to form a second data set; Calculate the similarity between the first data set and the second data set; Extract the value with the smallest difference among the relevant thresholds specified in the pre-configured relevant set for the similarity.
[0040] By quantifying each value to form the first data set and the second data set, and then comparing the value in the value library with the calibrated value in the way of similarity calculation. In addition, when retrieving relevant values, it is guided by the relevant set. Generally, the so-called relevance is the maximum similarity; the relevance can also be expressed as the maximum difference, that is, the minimum similarity is optimally zero; therefore, in this embodiment, the relevant set is used to determine the relevant situation. For example, 0%, 20%, 50%, and 100% are configured in the relevant set; in this way, 4 values need to be retrieved, namely the values with the maximum, minimum similarity and the values close to 20% and 50%; in addition, the cosine similarity calculation method can be used to calculate the similarity.
[0041] In one embodiment, the sample data can be collected by solicitation. The specific steps are as follows: By publishing scenario data to solicit social interaction data from the scenario; the data on the big data platform is existing and the acquisition speed is relatively fast; while the data obtained by publishing scenario data has a relatively slow speed. The present invention combines the two collection methods. The specific steps are as follows: Extract the data corresponding to the classification on the big data platform and determine the statistical parameters of the data; when the statistical parameters do not meet the preset conditions; publish the scenario data to obtain the solicited data; update the statistical parameters based on the solicited data, and when the preset conditions are met, end the solicitation; where the statistical parameters include: the total number of data, the number of data of each different interaction decision data, the number of interaction decision data, etc., one or more combinations thereof; the preset conditions include: the data threshold corresponding to the total number of data, the data threshold corresponding to each different interaction decision data, the quantity threshold corresponding to the number of interaction decision data, etc., one or more combinations thereof; In addition, when collecting data, the validity of the data needs to be verified. The specific verification method can be to analyze the users who provide the data, and indirectly verify the data through this. The specific analysis process is as follows: Obtain the permission information of the user, and determine the first score value based on the permission information and the pre-configured first scoring table; Obtain the data provided by the user corresponding to the historical collection in the collection module, calculate the similarity between the data provided in the historical collection and the currently provided data, and determine the second score value based on the similarity and the pre-configured second scoring table; Based on the time interval of the data provided by the user in the collection module for a preset number of times, construct an analysis vector, and retrieve the corresponding third score value from the pre-configured third scoring library according to the analysis vector; When the sum of the first score value, the second score value, and the third score value is greater than the preset scoring threshold, the verification passes; Among them, the second score value in the second scoring table corresponds one-to-one with the similarity. The specific corresponding relationship can be configured with an intermediate value (for example, similarity 0.8 corresponds to score 100). Starting from the similarity of the intermediate value, as the similarity increases or decreases, the second score value gradually decreases; The third score value in the third scoring library is associated with the analysis vector one-to-one. The third scoring library is pre-configured in advance; When configuring, follow the following rules: The shorter the average interval time, the lower the third score value; For the same average time interval, the smaller the item with the interval time, the lower the third score value; Regarding the permission information of the user, it is dynamically determined based on the initial configuration value and then analyzing the user's behavior to ensure the accuracy of the determination; The determination of the permission value corresponding to the specific permission information is to count the number of times the data provided by the user is determined to be accurate and the number of times it is determined to be inaccurate; When the number of times determined to be inaccurate is greater than the preset threshold, the permission value is set to zero; Otherwise, it is the sum of the initial configuration value, the value obtained by multiplying the number of times determined to be accurate by the first adjustment value, and the value obtained by multiplying the number of times determined to be accurate by the second adjustment value; Among them, the first adjustment value is positive, the second adjustment value is negative, and the absolute value of the first adjustment value is less than the absolute value of the second adjustment value. By setting different first and second adjustment values, it can adapt to the punitive reduction of the permission value for malicious provision; to avoid malicious data provision next time.
[0042] In actual use, different value functions are called according to different current scenarios to make targeted decisions, improving the anthropomorphic level; and different scenario configurations correspond to different value functions for different personalities to adapt to the interaction experiences of different personalities.
[0043] In one embodiment, the value extraction unit is further configured to extract the values in the value library labeled with representative tags.
[0044] Through the representative tags, synchronous output is performed to realize the feedback of representative values; The annotation of the representative tags is performed by professional personnel after professional analysis.
[0045] In one embodiment, to enable the autonomous intelligent system to communicate with the user, an input module is used to receive the input data of the user; an output module is used to output the output data of each agent for the input data.
[0046] The input and output of the autonomous intelligent system are realized through the input data and the output data. The input module transmits the input data to each agent respectively; the output module receives the output of each agent.
[0047] In one embodiment, the value calibration unit generates multiple calibrated values; when multiple values are generated, each value is output and weight configuration is received.
[0048] In reality, a person is complex, so a person has multiple thoughts, usually dominated by a certain thought; therefore, multiple feedback items can be selected for the value questionnaire items in the provided value calibration unit; multiple calibrated values are generated in this way, and then the multiple calibrated values are described and output, and the user assigns weights, and the sum of the weights is 1; in this process, through the pre-configured value-description correspondence library, through the generated calibrated values, the corresponding value descriptions are retrieved from the library for output. Because the value function itself is a data that not everyone can understand, by describing and exemplifying each data of the value function, users can have an intuitive impression; thus facilitating the user to assign weights. For example, in the value function, data A represents the choice at a fork in the road, 0 represents the right side, and 1 represents the left side; the corresponding descriptions are that when encountering a fork in the road, it tends to the right side, and when encountering a fork in the road, it tends to the left side. Configuring the bias weights of multiple values can further consider the randomness of the selection.
[0049] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A value-driven autonomous intelligent system, characterized in that, Including: An agent, where the agent includes: a value extraction module, an inverse sampling generation module, and a planning module; Among them, the value extraction module extracts values from the scenario to obtain predicted values; the inverse sampling generation module is used to generate a target state parsing graph by inverse sampling based on the given expected value and predicted value; the planning module is used to plan and generate an action trajectory based on the initial parsing graph and the target state parsing graph.
2. The value-driven autonomous intelligent system according to claim 1, wherein The value extraction module extracts values from the scenario to obtain predicted values and performs the following operations: Parse the scenario to generate a large parsing graph; Obtain the values of each region in the scenario based on the large parsing graph; Extract the value of the region of interest from the large parsing graph as the predicted value.
3. The value-driven autonomous intelligent system according to claim 2, wherein, The inverse sampling generation module is used to generate a target state parsing graph by inverse sampling based on the given expected value and predicted value, including: Determine the difference between the expected value and the predicted value; Establish the probability distribution of the target state parsing graph based on the difference; Sample the sub-parsing graph corresponding to the large parsing graph based on the probability distribution of the target state parsing graph to obtain the target state parsing graph.
4. The value-driven autonomous intelligent system according to claim 3, wherein The pre-configured algorithm rules include: genetic algorithm or monte carlo algorithm.
5. The value-driven autonomous intelligent system according to claim 4, wherein Before calculating the probability distribution, expand the probability distribution to obtain a function form.
6. The value-driven autonomous intelligent system according to claim 1, characterized in that The agent further includes: an iteration module; the iteration module is used to update the initial state and the target state after the actions of the action trajectory, and cyclically control the actions of the value extraction module, the inverse sampling generation module, and the control planning module; when the difference between the expected value and the predicted value is within a preset range, the iteration ends.
7. The value-driven autonomous intelligent system according to claim 1, wherein Also including: An input module, an agent construction module, and an output module; among them, the agent construction module includes: a value calibration unit, a value extraction unit, and a construction unit; Among them, the value calibration unit is used to generate calibrated values; the value extraction unit is used to retrieve relevant values from the value library based on the calibrated values; the construction unit is used to construct an agent according to the calibrated values and / or the retrieved values.
8. The value-driven autonomous intelligent system according to claim 7, characterized in that The construction steps of the values in the value library are as follows: Group the data samples to obtain multiple sample groups; Each sample group corresponds to training a value.
9. The value-driven autonomous intelligent system according to claim 8, characterized in that The training steps of the value are as follows: Extract the fluid state from the data samples; Discover value scalars from the fluid state; Train the pre-configured model according to the value scalars.
10. The value-driven autonomous intelligent system according to claim 1, characterized in that, The value extraction unit is further used to extract the values in the value library labeled with representative labels; the calibrated values generated by the value calibration unit are multiple; when multiple values are generated, each value is output and weight configuration is received.
Citation Information
Patent Citations
Underwater robot control method based on reinforcement learning and control method of tracking with underwater robot
CN109240091A
Method, device and storage medium for continuous space action planning of intelligent agent
CN112264999A
Intelligent agent training method, computer equipment and storage medium
CN115759284A
Method and device for generating agent social behaviors based on value driving
CN117350907A
Social interaction cognition method and system based on value function
CN119721112A