Robot control method, device, electronic device and storage medium
By updating and optimizing the weight matrix of the user action prediction model and combining the robot action determination model, the problem of low prediction accuracy of robot user behavior in the prior art is solved, and a higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202211032396.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-08-26
AI Technical Summary
In the prior art, robots have low accuracy in predicting user behavior.
By obtaining the weight matrix of the pre-constructed user action prediction model, the estimated reward based on the weight matrix is determined, and the weight matrix and the estimated reward are input into the robot action determination model, and a robot action that can improve the estimated reward is obtained. The control robot then performs the action, updates the weight matrix, and predicts the user action using the updated model.
Improves the accuracy of robot prediction of user behavior.
Smart Images

Figure CN115401693B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of robot technology, and in particular to a robot control method, device, electronic device and storage medium. Background Art
[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims. No description herein is admitted to be prior art by inclusion in this section.
[0003] In an environment where both intelligent robots and users exist, the robot's behavior will be affected not only by the current environment, but also by the user's behavior.
[0004] In the related art, for robots, it is usually considered to model the user so that the robot can predict the user's behavior, and then select and execute the robot behavior according to the user's behavior.
[0005] However, in the prior art, the accuracy of robots' predictions of user behavior is low. Summary of the invention
[0006] In view of this, the purpose of the present disclosure is to propose a robot control method, device, electronic device and storage medium, which at least to a certain extent solve one of the technical problems in the related art.
[0007] Based on the above purpose, the exemplary embodiment of the present disclosure provides a control method of a robot, the method comprising:
[0008] Obtaining a weight matrix of a pre-built user action prediction model, and determining an estimated reward for controlling the robot to perform a robot action based on the weight matrix;
[0009] Inputting the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model; wherein the first robot action can improve the estimated reward;
[0010] controlling the robot to execute the first robot action, obtaining an execution reward, and updating the weight matrix according to the execution reward;
[0011] The user action is predicted using the user action prediction model with the updated weight matrix to obtain a first predicted user action, the first predicted user action is input into the robot action determination model to obtain a second robot action output by the robot action determination model, and the robot is controlled to perform the second robot action.
[0012] In some exemplary embodiments, the method further comprises:
[0013] determining a user action determination model from a pre-built user action determination model library according to the first robot action;
[0014] Acquire local observation information, and input the local observation information into the user action determination model to obtain a reconstructed user action output by the user action determination model;
[0015] In a pre-built basic environment, the local observation information is updated according to a current user action to obtain first updated local observation information;
[0016] In a pre-constructed parallel environment, updating the local observation information according to the reconstructed user action to obtain second updated local observation information;
[0017] Calculating planning similarity according to the first updated local observation information and the second updated local observation information;
[0018] The weight matrix is updated according to the planning similarity.
[0019] In some exemplary embodiments, the method further comprises:
[0020] determining a user action determination model from a pre-built user action determination model library according to the first robot action;
[0021] Acquire local observation information, and input the local observation information into the user action determination model to obtain a reconstructed user action output by the user action determination model;
[0022] Determining a current user action, and determining a reconstruction accuracy based on the current user action and the reconstructed user action;
[0023] The weight matrix is updated according to the reconstruction accuracy.
[0024] In some exemplary embodiments, the user action includes a movement operation for a virtual character controlled by the user;
[0025] The robot action includes movement operations for a virtual character controlled by the robot.
[0026] In some exemplary embodiments, determining an estimated reward corresponding to controlling the robot to perform a robot action based on the weight matrix includes:
[0027] pass Determining an estimated reward corresponding to controlling the robot to perform a robot action based on the weight matrix;
[0028] in, represents the estimated reward, π∈Π means the robot action π belongs to the robot action set Π, It indicates that the user action τ belongs to the user action set T, β(τ) represents the weight matrix of the user action prediction model, and E[U|τ,π] represents the functional relationship between the reward U and the user action τ and the robot action π.
[0029] In some exemplary embodiments, inputting the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model includes:
[0030] The weight matrix and the estimated reward are input into a pre-built robot action determination model, by Obtaining a first robot action output by the robot action determination model;
[0031] Among them, π * represents the first robot action, π∈Π means the robot action π belongs to the robot action set Π, U max represents the maximum reward, Represents the estimated reward, indicates that the user action τ belongs to the user action set Τ, β(τ) represents the weight matrix of the user action prediction model, P(U + |τ,π) represents the increased reward U + The functional relationship between the user action τ and the robot action π.
[0032] In some exemplary embodiments, updating the weight matrix according to the execution reward includes:
[0033] pass updating the weight matrix according to the execution reward;
[0034] Among them, β k (τ) represents the weight matrix of round k, β k-1 (τ) represents the weight matrix of the k-1 round, P(u k |τ,π k ) represents the execution reward u of round k k With user action τ and robot action π in round k k The functional relationship between It means that the user action τ′ belongs to the user action set T.
[0035] Based on the same inventive concept, the exemplary embodiment of the present disclosure further provides a control device for a robot, the device comprising:
[0036] a user action prediction model determination module configured to obtain a weight matrix of a pre-built user action prediction model and determine an estimated reward for controlling the robot to perform a robot action based on the weight matrix;
[0037] A robot action determination module is configured to input the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model; wherein the first robot action can improve the estimated reward;
[0038] a user action prediction model updating module, configured to control the robot to execute the first robot action, obtain an execution reward, and update the weight matrix according to the execution reward;
[0039] The robot action update module is configured to predict the user action using the user action prediction model with the weight matrix updated to obtain a first predicted user action, input the first predicted user action into the robot action determination model to obtain a second robot action output by the robot action determination model, and control the robot to perform the second robot action.
[0040] Based on the same inventive concept, an exemplary embodiment of the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the methods described above is implemented.
[0041] Based on the same inventive concept, an exemplary embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute any of the above methods.
[0042] From the above, it can be seen that the control method, device, electronic device and storage medium of the robot provided by the embodiment of the present disclosure include: obtaining the weight matrix of the pre-constructed user action prediction model, and determining the estimated reward for controlling the robot to perform the robot action based on the weight matrix; inputting the weight matrix and the estimated reward into the pre-constructed robot action determination model to obtain the first robot action output by the robot action determination model; wherein the first robot action can improve the estimated reward; controlling the robot to perform the first robot action, obtaining the execution reward, and updating the weight matrix according to the execution reward; predicting the user action using the user action prediction model with the updated weight matrix, obtaining the first predicted user action, inputting the first predicted user action into the robot action determination model, obtaining the second robot action output by the robot action determination model, and controlling the robot to perform the second robot action. Through the present disclosure, the accuracy of the robot's prediction of user behavior can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0044] Figure 1 A schematic diagram of an application scenario of the robot control method provided according to an embodiment of the present disclosure;
[0045] Figure 2 A schematic diagram of a flow chart of a robot control method provided by an embodiment of the present disclosure;
[0046] Figure 3 Another schematic flow chart of a robot control method provided by an embodiment of the present disclosure;
[0047] Figure 4 Another schematic flow chart of a robot control method provided by an embodiment of the present disclosure;
[0048] Figure 5 A schematic diagram of an experimental scenario of a robot control method provided according to an embodiment of the present disclosure;
[0049] Figure 6 A schematic diagram of the accuracy results of the user model reconstruction action in the offline stage provided according to an embodiment of the present disclosure;
[0050] Figure 7 A schematic diagram of an offline phase performance model result provided according to an embodiment of the present disclosure;
[0051] Figure 8 A schematic diagram of an offline reconstruction accuracy model result provided according to an embodiment of the present disclosure;
[0052] Fig. 9 A schematic diagram of a similarity model result for offline phase planning provided according to an embodiment of the present disclosure;
[0053] Fig.10 A schematic diagram of a robot round reward result in an online phase provided according to an embodiment of the present disclosure;
[0054] Fig.11 A schematic diagram of the cumulative reward result of the online stage robot provided according to an embodiment of the present disclosure;
[0055] Fig.12 A schematic diagram of the robot recognition accuracy result in the online stage provided according to an embodiment of the present disclosure;
[0056] Fig.13 A schematic diagram of the effect of the user strategy switching interval on the algorithm performance during the online phase provided by an embodiment of the present disclosure;
[0057] Fig.14 A schematic diagram of a structure of a control device of a robot provided in an embodiment of the present disclosure;
[0058] Fig.15 A schematic diagram of a flow chart of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solutions and advantages of the present disclosure more clear, the principles and spirit of the present disclosure will be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0060] According to an embodiment of the present disclosure, a control method, device, electronic device and storage medium of a robot are proposed.
[0061] It should be understood herein that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction rather than having any limiting meaning.
[0062] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0063] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure.
[0064] refer to Figure 1 , which is a schematic diagram of an application scenario of the robot control method provided according to an embodiment of the present disclosure.
[0065] The application scenario includes a terminal device 101 , a server 102 , and a data storage system 103 .
[0066] The terminal device 101, the server 102 and the data storage system 103 may be connected via a wired or wireless communication network.
[0067] The terminal device 101 includes but is not limited to a desktop computer, a mobile phone, a mobile computer, a tablet computer, a media player, a smart wearable device, a personal digital assistant (PDA) or other electronic devices capable of realizing the above functions.
[0068] Server 102 and data storage system 103 can both be independent physical servers, or a server cluster or distributed system composed of multiple physical servers. They can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0069] The server 102 is used to provide interactive services of the robot to the user of the terminal device 101. The terminal device 101 is installed with a client that communicates with the server 102. The user can input a user behavior through the client. The server 102 executes the robot control method provided by the present disclosure and outputs a robot behavior through the client, thereby realizing the interaction between the robot and the user.
[0070] A large amount of training data is stored in the data storage system 103 , wherein the source of the training data includes but is not limited to an existing database, data crawled from the Internet, or data uploaded by a user when using a client.
[0071] Combine the following Figure 1 The control method of the robot according to the exemplary embodiment of the present disclosure is described by using the application scenario. It should be noted that the above application scenario is only shown to facilitate understanding of the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0072] refer to Figure 2 , which is a flow chart of the robot control method provided in the embodiment of the present disclosure.
[0073] The robot control method comprises the following steps:
[0074] Step S210: Obtain a weight matrix of a pre-built user action prediction model, and determine an estimated reward for controlling the robot to perform a robot action based on the weight matrix.
[0075] The user action prediction model is used to predict the user's actions, and the weight matrix of the user action prediction model represents the logic of the user action prediction model in predicting the user's behavior.
[0076] When the present disclosure is first implemented, the user action prediction model needs to be initialized. The prediction of user actions by the initialized user action prediction model is uniformly distributed. For example, if the number of user actions in the user action library maintained by the user action prediction model is N, the probability of each user action predicted by the initialized user action prediction model is 1 / N.
[0077] The estimated reward is the optimal reward that may be obtained by controlling the robot to perform robot actions based on the weight matrix.
[0078] In some exemplary embodiments, determining an estimated reward corresponding to controlling the robot to perform a robot action based on the weight matrix includes:
[0079] pass Determining an estimated reward corresponding to controlling the robot to perform a robot action based on the weight matrix;
[0080] in, represents the estimated reward, π∈Π means the robot action π belongs to the robot action set Π, It indicates that the user action τ belongs to the user action set T, β(τ) represents the weight matrix of the user action prediction model, and E[U|τ,π] represents the functional relationship between the reward U and the user action τ and the robot action π.
[0081] Step S220: input the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model; wherein the first robot action can improve the estimated reward.
[0082] In some exemplary embodiments, inputting the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model includes:
[0083] The weight matrix and the estimated reward are input into a pre-built robot action determination model, by Obtaining a first robot action output by the robot action determination model;
[0084] Among them, π * represents the first robot action, π∈Π means the robot action π belongs to the robot action set Π, U max represents the maximum reward, Represents the estimated reward, indicates that the user action τ belongs to the user action set Τ, β(τ) represents the weight matrix of the user action prediction model, P(U + |τ,π) represents the increased reward U + The functional relationship between the user action τ and the robot action π.
[0085] In the specific implementation, in the offline stage, the robot interacts with different user strategies, uses the A2C algorithm to learn the response strategies, and builds a response strategy library.
[0086] In the offline stage, the present disclosure uses the A2C algorithm to learn the response strategy. Specifically, during the interaction between the robot and the user, the robot's state s, action a, and reward r are collected and used to update the strategy model π θ , where θ represents the parameters of the strategy model. For simplicity, in the subsequent description, the present disclosure award strategy model π θ Abbreviated as π. A2C uses a parallel environment to break the correlation between experience samples and introduces advantage terms Accelerate the learning process. The loss function during A2C update can be expressed as:
[0087]
[0088] Among them, B is the experience data within one round, H represents the total number of time steps in a round, For the advantage, γ represents the discount factor.
[0089] Step S230: control the robot to execute the first robot action, obtain an execution reward, and update the weight matrix according to the execution reward.
[0090] In some exemplary embodiments, updating the weight matrix according to the execution reward includes:
[0091] pass updating the weight matrix according to the execution reward;
[0092] Among them, β k (τ) represents the weight matrix of round k, β k-1 (τ) represents the weight matrix of the k-1 round, P(u k |τ,πk) represents the execution reward u for round k k With user action τ and robot action π in round k k The functional relationship between It means that the user action τ′ belongs to the user action set T.
[0093] During specific implementation, a performance model is constructed.
[0094] The performance model refers to the robot using strategy π∈Π and the user using strategy , the probability distribution P(U|τ,π) of the robot obtaining the cumulative utility reward U in one round. Specifically, for each user strategy τ in the user strategy library, the robot uses each strategy π∈Π in the optimal response strategy library to interact with it for multiple rounds, collects the cumulative utility reward U of the robot in each round, and fits it to a normal distribution.
[0095] Step S240: predict the user action using the user action prediction model that has updated the weight matrix to obtain a first predicted user action, input the first predicted user action into the robot action determination model to obtain a second robot action output by the robot action determination model, and control the robot to perform the second robot action.
[0096] In addition, the present disclosure proposes a method of combining two models for identifying user strategies according to the degree to which user information is available.
[0097] refer to Figure 3 , which is another flow chart of the robot control method provided in the embodiment of the present disclosure.
[0098] The robot control method comprises the following steps:
[0099] Step S310: Obtain a weight matrix of a pre-built user action prediction model, and determine an estimated reward for controlling the robot to perform a robot action based on the weight matrix.
[0100] Step S320: input the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model; wherein the first robot action can improve the estimated reward.
[0101] Step S330: control the robot to execute the first robot action, and obtain execution reward and local observation information.
[0102] Step S340: determining a user action determination model from a pre-built user action determination model library according to the first robot action.
[0103] Among them, since each robot action corresponds to a user action determination model, it can be determined according to the current robot action π k , determine the user action determination model x from the user action determination model library k .
[0104] Optionally, the user action determination model is built based on a VAE model framework.
[0105] In specific implementation, the VAE model is trained using the behavior data of robots and users, and a VAE-based user model library is built.
[0106] This paper uses the VAE model to learn the mapping relationship between the robot's local observation information and user behavior, and uses the trained VAE model to build a VAE-based user model library. Specifically, the VAE model consists of two parts, one is the encoder (Encoder), and the other is the decoder (Decoder), which collects the robot's local observation information in a round. Encoder uses a variational Gaussian distribution Approximate the true posterior distribution p(z) and map the robot's local observation information to hidden variables z, where σ is the parameter of the variational distribution. The similarity between the approximate variational distribution and the true posterior distribution is measured by KL divergence. The decoder generates a distribution Map the latent variable z to the user's local observation information Where α is the parameter of the generative distribution. The goal of the Encoder-Decoder network when updating is:
[0107]
[0108] Step S350: input the local observation information into the user action determination model to obtain the reconstructed user action output by the user action determination model.
[0109] Step S360: in a pre-constructed basic environment, update the local observation information according to the current user action to obtain first updated local observation information; in a pre-constructed parallel environment, update the local observation information according to the reconstructed user action to obtain second updated local observation information; based on the first updated local observation information and the second updated local observation information, calculate the planning similarity.
[0110] During specific implementation, a planning similarity model is constructed.
[0111] The planning similarity model refers to the reconstruction accuracy model, which refers to the robot's strategy π∈Π and the user's strategy The user model is The probability distribution of the robot's planning similarity Q in one round is Specifically, when the user's actions cannot be observed, a base environment and a parallel environment are created. The initial states of the two environments are the same and known. In the base environment, the user uses real actions, and in the parallel environment, the user uses actions reconstructed by the robot using the user model. The observations of the robot at each step in the base environment and the observations of the robot at each step in the parallel environment are collected, and the mean square error between the two is calculated to obtain the planning similarity Q of each round, and it is fitted to a normal distribution.
[0112] Step S370: Update the weight matrix according to the planning similarity.
[0113] When neither the user's observation nor the user's action can be obtained, the present disclosure uses a planning method to infer the user's type and combines the performance model with the planning similarity model to update the belief. Specifically, it is necessary to use the existing basic environment env b With the parallel environment env p , their initial states are consistent and known, and the actions taken by the robot in both environments are the same. b In the example, the user uses the real action {a -1}, and in the parallel environment env p In the example, the robot thinks that the user uses the reconstructed action Robot in envp Planning starts with a known initial state of the environment and infers the observations within a round Then from env b Sampling real observations {o 1}, by calculating With {o 1}, and obtain the round planning similarity q k Finally, the performance model and planning accuracy model are combined to update the round belief:
[0114]
[0115] Step S380: predict the user action using the user action prediction model that has updated the weight matrix to obtain a first predicted user action, input the first predicted user action into the robot action determination model to obtain a second robot action output by the robot action determination model, and control the robot to perform the second robot action.
[0116] refer to Figure 4 , which is another flow chart of the robot control method provided in the embodiment of the present disclosure.
[0117] The robot control method comprises the following steps:
[0118] Step S410: Obtain a weight matrix of a pre-built user action prediction model, and determine an estimated reward for controlling the robot to perform a robot action based on the weight matrix.
[0119] Step S420: input the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model; wherein the first robot action can improve the estimated reward.
[0120] Step S430: Control the robot to execute the first robot action, and obtain execution reward and local observation information.
[0121] Step S440: determining a user action determination model from a pre-built user action determination model library according to the first robot action.
[0122] Step S450: input the local observation information into the user action determination model to obtain the reconstructed user action output by the user action determination model.
[0123] Step S460: determine the current user action, and determine the reconstruction accuracy according to the current user action and the reconstructed user action.
[0124] During specific implementation, a reconstruction accuracy model is constructed.
[0125] The reconstruction accuracy model refers to the robot using strategy π∈Π and the user using strategy The user model is The probability distribution of the robot's action reconstruction accuracy W in one round is Specifically, when the user's actions can be observed, the robot reconstructs the user's actions with the user model at each step, compares the reconstructed actions with the real user actions, obtains the accuracy W of the action reconstruction in each round, and fits it to a normal distribution.
[0126] Step S470: Update the weight matrix according to the reconstruction accuracy.
[0127] When the user's observation is not available, but the user's action is available, the present disclosure combines the performance model with the reconstruction accuracy model to update the belief. Specifically, the robot reconstructs the user's action with the user model at each step, and compares the reconstructed action with the real user action to obtain the accuracy of the action reconstruction in each round w k , and use the reconstruction accuracy model and the performance model to jointly update the belief. The update method can be expressed as:
[0128]
[0129] Step S480, using the user action prediction model with the updated weight matrix to predict the user action, obtain a first predicted user action, input the first predicted user action into the robot action determination model, obtain a second robot action output by the robot action determination model, and control the robot to perform the second robot action.
[0130] In some exemplary embodiments, the user action includes a movement operation for a virtual character controlled by the user;
[0131] The robot action includes movement operations for a virtual character controlled by the robot.
[0132] In some exemplary embodiments, the state data, action data and reward data of the robot in each round are obtained, and the robot action model is updated according to the state data, the action data and the reward data.
[0133] The present disclosure mainly consists of two stages, namely, offline strategy learning and model generation stage, and online user detection and strategy reuse stage.
[0134] In the offline phase, users can The robot selects a strategy to fight against the robot. The robot learns the corresponding strategy through the Advantage Actor-Critic (A2C) algorithm and builds a response strategy library Π. At the same time, the variational autoencoder is used to learn the relationship between the robot's local observation information and user behavior, and a user model library based on VAE is built. Using the existing strategy library and model library, the robot uses round rewards to fit the performance model Fitting the reconstruction accuracy model with behavioral data and planning similarity model The above models all obey Gaussian distribution.
[0135] In the online interaction phase, the robot selects a strategy from the response strategy library based on the belief β(τ) about the user type and interacts with users of unknown types using the VAE-based user model library. The user's actions are reconstructed based only on local observation information, and then the reconstruction accuracy and planning similarity are calculated. Finally, using the round reward, reconstruction accuracy, and planning similarity, the robot updates its beliefs based on the model built in the offline phase, completing user strategy type detection and its own response strategy reuse.
[0136] The following is a description of the effectiveness of the proposed method in conjunction with a specific application environment. The comparison algorithms include BPR+, Bayes-ToMoP, and Deep BPR+. This disclosure uses Bayes-Lab (A) to indicate that the robot can observe the user's actions during the online phase, and uses Bayes-Lab to indicate that the robot cannot observe the user's actions during the online phase. In all experiments, this disclosure assumes that the robot policy library Π contains some user policies Therefore, when the user uses an unknown strategy, the robot should identify the unknown strategy as quickly as possible and learn how to respond.
[0137] Experimental design:
[0138] Figure 5The hunter-prey environment is shown, which consists of a robot 510, three users 520 and two obstacles 530. The robot 510 acts as the prey and the three users 520 act as the hunters. The hunter's goal is to collide with the prey as much as possible, while the prey's goal is to avoid collisions with the hunter as much as possible. At each time step, if the prey does not collide with the hunter, the prey receives a reward of 0.1, if it collides once, the reward is -1, if it collides twice, the reward is -5, if it collides three times, the reward is -10, and the hunter's reward is completely opposite to the prey; the action space of the prey and the hunter is discrete, including five actions: standstill, left, right, down, and up. The environment is surrounded by impenetrable walls. Any action that allows the robot or user to pass through the wall will not succeed. The prey and the hunter can only fight within the wall. In the experiment, the present disclosure controls the prey, and the hunter selects a strategy from a series of fixed strategies. This environment is partially observable. The prey can only observe the hunters and obstacles within its field of view, while the hunter can observe all robots and obstacles. All robots are observed continuously. Each episode ends after 100 time steps.
[0139] The present disclosure designs four fixed strategies for hunters, namely: (1) vertical priority pursuit (2) horizontal priority pursuit (3) clockwise pursuit (4) counterclockwise pursuit. When the hunter adopts vertical priority pursuit, he will first take an upward or downward action to shorten the vertical distance with the prey. When the vertical distance with the prey is small enough, he will take a left or right action to shorten the horizontal distance with the prey. When the horizontal distance with the prey is small enough, he will take an upward or downward action to shorten the vertical distance with the prey. When the hunter adopts a clockwise pursuit, he may pursue the prey in the order of right, down, left, and up. When the hunter adopts a counterclockwise pursuit, he will pursue the prey in the order of left, down, right, and up. These four fixed strategies constitute the user strategy library. In the online interaction stage, the present invention enables the prey and the hunter to fight against each other for 100 rounds. The hunter switches strategies every 10 rounds. The order of strategy switching is (2, 3, 1, 3, 4, 1, 4, 1, 3, 1).
[0140] Experimental results:
[0141] Figure 6The accuracy results of the user model in reconstructing user actions in the offline stage are shown. When user strategy 1 interacts with response strategy 1, user model 1 can reconstruct the user's actions with an accuracy of 79%, and user model 2 can reconstruct the user's actions with an accuracy of 78%. However, the accuracy of user models 3 and 4 is only 44% and 52%. This is because user strategies 1 and 2 have few random actions and tend to chase prey up or down. However, the actions of user strategies 3 and 4 are more random, and the probabilities of taking the four actions of up, down, left, and right are very similar, which makes it more difficult for VAE to reconstruct user actions.
[0142] Figure 7 The performance model is shown. Since there are four user strategies and four response strategies, they are combined in pairs to obtain a total of sixteen Gaussian distributions. Each Gaussian distribution reflects the round reward information when the corresponding user strategy interacts with the response strategy. As can be seen from the figure, for each user strategy, there is an optimal response strategy, which obtains the highest average round reward, which is around 10, and has the smallest variance. In addition, when the response strategy interacts with different user strategies, the distribution of the round rewards obtained may be very similar. For example, when the response strategy is 1, the performance distribution when interacting with user strategy 3 is very close to the performance distribution when interacting with user strategy 4. If you only rely on round rewards to detect users, then due to the high similarity of the distribution, it is likely that accurate identification will not be possible.
[0143] Figure 8 The prediction accuracy model is shown. Like the performance model, the prediction accuracy model is also composed of sixteen Gaussian distributions. Each distribution reflects the accuracy of the corresponding user model in predicting user actions when the response strategy interacts with the user strategy. For each user strategy, the action prediction accuracy of the optimal response strategy is the highest. Response strategies 1 and 2 can both achieve an accuracy of about 80% when interacting with user strategies 1 and 2 respectively, while response strategies 3 and 4 can only achieve an accuracy of about 50% when interacting with user strategies 3 and 4. Compared with the performance model, the distribution difference in the prediction accuracy model is greater. For example, in the performance model, the distribution of response strategy 1 when interacting with user strategy 3 and the distribution when interacting with user strategy 4 are very similar, but in the prediction accuracy model, their distributions are more different. In addition, the prediction accuracy model is applicable to situations where the user's actions can be obtained during online interaction. It can be used in conjunction with the performance model.
[0144] Fig. 9The planning similarity model is presented. The present disclosure uses the mean square error to measure the similarity between local observation information. The smaller the mean square error, the higher the similarity between observations, and vice versa. As can be seen from the figure, the mean square error of the optimal response strategy is the smallest, and response strategies 1 and 2 can both reach about 0. Since the action reconstruction accuracy of user models 3 and 4 is lower than that of 1 and 2, the error of local observation information caused by the planning method will also be greater, so the minimum mean square error of response strategies 3 and 4 is about 200. In addition, the planning similarity model is suitable for situations where neither user observations nor user actions can be obtained during online interactions, and it can be used in conjunction with the performance model.
[0145] Fig.10The round reward results of the online reuse phase are shown. When the robot is not caught by the user, it can get a reward of 0.1 for each step. Therefore, within 100 steps in a round, the reward obtained by the robot for accurately identifying the user strategy and reusing the optimal response strategy will be around 10. It can be clearly observed that the round reward will drop rapidly every 10 rounds of user strategy switching. By observing whether the round rewards of different algorithms can quickly recover to around 10, the performance of the algorithm can be judged. The faster the recovery speed, the better the recognition and reuse effect of the algorithm. When the user's observation is not available, but the user's action is available, Bayes-Lab (A) uses detection mode 0 in Algorithm 2 to detect the user. When the user's observation and action are both unavailable, Bayes-Lab uses detection mode 1 in Algorithm 2 to detect the user. From the results, Bayes-Lab (A) regains the optimal reward at the fastest speed during each strategy switching process, which shows that the performance of Bayes-Lab (A) is better than other algorithms. Bayes-Lab is similar to Bayes-Lab(A) in most cases, but in the 50-60th and 70-80th rounds, that is, when the user strategy switches from 4 to 1, the round reward of Bayes-Lab is significantly lower than that of Bayes-Lab. By observing the Gaussian distribution, it is found that when the robot strategy is 4, the distribution difference between the prediction accuracy model 4 to 1 and 4 to 2 is larger, while the distribution difference between the planning similarity 4 to 1 and 4 to 2 is small. Therefore, the prediction accuracy model is more suitable for dealing with the situation where the user strategy 4 switches to the user strategy 1. The performance of Deep-BPR+ is worse than that of Bayes-Lab(A) and Bayes-Lab. This is because Deep-BPR+ needs to obtain the user's observations and actions, but in a local observable environment, it is assumed that Deep-BPR+ can only obtain the user's observations and actions within the robot's field of view, otherwise Deep-BPR+ will degenerate into BPR+. However, the performance of Deep-BPR+ is significantly better than that of BPR+. This is because BPR+ can only rely on round reward signals to update beliefs. However, through the Gaussian distribution of round rewards, we can see that the distributions of 1 to 3 and 1 to 4, 2 to 3 and 2 to 4 are relatively similar. BPR+ has a very low recognition accuracy in similar situations. Bayes-Tomop has the lowest performance among all algorithms. This is because the user strategy is switched randomly, but Bayes-Tomop assumes that the user will also use the BPR selection strategy. Therefore, it takes more time to determine whether the user is using BPR, resulting in very low round rewards during this period.
[0146] Fig.11The cumulative rewards of 100 rounds of interaction between the robot and the user are shown. It can be observed that each time the user strategy switches, the cumulative reward of the algorithm tends to decrease, but as the recognition and reuse are completed, the cumulative reward rises again. Bayes-Lab (A) obtains the highest cumulative reward because it can quickly detect each strategy switch and reuse the optimal response strategy. The cumulative reward of Bayes-Lab is not as good as Bayes-Lab (A), but it is higher than the other three algorithms. It proves that the planning similarity model can compensate for the impact of the lack of user observation and action to a certain extent. The cumulative rewards of Deep-BPR+ and Bayes-Lab are similar, but significantly higher than the cumulative reward of BPR+, which proves that Deep-BPR+ is more effective in detecting user types by using user behavior than BPR+ using only reward signals to detect user types. The Bayes-Tomop cumulative reward is the lowest and the variance is the largest, which proves that under different round rewards, the Bayes-Tomop process of detecting whether the user uses BPR is not stable. It assumes that the optimal rewards of all optimal response strategies are consistent. However, in this environment, even if the robot chooses the optimal response strategy, the optimal reward it obtains may fluctuate around 10, which makes the Bayes-Tomop detection more difficult.
[0147] Fig.12 The detection accuracy of 100 rounds of robot-user interaction is shown. Bayes-Lab (A) achieves relatively high detection accuracy compared to Bayes-Lab and Deep-BPR+, but Bayes-Lab (A) and Bayes-Lab lack complete user information. Deep-BPR+ requires complete user behavior to achieve high detection accuracy, so Bayes-Lab is more practical. BPR+ can also achieve 65% detection accuracy by relying solely on round reward signals, but when the performance distribution is similar, the detection accuracy of BPR+ will be greatly affected. Bayes-Tomop has the lowest detection accuracy because in many rounds of incorrect detection, Bayes-Tomop is detecting whether the user uses BPR.
[0148] Fig.13The effect of the switching interval on the algorithm performance is shown: In order to explore the effect of the switching frequency of the strategy on the algorithm, the present disclosure tries four different strategy switching intervals. It can be clearly seen from the results that the larger the switching interval, the higher the detection accuracy of all algorithms, among which Bayes-Lab (A) and Bayes-Lab have significantly better performance than other algorithms under the four switching intervals. When the switching interval is 12, the detection accuracy of all algorithms is relatively high, the detection accuracy of Deep-BPR+ and BPR+ exceeds 70%, and the detection accuracy of Bayes-Tomop exceeds 60%. Therefore, the longer the user strategy remains fixed, the more conducive it is to the effect of algorithm detection. When the switching interval is switched from 12 to 10, the performance of all algorithms decreases slightly. When the switching interval is switched from 12 to 8, the performance of Bayes-Lab (A) and Bayes-Lab decreases relatively little. The detection accuracy of Bayes-Lab (A) and Bayes-Lab decreases by about 7%, Deep-BPR+ decreases by 11%, BPR+ decreases by 13%, and Bayes-Tomop decreases by 14%. When the switching interval is from 8 to 4, the performance of all algorithms drops significantly, and the gap between Bayes-Lab(A) and other algorithms widens significantly. The results show that the switching interval of user strategies has an important impact on the detection accuracy of the algorithm. The larger the switching interval, the less difficult the detection is and the better the detection effect is. When users frequently switch strategies, it will bring great challenges to detection. Generally speaking, the performance of Bayes-Lab and Bayes-Lab(A) is relatively small when dealing with this challenge.
[0149] In summary, the Baye-Lab proposed in this disclosure can accurately identify user strategies and reuse the optimal response strategies in partially observable environments by relying only on the robot's local observation information.
[0150] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server. The method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the described method.
[0151] It should be noted that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0152] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a robot control device.
[0153] refer to Fig.14 , which is a structural schematic diagram of the control device of the robot provided in an embodiment of the present disclosure.
[0154] The robot's control device includes the following modules:
[0155] a user action prediction model determination module configured to obtain a weight matrix of a pre-built user action prediction model and determine an estimated reward for controlling the robot to perform a robot action based on the weight matrix;
[0156] A robot action determination module is configured to input the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model; wherein the first robot action can improve the estimated reward;
[0157] a user action prediction model updating module, configured to control the robot to execute the first robot action, obtain an execution reward, and update the weight matrix according to the execution reward;
[0158] The robot action update module is configured to predict the user action using the user action prediction model with the weight matrix updated to obtain a first predicted user action, input the first predicted user action into the robot action determination model to obtain a second robot action output by the robot action determination model, and control the robot to perform the second robot action.
[0159] In some exemplary embodiments, the method further comprises:
[0160] determining a user action determination model from a pre-built user action determination model library according to the first robot action;
[0161] Acquire local observation information, and input the local observation information into the user action determination model to obtain a reconstructed user action output by the user action determination model;
[0162] In a pre-built basic environment, the local observation information is updated according to a current user action to obtain first updated local observation information;
[0163] In a pre-constructed parallel environment, updating the local observation information according to the reconstructed user action to obtain second updated local observation information;
[0164] Calculating planning similarity according to the first updated local observation information and the second updated local observation information;
[0165] The weight matrix is updated according to the planning similarity.
[0166] In some exemplary embodiments, the method further comprises:
[0167] determining a user action determination model from a pre-built user action determination model library according to the first robot action;
[0168] Acquire local observation information, and input the local observation information into the user action determination model to obtain a reconstructed user action output by the user action determination model;
[0169] Determining a current user action, and determining a reconstruction accuracy based on the current user action and the reconstructed user action;
[0170] The weight matrix is updated according to the reconstruction accuracy.
[0171] In some exemplary embodiments, the user action includes a movement operation for a virtual character controlled by the user;
[0172] The robot action includes movement operations for a virtual character controlled by the robot.
[0173] In some exemplary embodiments, determining an estimated reward corresponding to controlling the robot to perform a robot action based on the weight matrix includes:
[0174] pass Determining an estimated reward corresponding to controlling the robot to perform a robot action based on the weight matrix;
[0175] in, represents the estimated reward, π∈Π means the robot action π belongs to the robot action set Π, It indicates that the user action τ belongs to the user action set T, β(τ) represents the weight matrix of the user action prediction model, and E[U|τ,π] represents the functional relationship between the reward U and the user action τ and the robot action π.
[0176] In some exemplary embodiments, inputting the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model includes:
[0177] The weight matrix and the estimated reward are input into a pre-built robot action determination model, by Obtaining a first robot action output by the robot action determination model;
[0178] Among them, π * represents the first robot action, π∈Π means the robot action π belongs to the robot action set Π, U max represents the maximum reward, Represents the estimated reward, indicates that the user action τ belongs to the user action set Τ, β(τ) represents the weight matrix of the user action prediction model, P(U + |τ,π) represents the functional relationship between the improved reward U+ and the user action τ and the robot action π.
[0179] In some exemplary embodiments, updating the weight matrix according to the execution reward includes:
[0180] pass updating the weight matrix according to the execution reward;
[0181] Among them, β k (τ) represents the weight matrix of round k, β k-1 (τ) represents the weight matrix of the k-1 round, P(u k |τ,π k ) represents the execution reward u of round k k With user action τ and robot action π in round k k The functional relationship between It means that the user action τ′ belongs to the user action set T.
[0182] For the convenience of description, the above device is described by dividing it into various modules according to its functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0183] The device of the above embodiment is used to implement the corresponding robot control method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0184] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the robot control method described in any of the above embodiments is implemented.
[0185] Fig.15 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.
[0186] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0187] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0188] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0189] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0190] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0191] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0192] The electronic device of the above embodiment is used to implement the corresponding robot control method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0193] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the robot control method described in any of the above embodiments.
[0194] The above-mentioned non-transitory computer-readable storage medium can be any available medium or data storage device that can be accessed by a computer, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)), etc.
[0195] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the robot control method described in any embodiment in the above exemplary method part, and have the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0196] Those skilled in the art will appreciate that the embodiments of the present disclosure may be implemented as a system, method, or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." In addition, in some embodiments, the present disclosure may also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.
[0197] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive examples) of computer-readable storage media may include, for example: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.
[0198] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0199] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0200] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0201] It should be understood that each box in the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine, and these computer program instructions are executed by a computer or other programmable data processing device to produce a device that implements the functions / operations specified in the boxes in the flowchart and / or block diagram.
[0202] These computer program instructions may also be stored in a computer-readable medium that enables a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable medium produce a product that includes an instruction device that implements the functions / operations specified in the blocks in the flowchart and / or block diagram.
[0203] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby enabling the instructions executed on the computer or other programmable device to provide a process for implementing the functions / operations specified in the blocks in the flowchart and / or block diagram.
[0204] In addition, although the operations of the disclosed method are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all the operations shown must be performed to achieve the desired results. On the contrary, the steps depicted in the flow chart can be performed in a different order. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step, and / or one step can be decomposed into multiple steps.
[0205] The use of the verbs "comprise", "include" and their conjugations mentioned in the application documents does not exclude the presence of elements or steps other than those recorded in the application documents. The article "a" or "an" before an element does not exclude the presence of a plurality of such elements.
[0206] Although the spirit and principle of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims. The scope of the attached claims conforms to the broadest interpretation, thereby including all such modifications and equivalent structures and functions.
Claims
1. A robot control method, It is characterized in that The method comprises: Obtaining a weight matrix of a pre-built user action prediction model, and determining an estimated reward for controlling the robot to perform a robot action based on the weight matrix; wherein the user action includes a movement operation for a virtual character controlled by the user; and the robot action includes a movement operation for a virtual character controlled by the robot; Inputting the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model; wherein the first robot action can improve the estimated reward; controlling the robot to execute the first robot action, obtaining an execution reward and local observation information, and updating the weight matrix according to the execution reward; Determine a user action determination model from a pre-constructed user action determination model library according to the first robot action; input the local observation information into the user action determination model to obtain a reconstructed user action output by the user action determination model; in a pre-constructed basic environment, update the local observation information according to the current user action to obtain first updated local observation information; in a pre-constructed parallel environment, update the local observation information according to the reconstructed user action to obtain second updated local observation information; calculate the planning similarity according to the first updated local observation information and the second updated local observation information; update the weight matrix according to the planning similarity; Determine a user action determination model from a pre-constructed user action determination model library according to the first robot action; input the local observation information into the user action determination model to obtain a reconstructed user action output by the user action determination model; determine a current user action, and determine a reconstruction accuracy according to the current user action and the reconstructed user action; update the weight matrix according to the reconstruction accuracy; The user action is predicted using the user action prediction model with the updated weight matrix to obtain a first predicted user action, the first predicted user action is input into the robot action determination model to obtain a second robot action output by the robot action determination model, and the robot is controlled to perform the second robot action.
2. The method according to claim 1, It is characterized in that The determining an estimated reward corresponding to controlling the robot to perform a robot action based on the weight matrix includes: pass Determining an estimated reward corresponding to controlling the robot to perform a robot action based on the weight matrix; in, represents the estimated reward, π∈Π means the robot action π belongs to the robot action set Π, It indicates that the user action τ belongs to the user action set T, β(τ) represents the weight matrix of the user action prediction model, and E[U|τ,π] represents the functional relationship between the reward U and the user action τ and the robot action π.
3. The method according to claim 1, It is characterized in that The step of inputting the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model comprises: The weight matrix and the estimated reward are input into a pre-built robot action determination model, by Obtaining a first robot action output by the robot action determination model; Among them, π * represents the first robot action, π∈Π means the robot action π belongs to the robot max Action set Π, U represents the maximum reward, Represents the estimated reward, indicates that the user action τ belongs to the user action set Τ, β(τ) represents the weight matrix of the user action prediction model, P(U + |τ,π) represents the increased reward U + The functional relationship between the user action τ and the robot action π.
4. The method according to claim 1, It is characterized in that The updating of the weight matrix according to the execution reward comprises: pass updating the weight matrix according to the execution reward; Among them, β k (τ) represents the weight matrix of round k, β k-1 (τ) represents the weight matrix of the k-1 round, P(u k |τ,π k ) represents the execution reward u of round k k With user action τ and robot action π in round k k The functional relationship between It means that the user action τ′ belongs to the user action set T.
5. A robot control device, It is characterized in that The device comprises: A user action prediction model determination module is configured to obtain a weight matrix of a pre-built user action prediction model and determine an estimated reward for controlling the robot to perform a robot action based on the weight matrix; wherein the user action includes a movement operation for a virtual character controlled by the user; and the robot action includes a movement operation for a virtual character controlled by the robot; A robot action determination module is configured to input the weight matrix and the estimated reward into a pre-built robot action determination model to obtain a first robot action output by the robot action determination model; wherein the first robot action can improve the estimated reward; a user action prediction model updating module, configured to control the robot to execute the first robot action, obtain an execution reward and local observation information, and update the weight matrix according to the execution reward; The user action prediction model updating module is further configured to determine a user action determination model from a pre-constructed user action determination model library according to the first robot action; input the local observation information into the user action determination model to obtain a reconstructed user action output by the user action determination model; in a pre-constructed basic environment, update the local observation information according to the current user action to obtain first updated local observation information; in a pre-constructed parallel environment, update the local observation information according to the reconstructed user action to obtain second updated local observation information; calculate the planning similarity according to the first updated local observation information and the second updated local observation information; and update the weight matrix according to the planning similarity; The user action prediction model updating module is further configured to determine a user action determination model from a pre-constructed user action determination model library according to the first robot action; input the local observation information into the user action determination model to obtain a reconstructed user action output by the user action determination model; determine a current user action, and determine a reconstruction accuracy according to the current user action and the reconstructed user action; and update the weight matrix according to the reconstruction accuracy; The robot action update module is configured to predict the user action using the user action prediction model with the weight matrix updated to obtain a first predicted user action, input the first predicted user action into the robot action determination model to obtain a second robot action output by the robot action determination model, and control the robot to perform the second robot action.
6. An electronic device, It is characterized in that The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 4 is implemented.
7. A non-transitory computer-readable storage medium, It is characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 4.