Robot intelligent decision control method based on adversarial imitation learning
Through the adversarial imitation learning method, the visual language pre-trained model is used to extract semantic features and build a teacher model group. Combined with knowledge distillation technology, the problem of insufficient decision-making ability in a dynamic changing environment is solved, and higher diversity of decision-making strategies and generalization ability are achieved.
Patent Information
- Application Number
- CN202510544847.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-06-24
AI Technical Summary
Traditional robot intelligent decision-making control methods are difficult to adapt to the dynamic changes of new situations and emergencies, and the behavioral cloning method has poor generalization ability when facing unoccupied states.
A robot intelligent decision-making control method based on adversarial imitation learning is proposed. By obtaining expert demonstration data, semantic features are extracted using visual language pre-trained models, a teacher model group is constructed, and the knowledge of multiple teacher models is transferred to a single student model through knowledge distillation technology.
It significantly improves the diversity and generalization capabilities of robotic intelligent decision-making strategies, and can handle complex and dynamic environments more effectively.
Smart Images

Figure CN120190823A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of automatic control technology, and in particular, to a robot intelligent decision-making control method based on adversarial imitation learning. Background Art
[0002] In the field of robotics, since robots need to make accurate decisions in real time in complex and dynamic environments, the research and application of robot intelligent decision-making control methods are particularly important. However, traditional robot intelligent decision-making control methods usually rely on expert data and preset rules, and it is difficult to adapt to the dynamic changes of new situations and emergencies.
[0003] To solve the above problems, imitation learning has been proposed as a new method for robot intelligent decision-making control in related technologies. Imitation learning learns decision-making control strategies from expert demonstrations, avoiding the dependence on preset rules and reward functions in traditional methods. Among them, behavior cloning is a classic method of imitation learning. It regards the behavior learning problem as a supervised learning task and realizes the replication of expert behaviors by establishing a mapping between environmental states and expert actions. However, behavior cloning can only simply copy the state-action pairs in expert data, and the model cannot make reasonable decisions in the face of states not appearing in expert data, resulting in poor generalization ability. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to propose a robot intelligent decision-making control method based on adversarial imitation learning, aiming to improve the diversity and generalization ability of robot intelligent decision-making strategies.
[0005] To achieve the above object, the first aspect of the embodiments of this application proposes a robot intelligent decision-making control method based on adversarial imitation learning, and the method includes:
[0006] Obtain the expert demonstration data of the robot, and perform semantic feature extraction on the image data corresponding to the expert demonstration data based on a vision-language pre-training model (Contrastive Language-Image Pretraining, CLIP) to obtain the semantic features of each trajectory data in the expert demonstration data;
[0007] Construct a group of teacher models based on the semantic features and the generative adversarial imitation learning strategy;
[0008] Obtain the distillation loss between multiple teacher models and a single student model in the group of teacher models;
[0009] Transfer the knowledge of the multiple teacher models to the single student model based on the distillation loss, and perform robot intelligent decision-making control based on the student model.
[0010] In some embodiments, the image data corresponding to the expert demonstration data includes the image data corresponding to each piece of trajectory data in the expert demonstration data;
[0011] The semantic feature extraction of the image data corresponding to the expert demonstration data by the vision-language pre-trained model to obtain the semantic features of each piece of trajectory data in the expert demonstration data includes:
[0012] Based on the vision-language pre-trained model, perform semantic feature extraction on the image data corresponding to the target trajectory data in the expert demonstration data to obtain the average value of the trajectory features of the target trajectory data, and use the average value of the trajectory features as the semantic feature of the target trajectory data;
[0013] Wherein, the target trajectory data is any one of the multiple pieces of trajectory data in the expert demonstration data.
[0014] In some embodiments, the method further includes:
[0015] Obtain the environmental dimension information associated with the expert demonstration data;
[0016] Based on the environmental dimension information and the action information of each trajectory in the expert demonstration data, perform image rendering processing on each piece of trajectory data to obtain the image data corresponding to each piece of trajectory data.
[0017] In some embodiments, the construction of the teacher model group based on the semantic features and the generative adversarial imitation learning strategy includes:
[0018] Divide the semantic features into multiple cluster sets;
[0019] Independently train each cluster set in the multiple cluster sets based on the generative adversarial imitation learning strategy to obtain a teacher model corresponding to each cluster set; the teacher model group includes the teacher models corresponding to each cluster set.
[0020] In some embodiments, the obtaining of the distillation loss between multiple teacher models in the teacher model group and a single student model includes:
[0021] Obtain the teacher actions generated by multiple teacher models in the teacher model group, and obtain the student actions generated by a single student model;
[0022] Calculate the distillation loss between each of the multiple teacher models and the single student model based on the teacher actions and the student actions.
[0023] In some embodiments, the obtaining of the student actions generated by a single student model includes:
[0024] Obtain the data generated by the interaction between the robot and the environment;
[0025] Initialize the single student model based on the data generated by the interaction between the robot and the environment, and obtain the student actions generated by the single student model.
[0026] In some embodiments, the transferring the knowledge of the multiple teacher models to the single student model based on the distillation loss includes:
[0027] Optimize the policy network parameters of the single student model based on the distillation loss and the reward signal of a preset discriminator, so as to transfer the knowledge of the multiple teacher models to the single student model;
[0028] Wherein, the discriminator generates a reward signal based on the expert demonstration data and the student actions generated by the single student model.
[0029] To achieve the above object, a second aspect of the embodiments of the present application proposes a robot intelligent decision-making control device based on adversarial imitation learning, and the device includes:
[0030] A data acquisition and feature extraction module, configured to acquire expert demonstration data of the robot, and perform semantic feature extraction on the image data corresponding to the expert demonstration data based on a vision-language pre-training model, so as to obtain the semantic features of each trajectory data in the expert demonstration data;
[0031] An adversarial imitation learning module, configured to construct a group of teacher models based on the semantic features and an adversarial imitation learning strategy;
[0032] A knowledge distillation module, configured to obtain the distillation loss between multiple teacher models and a single student model in the group of teacher models; and transfer the knowledge of the multiple teacher models to the single student model based on the distillation loss, and perform robot intelligent decision-making control based on the student model.
[0033] To achieve the above object, a third aspect of the embodiments of the present application further proposes a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the robot intelligent decision-making control method based on adversarial imitation learning described in the first aspect above.
[0034] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the robot intelligent decision-making control method based on adversarial imitation learning described in the first aspect above.
[0035] To achieve the above object, a fifth aspect of the embodiments of the present application proposes a computer program product. The computer program product stores a computer program, and when the computer program is executed by a processor, it implements the robot intelligent decision-making control method based on adversarial imitation learning described in the first aspect above.
[0036] The robot intelligent decision-making control method based on adversarial imitation learning, the robot intelligent decision-making control device based on adversarial imitation learning, the computer device, the computer-readable storage medium, and the computer program product proposed by the embodiments of the present application obtain expert demonstration data of the robot, and based on a vision-language pre-training model, extract semantic features from the image data corresponding to the expert demonstration data to obtain the semantic features of each trajectory data in the expert demonstration data; construct a teacher model group based on the semantic features and a generative adversarial imitation learning strategy; obtain the distillation loss between multiple teacher models and a single student model in the teacher model group; and based on the distillation loss, transfer the knowledge of the multiple teacher models to the single student model, and perform robot intelligent decision-making control based on the student model.
[0037] In this way, compared with the method of robot intelligent decision-making control based on simple behavior cloning in the related art, the embodiments of the present application perform image rendering processing on the collected expert demonstration data, use CLIP to extract semantic features, then train an independent generative adversarial imitation learning strategy based on the extracted semantic features to construct a teacher model group, and transfer the knowledge of multiple teacher models to a single student model through knowledge distillation technology. In this way, the limitations of a single strategy model can be effectively overcome based on the multi-teacher knowledge distillation mechanism, thereby significantly improving the diversity and generalization ability of the robot intelligent decision-making strategy.
[0038] In addition, by combining the two technical means of image rendering processing and CLIP feature extraction, the embodiments of the present application can also achieve efficient representation of multi-modal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic flow chart of the steps of the robot intelligent decision-making control method based on adversarial imitation learning provided by the embodiments of the present application in some embodiments;
[0040] Figure 2 It is a schematic flow chart of the steps of the robot intelligent decision-making control method based on adversarial imitation learning provided by the embodiments of the present application in other embodiments;
[0041] Figure 3 For Figure 1 It is a schematic flow chart of the refined steps of step S102 in
[0042] Figure 4For Figure 1 Schematic diagram of the refined step process of step S103 in
[0043] Figure 5 Flowchart of the robot intelligent decision-making control method based on adversarial imitation learning provided by an embodiment of the present application in a complete embodiment;
[0044] Figure 6 Flowchart of the generative adversarial imitation learning method involved in the robot intelligent decision-making control method based on adversarial imitation learning provided by an embodiment of the present application in a complete embodiment;
[0045] Figure 7 Flowchart of the knowledge distillation method involved in the robot intelligent decision-making control method based on adversarial imitation learning provided by an embodiment of the present application in a complete embodiment;
[0046] Figure 8 Screenshot of the visual graphical interface of three experimental tasks of the robot intelligent decision-making control method based on adversarial imitation learning proposed by an embodiment of the present application;
[0047] Figure 9 Schematic diagram of the comparison curve of the robot intelligent decision-making control method based on adversarial imitation learning proposed by an embodiment of the present application with other methods in the HalfCheetah physical environment;
[0048] Figure 10 Schematic diagram of the comparison curve of the robot intelligent decision-making control method based on adversarial imitation learning proposed by an embodiment of the present application with other methods in the Hopper physical environment,
[0049] Figure 11 Schematic diagram of the comparison curve of the robot intelligent decision-making control method based on adversarial imitation learning proposed by an embodiment of the present application with other methods in the Walker physical environment;
[0050] Figure 12 Schematic diagram of the structure of the robot intelligent decision-making control device based on adversarial imitation learning provided by an embodiment of the present application;
[0051] Figure 13 Schematic diagram of the hardware structure of the computer device provided by an embodiment of the present application. Detailed implementation manners
[0052] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0053] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the description, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0055] First, a brief explanation of the professional technical terms involved in the embodiments of this application is given.
[0056] Intelligent decision-making control:
[0057] Intelligent decision-making control refers to a technology that analyzes and processes a large amount of data through artificial intelligence technology and intelligent algorithms to obtain the optimal decision-making result. The core of intelligent decision-making control lies in integrating advanced algorithms such as machine learning, deep learning, and reinforcement learning, imitating the human decision-making process, and realizing the automation and intelligence of decision-making.
[0058] Next, the overall concept of the embodiments of this application is described.
[0059] In the field of robotics, since robots need to make accurate decisions in complex and dynamic environments in real time, the research and application of robot intelligent decision-making control methods are particularly important. However, traditional robot intelligent decision-making control methods usually rely on expert data and preset rules and are difficult to adapt to the dynamic changes of new situations and emergencies. Secondly, the scalability and computational efficiency of the models in the related technologies in the high-dimensional state space are relatively low, and it is difficult to balance the real-time performance and accuracy of decision-making at the same time.
[0060] To solve the above problems, imitation learning provides a new method for intelligent decision-making control of robots. Imitation learning learns decision-making control strategies from expert demonstrations, avoiding the dependence on preset rules and reward functions in traditional methods. Behavioral cloning is a classic method of imitation learning. It regards the behavior learning problem as a supervised learning task and replicates expert behavior by establishing a mapping between environmental states and expert actions. However, behavioral cloning can only simply copy the state-action pairs in expert data. Facing states not present in expert data, the model cannot make reasonable decisions, resulting in poor generalization ability. Inverse reinforcement learning is another imitation learning method. Its core is to infer a potential reward function from expert demonstrations and use this reward function to train the policy. However, the reward function inferred by inverse reinforcement learning is overly dependent on expert data, difficult to generalize to new environments, and the process of setting the reward function is complex with high computational costs, limiting its widespread use in practical applications.
[0061] Generative adversarial imitation learning combines the advantages of generative adversarial networks and imitation learning. Through the adversarial game between the generator and the discriminator, it directly learns the policy from expert data. The generator attempts to generate a policy that matches the expert data, while the discriminator distinguishes between the generated data and the real expert data. Through this adversarial mechanism, generative adversarial imitation learning can effectively avoid the complexity of reward function design while improving the generalization ability of the policy. However, when dealing with high-dimensional and multi-modal data, generative adversarial imitation learning often has difficulty effectively capturing the complex relationships of state-action sequences, and a single policy model has difficulty adapting to diverse decision-making needs when facing complex tasks.
[0062] Based on this, the embodiments of this application propose a robot intelligent decision-making control method based on adversarial imitation learning, aiming to achieve efficient representation of multi-modal data and improve the diversity and generalization ability of decision-making strategies.
[0063] In some embodiments, the robot intelligent decision-making control method based on adversarial imitation learning proposed in the embodiments of this application first performs image rendering processing on the collected expert demonstration data and uses the vision-language pre-training model CLIP to extract semantic features; secondly, the clustering algorithm Kmeans is used to divide the extracted semantic feature data into multiple cluster sets, and an independent generative adversarial imitation learning strategy is trained for each cluster set to construct a group of teacher models; then, through knowledge distillation technology, the knowledge of multiple teacher models is transferred to a single student model, and the discriminator reward signal is fused to optimize the student policy. In this way, through the combination of image rendering and CLIP feature extraction, efficient representation of multi-modal data is achieved; and by adopting the multi-teacher knowledge distillation mechanism, the limitations of a single policy model are effectively overcome, significantly improving the diversity and generalization ability of decision-making strategies.
[0064] The robot intelligent decision-making control method based on adversarial imitation learning proposed in the embodiments of this application can be widely applied to the field of robot intelligent decision-making control and has important practical value.
[0065] Next, based on the overall concept of the above embodiments of this application, specific embodiments of the robot intelligent decision-making control method based on adversarial imitation learning, the robot intelligent decision-making control device based on adversarial imitation learning, the computer device, the computer-readable storage medium, and the computer program product provided by the embodiments of this application are proposed. First, each specific embodiment of the robot intelligent decision-making control method in the embodiments of this application is described in detail.
[0066] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0067] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0068] It should be noted that in each specific implementation manner of this application, when it comes to relevant processing that needs to be carried out based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.
[0069] In addition, the robot intelligent decision-making control method based on adversarial imitation learning provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on the terminal or the server side. In some embodiments, the terminal can be a terminal device capable of controlling the operation of the robot, such as a smart phone, a tablet computer, a laptop computer, a desktop computer, etc. In some embodiments, the terminal can also be the robot itself; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the robot intelligent decision-making control method based on adversarial imitation learning, etc., but is not limited to the above forms.
[0070] Alternatively, the embodiments of the present application can also be used in many general-purpose or special-purpose computer system environments or configurations. For example: robots, robot systems, personal computers, server computers, handheld devices or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0071] For the convenience of understanding and elaboration, in the following text, the terminal device applying the robot intelligent decision-making control method based on adversarial imitation learning provided by the embodiments of the present application will be taken as an example for detailed description. The implementation of the robot intelligent decision-making control method based on adversarial imitation learning provided by the embodiments of the present application for any of the above forms of the subject can refer to the process of the terminal device applying the robot intelligent decision-making control method based on adversarial imitation learning described later.
[0072] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the step flow of the robot intelligent decision-making control method based on adversarial imitation learning provided by the embodiments of the present application in some embodiments. It should be understood that although Figure 1The execution order of some method steps is shown, but based on different design requirements of actual applications, the robot intelligent decision-making control method provided by the embodiments of the present application can of course adopt an execution order different from that shown in the figure. That is, Figure 1 The order of the shown method steps does not constitute a limitation on the execution logic order of the robot intelligent decision-making control method provided by the embodiments of the present application. Any other Figure 1 reasonable changes to the shown method step order should be included within the protection scope of the robot intelligent decision-making control method provided by the embodiments of the present application.
[0073] For example, Figure 1 as shown, in some embodiments, the robot intelligent decision-making control method provided by the embodiments of the present application may include, but is not limited to, steps S101 to S104.
[0074] Step S101: Obtain the expert demonstration data of the robot, and perform semantic feature extraction on the image data corresponding to the expert demonstration data based on the vision-language pre-training model to obtain the semantic features of each trajectory data in the expert demonstration data.
[0075] Before the terminal device performs robot intelligent decision-making control, it first obtains the expert demonstration data of the robot, and then performs image rendering processing on the expert demonstration data, so as to render the expert demonstration data into an image (image data corresponding to the expert demonstration data). Then, the vision-language pre-training model CLIP is used to perform semantic feature extraction on the image data, so as to obtain the semantic features of each trajectory data in the expert demonstration data.
[0076] It should be noted that in the expert demonstration data obtained by the terminal device, each trajectory data is composed of a series of consecutive states and actions. For example, the expert demonstration data is {obs:[s1,s2,…,s T ,action:[a1,a2,…,a T}, where s t is the state at the t-th step, and a t is the action at the t-th step.
[0077] Step S102: Construct a teacher model group based on the semantic features and the generative adversarial imitation learning strategy.
[0078] After obtaining the semantic features of each trajectory data in the expert demonstration data, the terminal device can train independent Generative Adversarial Imitation Learning (GAIL) strategies based on the semantic features respectively, so as to construct a group of teacher models including multiple teacher models.
[0079] It should be noted that the Generative Adversarial Imitation Learning (GAIL) strategy can also be called adversarial generative imitation learning. The terminal device trains independent teacher models based on training independent Generative Adversarial Imitation Learning strategies, which can capture the characteristics of different decision-making modes. Since each of the multiple teacher models in the group of teacher models is obtained based on training independent Generative Adversarial Imitation Learning strategies, the knowledge of each of the multiple teacher models is different.
[0080] In some embodiments, when the terminal device trains an independent teacher model based on training an independent Generative Adversarial Imitation Learning strategy, the Generative Adversarial Imitation Learning (GAIL) strategy may include, but is not limited to, the following steps S1 to S6.
[0081] S1: Sample state-action pairs from the expert demonstration data The agent (robot) interacts with the environment to generate state-action pairs
[0082] S2: Initialize the parameters of the policy network π θ , discriminator network D φ and value network parameters;
[0083] S3: Calculate the discriminator loss Update the discriminator parameters φ;
[0084] Among them, the goal of the discriminator D φ is to distinguish expert behavior from agent behavior, and its loss function is binary cross-entropy loss:
[0085]
[0086] Exemplarily, the discriminator parameters φ can be updated by gradient descent:
[0087]
[0088] Among them, η D is the learning rate of the discriminator.
[0089] S4: Calculate the policy network loss Update the policy parameters θ;
[0090] Among them, the policy network π θThe goal is to maximize the reward signal provided by the discriminator:
[0091]
[0092] Exemplarily, the trust region policy optimization (TRPO) algorithm can be used to update the policy parameter θ:
[0093]
[0094] Among them, η π is the learning rate of the policy network.
[0095] S5: Calculate the value network loss Update the value network parameter ψ;
[0096] Among them, the value network aims to accurately estimate the state value function, and its loss error is the mean square error loss:
[0097]
[0098] Among them, R(s) is the cumulative discounted reward of state s.
[0099] Exemplarily, the value network parameter can be updated by gradient descent :
[0100]
[0101] Among them, η V is the learning rate of the value network.
[0102] S6: Repeat steps S2 - S5 until the policy network converges.
[0103] Step S103: Obtain the distillation loss between multiple teacher models in the teacher model group and a single student model.
[0104] After the terminal device constructs the teacher model group, for multiple teacher models in the teacher model group, calculate the distillation loss between each of the multiple teacher models and the single student model respectively.
[0105] Step S104: Transfer the knowledge of the multiple teacher models to the single student model based on the distillation loss, and perform robot intelligent decision control based on the student model.
[0106] After the terminal device calculates the distillation losses between each of multiple teacher models and a single student model, it uses knowledge distillation technology to transfer the knowledge of the multiple teacher models to the single student model based on each distillation loss. Then, the terminal device can perform robot intelligent decision-making control based on this student model.
[0107] In some embodiments, after the terminal device transfers the knowledge of multiple teacher models to a single student model through knowledge distillation technology, it can also determine whether the student model has learned well. If the student model has learned well, the training ends, and robot intelligent decision-making control is performed based on this student model. If the student model has not learned well, the strategy of the student model is updated and the foregoing steps are repeated until it is determined that the student model has learned well.
[0108] It should be noted that the terminal device uses knowledge distillation technology to transfer the knowledge of multiple teacher models to a single student model, thereby improving the generalization ability and decision-making diversity of the student model. Among them, the knowledge distillation technology may include, but is not limited to, the following steps S1 to S3.
[0109] S1: Obtain the action predictions of each teacher model on the expert trajectory from multiple teacher models;
[0110] S2: Calculate the action prediction differences between the student model and each teacher model on the expert trajectory;
[0111] S3: Update the policy network parameters of the student model by accumulating the distillation losses of multiple teacher models.
[0112] In the embodiments of the present application, before the terminal device performs robot intelligent decision-making control, it first obtains the expert demonstration data of the robot, performs image rendering processing on the expert demonstration data to obtain the image data corresponding to the expert demonstration data, and then uses the vision-language pre-training model CLIP to extract semantic features from the image data, so as to obtain the semantic features of each trajectory data in the expert demonstration data. For this semantic feature, the terminal device trains an independent generative adversarial imitation learning strategy GAIL respectively to construct a teacher model group including multiple teacher models. For the multiple teacher models in the teacher model group, the terminal device calculates the distillation losses between each of the multiple teacher models and a single student model respectively. Then, it uses knowledge distillation technology to transfer the knowledge of the multiple teacher models to the single student model based on each distillation loss. In this way, the terminal device can perform robot intelligent decision-making control based on this student model.
[0113] Compared with the method of robot intelligent decision-making control based on simple behavior cloning in the related art, in the embodiments of the present application, the collected expert demonstration data is subjected to image rendering processing, and the visual language pre-trained model CLIP is used to extract semantic features. Then, an independent generative adversarial imitation learning strategy is trained based on the extracted semantic features to construct a group of teacher models, and the knowledge of multiple teacher models is transferred to a single student model through the knowledge distillation technology. In this way, the limitations of a single policy model can be effectively overcome based on the multi-teacher knowledge distillation mechanism, thereby significantly improving the diversity and generalization ability of the robot intelligent decision-making strategy.
[0114] In addition, by combining the two technical means of image rendering processing and CLIP feature extraction, the embodiments of the present application can also achieve efficient representation of multi-modal data.
[0115] In some embodiments, the image data corresponding to the expert demonstration data may be the image data corresponding to each trajectory data in the expert demonstration data. In this case, in step S101 above, the step of "extracting semantic features from the image data corresponding to the expert demonstration data based on the visual language pre-trained model to obtain the semantic features of each trajectory data in the expert demonstration data" may include the following steps:
[0116] Based on the visual language pre-trained model, extract semantic features from the image data corresponding to the target trajectory data in the expert demonstration data to obtain the mean value of the trajectory features of the target trajectory data, and use the mean value of the trajectory features as the semantic features of the target trajectory data;
[0117] Wherein, the target trajectory data is any one of the multiple trajectory data in the expert demonstration data.
[0118] When the terminal device uses CLIP to extract semantic features from the image data corresponding to the expert demonstration data, it can use this CLIP to extract the mean value of the trajectory features of the image data corresponding to each target trajectory data in the expert demonstration data respectively, and then use the mean value of the semantic features as the semantic features of each target trajectory data.
[0119] In this embodiment, the terminal device uses CLIP to extract the mean value of the semantic features of each trajectory, which can enhance the data representation ability.
[0120] In some embodiments, the terminal device may perform image rendering processing on each trajectory data in the expert demonstration data before using CLIP to extract the mean value of the trajectory features of the image data corresponding to each target trajectory data in the expert demonstration data, so as to obtain the image data corresponding to each trajectory data.
[0121] Please refer toFigure 2 , Figure 2 It is a schematic flowchart of steps in other embodiments of the robot intelligent decision-making control method based on adversarial imitation learning provided by the embodiments of the present application.
[0122] As Figure 2 shown, in some embodiments, the robot intelligent decision-making control method based on adversarial imitation learning provided by the embodiments of the present application may further include the following steps S201 and S202.
[0123] Step S201: Obtain environmental dimension information associated with the expert demonstration data.
[0124] When the terminal device performs image rendering processing on each piece of trajectory data in the expert demonstration data, it can initialize the simulation environment of the robot (such as three environments: HalfCheetah, Hopper, and Walker), so as to obtain the corresponding environmental dimension information and prepare for setting the environmental state when performing image rendering processing subsequently. And the terminal device uses this environmental dimension information as the environmental dimension information associated with the expert demonstration data of the robot.
[0125] It should be noted that the environmental dimension information obtained by the terminal device may include: position dimension nq (such as joint angle) and speed dimension nv (such as joint angular velocity)
[0126] Step S202: Based on the environmental dimension information and the action information of each trajectory in the expert demonstration data, perform image rendering processing on each piece of trajectory data to obtain the image data corresponding to each piece of trajectory data.
[0127] After the terminal device obtains the environmental dimension information, it combines this environmental dimension information with the action information of each trajectory in the expert demonstration data, sets the environmental state and executes the action respectively to render the current frame image, so as to realize the image rendering processing of each piece of trajectory data and obtain the image data corresponding to each piece of trajectory data.
[0128] Exemplarily, for the expert demonstration data, the terminal device can traverse each piece of trajectory data in the expert demonstration data, that is: the input is all the trajectory data in the expert demonstration data, traverse each piece of trajectory data and display a progress bar. During this process, for each piece of trajectory data, the terminal device initializes a feature accumulator, then traverses each step action in the trajectory data, and uses the environmental dimension information to set the environmental state, execute the action, and render the current frame image, so that the image data corresponding to each piece of trajectory data can be obtained.
[0129] Moreover, after rendering the current frame image, the terminal device can use the CLIP model to extract image features, and then calculate the mean value of the trajectory features after accumulating the features. The mean value of the trajectory features can be regarded as the "overall feature" of the trajectory (i.e., the aforementioned semantic feature).
[0130] Please refer to Figure 3 , Figure 3 for Figure 1 the detailed step flow diagram of step S102 in
[0131] As Figure 3 shown, in some embodiments, the above step S102: constructing a teacher model group based on the semantic feature and the generative adversarial imitation learning strategy may include, but is not limited to, step S301 and step S302 shown below.
[0132] Step S301: Divide the semantic feature into multiple cluster sets.
[0133] Step S302: Independently train each cluster set in the multiple cluster sets based on the generative adversarial imitation learning strategy to obtain a teacher model corresponding to each cluster set; the teacher model group includes the teacher models corresponding to each cluster set.
[0134] When the terminal device trains an independent generative adversarial imitation learning strategy based on the semantic feature to construct a teacher model group, it can first perform clustering analysis on the semantic feature, that is, use the clustering algorithm Kmeans to divide the semantic feature data into multiple cluster sets. Then, the terminal device trains an independent generative adversarial imitation learning strategy for each cluster set respectively, so as to train a teacher model corresponding to each cluster set, and the teacher models corresponding to each cluster set constitute a teacher model group.
[0135] It should be noted that after the terminal device uses the clustering algorithm Kmeans to divide the extracted semantic feature into multiple cluster sets, each cluster set in the multiple cluster sets represents a specific decision-making mode. The KMeans algorithm iteratively optimizes the position of the cluster center u k , calculates the distance between the semantic feature of each trajectory data and the cluster center, and assigns it to the cluster with the closest distance, and finally returns the cluster label of each trajectory;
[0136]
[0137] where C k is the kth cluster, and u k is the center of the kth cluster.
[0138] Please refer to Figure 4 , Figure 4 for Figure 1 the detailed step flow diagram of step S103 in
[0139] As Figure 4 shown, in some embodiments, step S103 above: obtaining the distillation loss between multiple teacher models in the teacher model group and a single student model may include, but is not limited to, steps S401 and S402 shown below.
[0140] Step S401: Obtain the teacher actions generated by each of the multiple teacher models in the teacher model group, and obtain the student actions generated by a single student model.
[0141] When the terminal device obtains the distillation loss between each of the multiple teacher models in the teacher model group and a single student model, it can load the multiple teacher models respectively to obtain the teacher actions generated by each of the multiple teacher models.
[0142] Exemplarily, the terminal device can generate teacher actions for each teacher model by loading the policy network value network and discriminator of each teacher model as follows:
[0143]
[0144] where a i is the action generated by the i-th teacher model.
[0145] While the terminal device obtains the teacher actions generated by each of the multiple teacher models, the terminal device can also initialize the student model to obtain the student actions generated by the student model.
[0146] In some embodiments, the step of "obtaining the student actions generated by a single student model" in step S401 above may include, but is not limited to, the steps described below:
[0147] Obtain the data generated by the interaction between the robot and the environment;
[0148] Based on the data generated by the interaction between the robot and the environment, perform initialization processing on the single student model to obtain the student actions generated by the single student model.
[0149] When the terminal device initializes the student model to generate student actions, it can first obtain the data generated by the interaction between the robot and the environment, and then perform initialization processing on the network parameters of the student model based on the data generated by the interaction between the robot and the environment, so as to obtain the student actions generated by the student model.
[0150] Exemplarily, the terminal device can initialize the policy network of the student model with the data generated by the interaction between the robot and the environment Value network and discriminator to obtain the student action where a s is the action generated by the student model.
[0151] Step S402: Calculate the distillation loss between each of the multiple teacher models and the single student model based on the teacher action and the student action.
[0152] After obtaining the teacher action and the student action, the terminal device can calculate the distillation loss between each of the multiple teacher models and the student model based on the Mean Square Error (MSE) loss function, with the teacher action and the student action as parameters.
[0153] Exemplarily, the terminal device can use the student action a s and the teacher action a i as inputs and calculate the MSE distillation loss between the student model and the teacher model according to the following formula:
[0154]
[0155] where N represents the number of teacher models.
[0156] In some embodiments, when calculating the MSE distillation loss between the student model and the teacher model, the terminal device can also update the policy network parameters θ s of the student model through backpropagation and an optimizer, satisfying:
[0157]
[0158] where η s is the learning rate of the student model.
[0159] In some embodiments, when the terminal device uses the knowledge distillation technique to transfer the knowledge of multiple teacher models to a single student model, it can also optimize the student policy by integrating the discriminator reward signal.
[0160] It should be noted that the reward signal of the preset discriminator is used to optimize the policy of the student model. The input of the discriminator is the student action generated by the student model and the expert demonstration data. The goal of the discriminator reward signal is to minimize the negative log probability of the discriminator for the student action:
[0161]
[0162] In this way, by combining the distillation loss and the discriminator reward signal, the student model can learn diverse behavioral patterns from multiple teacher models.
[0163] Based on this, the step of "transferring the knowledge of the multiple teacher models to the single student model based on the distillation loss" in the above step S104 may include, but is not limited to, the following steps:
[0164] Optimize the policy network parameters of the single student model based on the distillation loss and the reward signal of a preset discriminator to transfer the knowledge of the multiple teacher models to the single student model;
[0165] Wherein, the discriminator generates a reward signal based on the expert demonstration data and the student actions generated by the single student model.
[0166] After the terminal device calculates the distillation loss between the student model and the teacher models, it can adopt the knowledge distillation technology to optimize the policy network parameters of the single student model based on this distillation loss and by integrating the reward signal of the discriminator, so as to realize transferring the knowledge of the multiple teacher models to the single student model.
[0167] In some embodiments, after the terminal device transfers the knowledge of the multiple teacher models to the single student model, it can further optimize the policy for the student model during the interaction between the robot and the environment. For example, update the policy of the student model through the data generated by the interaction between the robot and the environment until the student model reaches a predetermined learning effect.
[0168] Next, a complete embodiment of the robot intelligent decision-making control method based on adversarial imitation learning provided by the embodiments of the present application is presented.
[0169] Please refer to Figure 5 , Figure 5 which is the flowchart of the robot intelligent decision-making control method based on adversarial imitation learning provided by the embodiments of the present application in a complete embodiment.
[0170] As Figure 5 shown, the robot intelligent decision-making control method based on adversarial imitation learning provided by the embodiments of the present application may include the following steps (1) to (8).
[0171] Step (1): Render the collected expert demonstration data into images, and use CLIP to extract the trajectory feature means of each trajectory data.
[0172] For step (1), it respectively includes the following steps (1.1) to (1.3).
[0173] Step (1.1): Load the expert demonstration data: The input is the expert demonstration data {obs: [s1, s2, …, s T , action: [a1, a2, …, aT} and load this expert demonstration data through pickle.load in the Python pickle module. These data contain multiple trajectories, and each trajectory consists of a series of consecutive states and actions, where s t is the state at the t-th step, and a t is the action at the t-th step.
[0174] Step (1.2): Initialize the environment: Initialize three environments, HalfCheetah, Hopper, and Walker, and obtain the dimension information of the environment, including the position dimension (nq, such as joint angles) and the velocity dimension (nv, such as joint angular velocities), to prepare for setting the environment state later.
[0175] Step (1.3): Traverse each trajectory: The input is all the trajectories in the expert demonstration data. Traverse each trajectory τ = {(s1, a1), (s2, a2), …, (s T , a T )} and display a progress bar. For each trajectory, initialize a feature accumulator, traverse each step t in the trajectory, set the environment state, execute the action, render the current frame image, and use the CLIP model to extract the image features. After accumulating the features, calculate the mean of the trajectory features, which can be regarded as the "overall feature" of the trajectory and is used for subsequent clustering analysis.
[0176] Step (2): Use the Kmeans clustering algorithm to cluster the overall features of the trajectories to generate multiple clusters. The KMeans algorithm iteratively optimizes the position of the cluster center u k and calculates the distance between each trajectory feature and the cluster center, and assigns it to the cluster with the closest distance, and finally returns the cluster label of each trajectory:
[0177]
[0178] where C k is the k-th cluster, and u k is the center of the k-th cluster.
[0179] Step (3): Train an independent generative adversarial imitation learning strategy for each cluster set respectively to construct a group of teacher models.
[0180] As Figure 6 shown, for Step (3), it respectively includes the following steps (3.1) to (3.6).
[0181] Step (3.1): Sample state-action pairs from the expert demonstration data and the agent interacts with the environment to generate state-action pairs
[0182] Step (3.2): Initialize the parameters of the policy network π θ , the discriminator network D φ and the value network ;
[0183] Step (3.3): Calculate the discriminator loss Update the discriminator parameters φ;
[0184] The goal of the discriminator D φ is to distinguish expert behavior from agent behavior, and its loss function is the binary cross-entropy loss:
[0185]
[0186] Update the discriminator parameters φ through gradient descent:
[0187]
[0188] where η D is the learning rate of the discriminator.
[0189] Step (3.4): Calculate the policy network loss Update the policy parameters θ using the TRPO algorithm;
[0190] The goal of the policy network π θ is to maximize the reward signal provided by the discriminator:
[0191]
[0192] Update the policy parameters θ using the TRPO algorithm:
[0193]
[0194] where η π is the learning rate of the policy network.
[0195] Step (3.5): Calculate the value network loss Update the value network parameters ψ;
[0196] The goal of the value network is to accurately estimate the state value function, and its loss error is the mean squared error loss:
[0197]
[0198] where R(s) is the cumulative discounted reward of state s.
[0199] Update the value network parameters through gradient descent :
[0200]
[0201] Among them, η V is the learning rate of the value network.
[0202] Step (3.6): Repeat steps (3.2) - (3.5) until the policy network converges.
[0203] Step (4): Load the policy networks of multiple teacher models the value network and the discriminator for each teacher model to generate teacher actions where a i is the action generated by the i-th teacher model.
[0204] Step (5): Initialize the policy network, the value network and the discriminator of the student model with the data generated by interacting with the environment to generate student actions where a is the action generated by the student model. s is the action generated by the student model.
[0205] Step (6): Input the student action a s and the teacher action a i , and calculate the MSE distillation loss between the student model and the teacher model;
[0206]
[0207] where N represents the number of teacher models.
[0208] Step (7): Transfer the knowledge of multiple teacher models to a single student model through knowledge distillation technology, and optimize the student policy by integrating the reward signal of the discriminator. First, update the student policy parameter θ s :
[0209]
[0210] where η s is the learning rate of the student model. In addition, the reward signal r(s,a) provided by the discriminator is also used to optimize the policy, and its goal is to minimize the negative log probability of the discriminator for the student action:
[0211]
[0212] By combining the distillation loss and the discriminator reward signal, the student model can learn diverse behavioral patterns from multiple teacher models and further optimize the policy during interaction with the environment, as Figure 7 shown.
[0213] Step (8): Determine whether the student model has learned well. If so, end; otherwise, update the policy and repeat the above steps.
[0214] In this embodiment, through image rendering technology, the environmental state (such as robot sensor data, scene information) is converted into visual information, enhancing the environmental perception ability. In addition, combined with the CLIP feature extraction technology, the unified representation of multi-modal data is realized, improving the data utilization efficiency. And by using the multi-teacher knowledge distillation mechanism, the limitations of a single policy model are effectively overcome, significantly enhancing the diversity and generalization ability of the decision-making strategy.
[0215] Next, a specific experimental analysis description of the robot intelligent decision control method based on adversarial imitation learning provided by the embodiments of the present application is presented.
[0216] The robot intelligent decision control method based on adversarial imitation learning provided by the embodiments of the present application is experimented on three control tasks with high-dimensional and continuous action spaces based on physics. The experimental results show that the method proposed by the embodiments of the present application will learn a better policy than the policy GAIL and the policy OODIL (Observation-Orientation-Decision Iterative Learning) under the same circumstances.
[0217] Next, the technical effects of the robot intelligent decision control method based on adversarial imitation learning proposed by the embodiments of the present application are described in detail in combination with the experiments.
[0218] Please refer to Figure 8 , Figure 8 , which is a screenshot of the visualization graphical interface for the three experimental tasks of the robot intelligent decision control method based on adversarial imitation learning proposed by the embodiments of the present application.
[0219] As Figure 8 shown, the three experimental tasks of the robot intelligent decision control method based on adversarial imitation learning proposed by the embodiments of the present application include:
[0220] HalfCheetah PyBullet: The goal is to make the HalfCheetah run forward quickly. This task has a twenty-six-dimensional continuous state space and a six-dimensional continuous action space. The reward is given based on each step the object takes, according to its movement relative to the floor, control cost, etc.
[0221] Hopper PyBullet: The goal is to make the Hopper move forward continuously without falling. This task has an eleven-dimensional continuous state space and a three-dimensional continuous action space. The reward is given based on each step the object takes, according to its movement relative to the floor, control cost, etc.
[0222] Walker PyBullet: The goal is to enable the Walker to walk forward while maintaining balance. This task has a 17-dimensional continuous state space and a 6-dimensional continuous action space. The reward is given based on the movement of the object per step, according to its movement relative to the floor, control cost, etc.
[0223] The experimental settings of the robot intelligent decision-making control method based on adversarial imitation learning proposed in the embodiments of this application are as follows:
[0224] Feedforward neural networks are used as approximators with different architectures: the policy network - two hidden layers, each with 128 units, the activation term between them is tanh, and the output layer is processed by softmax; the value network and discriminator network in the TRPOs algorithm consist of two hidden layers with 128 units each and an output layer, and the activation term between them is tanh. All networks are randomly initialized at the start of each trial. In each task, three random seeds are used for the random initialization of the environment simulator and the model, and the policy learning curves of the robot under three methods in the HalfCheetah, Hopper, and Walker environments are plotted for comparison.
[0225] The experimental results of the robot intelligent decision-making control method based on adversarial imitation learning proposed in the embodiments of this application are as Figure 9 , Figure 10 and Figure 11 shown, where Figure 9 is a schematic diagram of the comparison curves of the three methods in the HalfCheetah physical environment, Figure 10 is a schematic diagram of the comparison curves of the three methods in the Hopper physical environment, Figure 11 is a schematic diagram of the comparison curves of the three methods in the Walker physical environment. From the results of the three figures, it can be seen that the robot intelligent decision-making control method based on adversarial imitation learning proposed in the embodiments of this application has a faster learning speed and better learning effect compared with Gail and OODIL. This shows that by introducing CLIP feature extraction and knowledge distillation, the robot intelligent decision-making control method based on adversarial imitation learning proposed in the embodiments of this application can learn strategies more effectively, thus obtaining higher rewards. The OODIL method also shows better performance than the traditional Gail through contrastive learning clustering and weight assignment, but it is not as good as the robot intelligent decision-making control method based on adversarial imitation learning proposed in the embodiments of this application. Generally speaking, introducing additional feature extraction and knowledge distillation steps can significantly improve the effect of imitation learning, making the model perform better in complex tasks.
[0226] Next, please refer to Figure 12, embodiments of the present application further provide a robot intelligent decision-making control device based on adversarial imitation learning. The robot intelligent decision-making control device based on adversarial imitation learning can implement the above-mentioned robot intelligent decision-making control method based on adversarial imitation learning.
[0227] As Figure 12 shown, in some embodiments, the robot intelligent decision-making control device based on adversarial imitation learning provided by embodiments of the present application may include:
[0228] A data acquisition and feature extraction module, configured to acquire expert demonstration data of the robot, and perform semantic feature extraction on the image data corresponding to the expert demonstration data based on a vision-language pre-trained model, so as to obtain the semantic features of each trajectory data in the expert demonstration data;
[0229] An adversarial imitation learning module, configured to construct a group of teacher models based on the semantic features and a generative adversarial imitation learning strategy;
[0230] A knowledge distillation module, configured to obtain the distillation loss between multiple teacher models in the group of teacher models and a single student model; and transfer the knowledge of the multiple teacher models to the single student model based on the distillation loss, and perform robot intelligent decision-making control based on the student model.
[0231] In some embodiments, the image data corresponding to the expert demonstration data includes the image data corresponding to each trajectory data in the expert demonstration data;
[0232] The data acquisition and feature extraction module is further configured to perform semantic feature extraction on the image data corresponding to the target trajectory data in the expert demonstration data based on a vision-language pre-trained model, so as to obtain the mean value of the trajectory features of the target trajectory data, and use the mean value of the trajectory features as the semantic feature of the target trajectory data;
[0233] Wherein, the target trajectory data is any one of multiple trajectory data in the expert demonstration data.
[0234] In some embodiments, the data acquisition and feature extraction module is further configured to acquire environmental dimension information associated with the expert demonstration data; and perform image rendering processing on each trajectory data based on the environmental dimension information and the action information of each trajectory in the expert demonstration data, so as to obtain the image data corresponding to each trajectory data.
[0235] In some embodiments, the adversarial imitation learning module is further configured to divide the semantic features into multiple cluster sets; and, independently train each of the multiple cluster sets based on the generative adversarial imitation learning strategy to obtain a teacher model corresponding to each cluster set; the teacher model group includes the teacher models corresponding to each cluster set.
[0236] In some embodiments, the knowledge distillation module is further configured to obtain the teacher actions generated by each of the multiple teacher models in the teacher model group, and, obtain the student actions generated by a single student model; and, calculate the distillation loss between each of the multiple teacher models and the single student model based on the teacher actions and the student actions.
[0237] In some embodiments, the knowledge distillation module is further configured to obtain the data generated by the interaction between the robot and the environment; and, perform an initialization process on the single student model based on the data generated by the interaction between the robot and the environment to obtain the student actions generated by the single student model.
[0238] In some embodiments, the knowledge distillation module is further configured to optimize the policy network parameters of the single student model based on the distillation loss and the reward signal of a preset discriminator, so as to transfer the knowledge of the multiple teacher models to the single student model;
[0239] wherein, the discriminator generates a reward signal based on the expert demonstration data and the student actions generated by the single student model.
[0240] It should be noted that the specific implementation manners of the robot intelligent decision-making control device based on adversarial imitation learning provided in the embodiments of the present application are basically the same as the specific embodiments of the above-mentioned robot intelligent decision-making control method based on adversarial imitation learning, and will not be elaborated herein.
[0241] Embodiments of the present application further provide a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned robot intelligent decision-making control method based on adversarial imitation learning and the robot intelligent decision-making control method based on adversarial imitation learning. This computer device can be any intelligent terminal device including a tablet computer, a personal computer, etc.
[0242] Please refer to Figure 13 , Figure 13 which illustrates the hardware structure of the computer device in some embodiments. The computer device may include:
[0243] The processor 1301 can be implemented in the form of a general - purpose CPU (Central Processing Unit), a microprocessor, an application - specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0244] The memory 1302 can be implemented in the form of a read - only memory (ROM), a static storage device, a dynamic storage device, or a random - access memory (RAM), etc. The memory 1302 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1302 and are called by the processor 1301 to execute the robot intelligent decision - making control method based on adversarial imitation learning in the embodiments of the present application;
[0245] The input / output interface 1303 is used to implement information input and output;
[0246] The communication interface 1304 is used to implement communication interaction between this device and other devices. It can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0247] The bus 1305 transmits information between the various components of the device (such as the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304);
[0248] Among them, the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304 achieve communication connections with each other inside the device through the bus 1305.
[0249] The embodiments of the present application also provide a computer - readable storage medium. The computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above - mentioned robot intelligent decision - making control method based on adversarial imitation learning and the robot intelligent decision - making control method based on adversarial imitation learning.
[0250] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0251] An embodiment of the present application also provides a computer program product. The computer program product stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned robot intelligent decision-making control method and robot intelligent decision-making control method based on adversarial imitation learning.
[0252] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0253] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0254] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0255] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0256] In the description of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0257] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0258] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.
[0259] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0260] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0261] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0262] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. This does not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.
Claims
1. A robot intelligent decision-making control method based on adversarial imitation learning, characterized in that: The method comprises: Acquire expert demonstration data of the robot, and extract semantic features of image data corresponding to the expert demonstration data based on a visual language pre-training model to obtain semantic features of each trajectory data in the expert demonstration data; Constructing a teacher model group based on the semantic features and the generative adversarial imitation learning strategy; Obtaining the distillation loss between multiple teacher models in the teacher model group and a single student model; The knowledge of the multiple teacher models is transferred to the single student model based on the distillation loss, and the robot intelligent decision control is performed based on the student model.
2. The robot intelligent decision-making control method based on adversarial imitation learning according to claim 1 is characterized in that: The image data corresponding to the expert demonstration data includes image data corresponding to each track data in the expert demonstration data; The extracting semantic features of the image data corresponding to the expert demonstration data based on the visual language pre-training model to obtain the semantic features of each trajectory data in the expert demonstration data includes: Based on the visual language pre-training model, semantic features are extracted from the image data corresponding to the target trajectory data in the expert demonstration data to obtain a trajectory feature mean of the target trajectory data, and the trajectory feature mean is used as the semantic feature of the target trajectory data; The target trajectory data is any one of the multiple trajectory data of the expert demonstration data.
3. The robot intelligent decision-making control method based on adversarial imitation learning according to claim 2 is characterized in that: The method further comprises: Acquiring environmental dimension information associated with the expert demonstration data; Based on the environmental dimension information and the action information of each trajectory in the expert demonstration data, image rendering processing is performed on each trajectory data to obtain image data corresponding to each trajectory data.
4. The robot intelligent decision-making control method based on adversarial imitation learning according to claim 1 is characterized in that: The method of constructing a teacher model group based on the semantic features and the generative adversarial imitation learning strategy includes: dividing the semantic features into a plurality of clusters; Based on the generative adversarial imitation learning strategy, each cluster in the multiple clusters is independently trained to obtain a teacher model corresponding to each cluster; the teacher model group includes the teacher model corresponding to each cluster.
5. The robot intelligent decision-making control method based on adversarial imitation learning according to claim 1 is characterized in that: The obtaining of the distillation loss between the multiple teacher models in the teacher model group and the single student model includes: Obtaining teacher actions generated by each of the plurality of teacher models in the teacher model group, and obtaining student actions generated by a single student model; A distillation loss between each of the plurality of teacher models and the single student model is calculated based on the teacher action and the student action.
6. The robot intelligent decision-making control method based on adversarial imitation learning according to claim 5 is characterized in that: The step of obtaining the student actions generated by the single student model includes: Acquiring data generated by the interaction between the robot and the environment; The single student model is initialized based on the data generated by the interaction between the robot and the environment to obtain the student movements generated by the single student model.
7. The robot intelligent decision-making control method based on adversarial imitation learning according to claim 1 is characterized in that: The step of transferring the knowledge of the multiple teacher models to the single student model based on the distillation loss comprises: Optimizing the policy network parameters of the single student model based on the distillation loss and a reward signal of a preset discriminator to transfer the knowledge of the multiple teacher models to the single student model; Wherein, the discriminator generates a reward signal based on the expert demonstration data and the student actions generated by the single student model.
8. A robot intelligent decision-making control device based on adversarial imitation learning, characterized in that: The device comprises: A data acquisition and feature extraction module is used to acquire expert demonstration data of the robot, and perform semantic feature extraction on image data corresponding to the expert demonstration data based on a visual language pre-training model to obtain semantic features of each trajectory data in the expert demonstration data; An adversarial imitation learning module, used for constructing a teacher model group based on the semantic features and the generative adversarial imitation learning strategy; A knowledge distillation module is used to obtain the distillation loss between multiple teacher models in the teacher model group and a single student model; and based on the distillation loss, the knowledge of the multiple teacher models is transferred to the single student model, and the robot intelligent decision-making control is performed based on the student model.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the robot intelligent decision-making control method based on adversarial imitation learning as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the robot intelligent decision-making control method based on adversarial imitation learning described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Control strategy imitation learning method and device based on adversarial learning
CN111488988A
Method and device for evaluating distillation strategy based on reinforcement learning and medium
CN115057006A
Incremental relation extraction method based on knowledge distillation
CN115203404A
Semantic segmentation model training method, semantic segmentation method and device
CN115511892A
Robot learning method for model-based generative adversarial interactive imitation learning
CN116663651A
Cited By
Multi-task dexterous hand operation method based on self-adaptive two-way distillation
CN121589820A
A multi-task dexterous hand operation method based on adaptive bidirectional distillation
CN121589820B