A Visual Servo Control Method and System Based on Novelty Metric for SAC Reinforcement Learning

By applying SAC reinforcement learning method in visual servo, combining cluster novelty measurement and hybrid sampling technology, the problems of large amount of manual design and low exploration efficiency are solved, and efficient end-to-end visual servo control is achieved.

CN116382089BActive Publication Date: 2025-05-30NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310439089.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-21
Publication Date
2025-05-30
Estimated Expiration
2043-04-21

AI Technical Summary

Technical Problem

The existing visual servo method based on visual feature design has the problem of large amount of manual design, and the visual servo based on reinforcement learning is slow to explore and utilize interactive data in continuous state space.

Method used

Using reinforcement learning method based on Soft Actor-Critic (SAC), we use the Actor-Critic network structure to define the objective function and the external reward function, and design a clustering novel metric method and a mixed probability sampling method to improve learning efficiency and exploration ability.

Benefits of technology

The manual design link is reduced, the exploration efficiency of reinforcement learning in continuous state space is improved, the end-to-end learning from visual features to control amount is realized, and the automation level of visual servo control is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116382089B_ABST
    Figure CN116382089B_ABST
Patent Text Reader

Abstract

The present invention relates to a visual servo control and system for novelty metric-based SAC reinforcement learning, and relates to the application field of reinforcement learning in visual servo. The method includes steps such as the construction of a virtual environment, the definition of an objective function, the definition of a reward function during the interaction process, and the definition of a novelty metric function. In the execution stage, the target position information is input into the policy network of the reinforcement learning, and the quadrotor is position-controlled according to the control speed information output by the policy network. The present invention solves the problem of many links between visual features and control speed information in related theories, realizes end-to-end control from features to control quantities, and thus reduces the artificial design links.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the application field of reinforcement learning in visual servoing. Specifically, it is an end-to-end visual servo control method based on reinforcement learning SAC (Soft Actor-Critic), which also includes a clustering-based novelty metric method to encourage the exploration space of reinforcement learning and a hybrid probability sampling method to improve learning efficiency. Background Art

[0002] Visual servoing (VS) is a key method that generates control commands by comparing the currently observed visual features with the expected visual features, enabling the robot to move towards the expected pose. Since visual perception and control form an important theoretical basis for manipulating robots to complete tasks, visual servoing has a wide range of applications in the field of robotics, such as autonomous grasping, target tracking, target grasping, etc. Classical visual servoing theory is mainly divided into image-based visual servoing (IBVS) and position-based visual servoing (PBVS) according to the attributes of feature errors.

[0003] Reinforcement learning (RL) is a way to maximize the expected reward value based on the current observation value through continuous interactive trial and error of an agent in a virtual environment. The development of RL is driven by CNN fitting complex behavioral policies. To simplify the data acquisition and loss function design process, RL-based visual servoing has gradually become one of the hotspots in this field, and it is often used to adaptively adjust control laws or gain parameters. With the development of RL theory, it is possible to control the end-to-end learning behavior through the interaction between the agent and the environment. However, different from adaptive parameter adjustment, RL-based visual servoing in continuous space has problems such as difficult sample exploration and low sampling efficiency. Due to the huge number of buffer pools or limited interaction time of the agent during the training process, the exploration and utilization of the buffer pool are closely related to the learning efficiency of the agent. Among them, Soft Actor-Critic (SAC) is a relatively novel method in reinforcement learning theory. Its characteristic lies in increasing the policy entropy, enhancing the exploration ability of the Actor for the environment, and maximizing the entropy of the policy as the objective function for enhancing exploration while maximizing the reward expectation. Summary of the Invention

[0004] The technical problems to be solved by the present invention are:

[0005] In existing position-based visual servoing or image-based visual servoing designed based on visual features, there is a problem of a large amount of manual design. In contrast, visual servoing based on reinforcement learning can learn visual servoing strategies directly applicable to robot control by designing an objective function and constructing a simulation environment. The present invention provides a SAC reinforcement learning visual servoing control method based on novelty measurement to solve the problem in related technologies that servo control needs to be achieved through camera parameter acquisition, manual feature design, and manual control law design, resulting in a large amount of manual work.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0007] A SAC reinforcement learning visual servoing control method based on novelty measurement, characterized by comprising:

[0008] Create a virtual environment for a quadrotor to track a target:

[0009] Construct an Actor-Critic network structure for describing training in SAC reinforcement learning. The Actor-Critic network structure includes a policy network Actor and a value network Critic. The two networks respectively learn visual servoing strategies and fit the value function in reinforcement learning;

[0010] Define an objective function and an extrinsic reward function during the training process of the network. The objective function adopts a function consistent with the reinforcement learning Soft Actor-Critic method, and the extrinsic reward function is defined by the position of the target in the field of view;

[0011] During the training process, design a clustering method to measure the novelty degree of visited states. Measure the novelty of visited states through the proposed centralized clustering metric as an intrinsic reward;

[0012] Through continuous interaction of a reinforcement learning agent in the virtual environment, obtain training data including target position, extrinsic reward and intrinsic reward, control amount, and target position at the next moment, and optimize the network using SAC reinforcement learning;

[0013] During actual use, first obtain the position information of the target detection result, and then input this information into the Actor network to obtain the control amount, and transmit the control amount to the quadrotor for servo control.

[0014] A further technical solution of the present invention: construct a target detection model and a quadrotor centroid model in the V-rep virtual environment according to the task background, that is, a virtual tracking scenario.

[0015] Further technical solution of the present invention: Both the policy network Actor and the value network Critic are composed of multi-layer neural networks, and the activation function uses LeakyRelu; the input and output of the Actor network are the target detection result and the control amount respectively, and the input of the Critic network is the target detection result and the control amount, and the output is the fitted Q value.

[0016] Further technical solution of the present invention: The external reward function is:

[0017]

[0018] where (x tar , y tar ) and (x g , y g ) are the coordinates of the target and the quadrotor aircraft in the xy plane respectively, and the corresponding objective function is defined as:

[0019]

[0020] where H(π(·|s t ) is the policy entropy representing the uncertainty of the behavior policy, α is a hyperparameter used to adjust the proportion of the policy entropy, π represents the policy of the Actor network, represents the expected value of the reward under the policy, R(s t , a t ) represents the reward value at the current moment, s t represents the state in reinforcement learning, i.e., the target position information, a t represents the control amount, and ρ π represents the state transition probability when the policy is π.

[0021] Further technical solution of the present invention: A method for measuring the novelty of the visited state by a clustering method, and its objective function is:

[0022]

[0023] where s t is the state observed by the agent, center is the defined cluster center, f(s t ; θ) represents the feature output by the network, θ represents the network parameters, and r int is defined as the intrinsic reward during the training process.

[0024] Further technical solution of the present invention: A mixed probability sampling method is also designed to obtain training data, and the probability is defined as:

[0025]

[0026]

[0027] Among them, is a hyperparameter, and the reward given by the environment is the external reward r ext , while the reward given by the novelty network is the internal reward r int , is the temporal difference error of the i-th sample, is the novelty degree of the i-th sample; where ε is a minimum positive number to prevent the probability from being 0;

[0028] The importance sampling weight is used to correct the backpropagation gradient:

[0029]

[0030] Among them, N is the number of samples in the cache pool, and β is a parameter used to correct the gradient, which is obtained through learning.

[0031] A SAC reinforcement learning visual servo control system based on novelty measurement, characterized by comprising:

[0032] A creation module, used to create a simulation environment describing the relationship between distance and pose;

[0033] An acquisition module, used to generate data during the interaction process;

[0034] A construction module, used to construct neural networks to form Actor and Critic networks; define the objective function, novelty measurement method, and hybrid sampling method, so as to optimize the training process during training; during the execution process, the converged Actor network is regarded as a controller, inputting the observation features and outputting the control quantity, so as to be used for quadrotor control.

[0035] A computer system, characterized by comprising: one or more processors, and a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.

[0036] A computer-readable storage medium, characterized by storing computer-executable instructions, which are used to implement the above method when executed.

[0037] The beneficial effects of the present invention are as follows:

[0038] The present invention designs an RL-based method to learn an end-to-end VS strategy, which takes the visual target detection result as input and directly outputs the planned motion information to the quadrotor robot, keeping the target always at the center of the quadrotor robot's field of view. In addition, a novelty measurement method and a data sampling method in the cache pool are proposed to solve the problem that RL explores and utilizes interactive data relatively slowly in continuous state space and action space. Moreover, a novel novelty measurement method for measuring visited states and a data sampling method in the cache pool are proposed to solve the problem of low efficiency of RL in exploring and utilizing interactive data in continuous states and action spaces. In comparison, the beneficial effects of the present invention are as follows:

[0039] 1. Reduce the manual design process. Manually designing visual features and control laws is the most laborious part in traditional visual servoing. In the method of the present invention, the visual servo strategy can be continuously learned through interaction in the simulation environment;

[0040] 2. Encourage the exploration of the state space by the Actor network in reinforcement learning. By designing a state novelty measurement method, the novelty of the visited states is measured through a neural network and an objective function during the interaction process, thus encouraging the Actor network to explore the states with fewer visits;

[0041] 3. Design a method based on hybrid sampling probability to reduce the degree of dependence on time difference during sampling in the reinforcement learning process. Description of the Drawings

[0042] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components.

[0043] Figure 1 Reinforcement learning visual servo architecture diagram;

[0044] Figure 2 Schematic diagram of novelty measurement;

[0045] Figure 3 Experimental scenario;

[0046] Figure 4 MNIST novelty test results;

[0047] Figure 5 Influence of sampling on the training process: (a) Comparing the influence of different reward methods; (b) Comparing the influence of different sampling methods;

[0048] Figure 6 Change of feature error during servoing;

[0049] Figure 7 Action convergence process during servoing

[0050] Figure 8 State and action changes when the target moves;

[0051] Figure 9 Quadrotor trajectory when the target is stationary;

[0052] Figure 10 Quadrotor trajectory when the target moves;

[0053] Figure 11 Flowchart of the visual servo control method based on SAC reinforcement learning. Detailed implementation manners

[0054] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0055] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the accompanying drawings are used to distinguish different objects, rather than to limit a specific order.

[0056] According to one aspect of the embodiments of the present invention, a SAC reinforcement learning visual servo control method based on novelty metric is provided, including: constructing a simulation environment for a quadrotor to track a target; constructing a multi-layer neural network representing the Actor and Critic in SAC, with the network input being the target position coordinates and the outputs being the control quantity and the Q value respectively; designing a reinforcement learning objective function based on the visual servo input and output information; defining an extrinsic reward function; generating training data during the interaction process for policy optimization; after training convergence, inputting the target detection result into the Actor network to obtain the control quantity for quadrotor control, thereby performing servo control.

[0057] Optionally, in the V-rep simulation software environment, a simulation environment of the relationship between the position target and the control quantity (speed information) is constructed according to the three-dimensional space mathematical model of the camera on the quadrotor and the target.

[0058] Optionally, the dataset for training is generated by continuously interacting the Actor network with the simulation environment, where the input is the target position and the output is the control quantity that enables the quadrotor plane to move to keep the target at the visual center.

[0059] According to another aspect of the embodiments of the present invention, there is provided a method for measuring state novelty as an intrinsic reward method and a method for sampling a sample of mixed probabilities from a cache pool. It includes: by defining a feature center, defining the features extracted by the state through a neural network as features, and defining the distance from the feature to the center as a measure for measuring state novelty, that is, the intrinsic reward; since the essence of sampling learning in reinforcement learning is to reduce the temporal difference error, in the present invention, a probability function for mixed sampling in the training stage is defined by mixing the temporal difference and the novelty metric probability.

[0060] In the embodiments of the present invention, it includes: creating a simulation environment for describing the target position coordinates and the quadrotor plane movement coordinates, defining the reinforcement learning optimization objective function during the training process, designing a neural network to define the novelty metric function, and defining the mixed sampling probability function according to the temporal difference and the novelty. The present invention solves the problem of a large amount of manual participation in the traditional visual servo method in the artificial design of features and the artificial design of control laws, and improves the reinforcement learning efficiency through novelty measurement and mixed sampling, so as to complete the end-to-end learning from visual features to control quantities.

[0061] As Figure 11 shown, it includes the following steps:

[0062] Step 1, create a simulation environment: According to the target detection algorithm and the quadrotor centroid model, simulate the quadrotor movement in the V-rep simulation software to make the target located in the middle of the field of view;

[0063] Step 2, construct a policy network Actor and a value network Critic for describing training in reinforcement learning: The two networks respectively learn the visual servo policy and fit the value function in reinforcement learning; where the input and output of the Actor network are the target detection result and the control quantity respectively, and the input of the Critic network is the target detection result and the control quantity, and the output is the fitted Q value; both networks are composed of a multi-layer neural network, and the activation function uses LeakyRelu;

[0064] Step 3, define the objective function: Taking the visual servo task as the background, define variables according to the input and output, define the external reward function, optimize the network with the Bellman update formula, and add the policy entropy definition to optimize the exploration process;

[0065] Step 4, define the novelty measurement method, that is, the intrinsic reward: Perform Euclidean distance measurement on the distance between the accessed state and the feature center, and the obtained result is used as the degree of novelty;

[0066] Step 5, mixed sampling probability: When sampling training samples from the cache pool, calculate the sampling probability according to the defined temporal difference and novelty, so as to determine the probability distribution of sample sampling;

[0067] Step 6: Save the converged visual servo policy model;

[0068] Step 7: According to the saved visual servo policy network, use the position of the detected target in the camera image during actual execution as the input of the network, and the output control amount is used to control the position of the quadrotor for visual servo control.

[0069] The following combines specific examples to illustrate this embodiment.

[0070] The present invention constructs a visual servo based on novelty metric reinforcement learning SAC, and the framework diagram is as Figure 1 shown. The reinforcement learning-based solution can be described by a Markov decision process. This is because the output of the visual servo control policy only depends on the current state. P(s t+1 |s t ) = P(s t+1 |s 1 s 2 ...s t ), which has the Markov property. In the reinforcement learning-based VS task, the image features or errors are regarded as the state. s While the control instruction of the manipulator is defined as the action a, where s ∈ S and a ∈ A. At time t, the servo policy accepts the observed state s t , outputs the control instruction a t , and transfers it from s t to s t+1 according to the state transition function P s,a . According to the artificially designed reward function, the reward r is obtained. γ is expressed as the discount factor to balance the importance of long-term rewards and current rewards. The above five elements (S, A, P s,a , R, γ) constitute the MDP tuple, and (s t , a t , s t+1 , r t , d) constitute the transition of the end-to-end visual servo system. In the end-to-end VS, it is considered as an agent that controls the robot to continuously execute a sequence of actions until it reaches the target pose. Visual servo based on SAC: The present invention develops an end-to-end VS method for target tracking based on SAC. At each step, a sub-goal g t = (x t , g t ) is set at the intermediate position between the current state and the center, and its distance Done = True gradually approaches 0. The state s t = (x t , y t , x g , y gis the target position observed by the on-board camera of the robot in the continuous space, and the action a t is the control instruction v of the robot in the continuous space t = (v x , v y ) is the control instruction of the robot in the continuous space.

[0071] The definition of the one-sided bounded external reward is:

[0072]

[0073] where, (x tar , y tar ) and (x g , y g ) are the coordinates of the target and the quadrotor on the xy plane respectively. The Bellman equation corresponding to the maximum entropy SAC is:

[0074]

[0075] where, is the state transition probability. r(s, a) is the reward value, and γ represents the discount factor of the reward. The corresponding SAC Bellman backup update formula is:

[0076]

[0077] The parameter α balances the reward expectation value and the maximum entropy. The update method is:

[0078]

[0079] where the entropy target is set to

[0080] Measuring the novelty of the state is crucial for the agent to explore the unknown environment. The present invention introduces a novelty measurement method based on a neural network, which has access to the centralized features of the state. A simple neural network trained by a centralized loss function is designed to measure the novelty of the state. The accessed state is input into the neural network, and then the corresponding features are output. The novelty of the state and the intrinsic reward are obtained by calculating the distance between the features and the cluster center. The novelty measurement architecture is as Figure 3 shown.

[0081] According to the learning property, the state with more access times is closer to the center. Therefore, the farther a state is from the center, the more novel it is. The loss function for training the network is defined as:

[0082]

[0083] where, s tis the state observed by the agent, center is the defined cluster center, which is a constant and does not participate in the update, θ represents the network parameters, and f(s t ; θ) represents the features output by the network. The reward given by the environment is the extrinsic reward, while the reward given by the novelty network is the intrinsic reward.

[0084] The update gradient of the novelty network is:

[0085]

[0086] As the network is trained, the features extracted by the network from the states gradually approach the cluster center.

[0087] According to the definition of learning, when a state appears more frequently in the dataset, after several rounds of learning, the extracted features should be closer to the center. Similarly, when the number of appearances is small, that is, the state is relatively novel and has not reached, the features extracted through the network are farther from the clustering center. The distance from the mapped features to the center is the intrinsic reward. During the training process, the distance from the features to the center is stored as the intrinsic reward in the transition and participates in the training of the RL scheme to represent the novelty of the state.

[0088] Therefore, during the learning process, the total reward is:

[0089]

[0090] where is a hyperparameter set to 0.1. The reward given by the environment is the extrinsic reward r ext , while the reward given by the novelty network is the intrinsic reward r int .

[0091] Based on the novelty of the state and the TD error property of the transition samples, hybrid probability sampling and maximum probability sampling are proposed, and the sampling probability is optimized according to the distribution property of the samples in the replay buffer. First, the intrinsic reward obtained during the training process is normalized by using the values of the previous steps:

[0092]

[0093] where, is the average value of the previous steps, output by the novelty network, and δ int is the corresponding standard deviation. After the intrinsic reward is normalized, the overall reward for each step is calculated by formula (7).

[0094] When using r ext to calculate the probability, since r ext is a negative number, it is passed through -1 / r ext . Combining TD to calculate the probability of each sample, given by:

[0095]

[0096] Among them, ε is the smallest positive number to prevent the probability (as the denominator) from being 0. Generally, the TD error δ t is defined as:

[0097]

[0098] The probability of the sampling probability P(i) transitioning to ii is proportional to the priority P(i) and is defined as

[0099]

[0100] where k is the number of minibatches. The exponent α determines how much priority is used. The importance sampling weight is used to correct the backpropagation gradient:

[0101]

[0102] where N is the batch size. β is a parameter used to correct the gradient and is obtained through learning.

[0103] Experimental analysis

[0104] The proposed VS method was trained in a simulation study and then verified in an experimental environment. First, the parameter settings during the training process are introduced, and then the convergence data, completion rate, and reward value of the network during the training process are compared with other RL algorithms. In addition, a novelty test was conducted on the sampled states and the dataset MNIST. Finally, the end-to-end VS algorithm was demonstrated and illustrated in a simulation environment and a real environment. To describe in detail the experimental results of the algorithm proposed in the present invention, the experimental results are presented in detail through Figure 4 、 5 、6, 7, 8, 9, 10.

[0105] The present invention verified the novelty measurement method on the open-source dataset MNIST. The size of each image is 28×28, reflecting the performance of our algorithm on high-dimensional data. Note: This test uses the samples of each molecule in MNIST as input, and each image is different. Therefore, similar observed states are still valid in this algorithm. Figure 4 The subgraphs in emphasize that our algorithm can achieve the effect of sample novelty measurement on datasets of different sizes. Although the process of the intrinsic reward decreasing shows specific non-linearity, the overall trend is consistent (the more input times, the smaller the output value).

[0106] The present invention set different rewards and different sampling methods for training, and the results are as Figure 5 shown.Figure 5 (a) The impacts of different reward methods were mainly compared. When the baseline was SAC, if there was only intrinsic reward, the effect was the worst. Because intrinsic reward only represents the novelty of the state and cannot reflect whether the task is completed. When sampling and hybrid rewards were used, the effect would be improved. Figure 5 (b) The impacts of different sampling methods were mainly compared. Similarly, sampling with intrinsic reward led to the worst effect. Sampling with only extrinsic reward and sampling with TD error obtained similar results. Generally speaking, using intrinsic reward can improve the baseline to a certain extent. Because it has nothing to do with task attributes, using intrinsic reward alone will lead to bad results in terms of reward or sampling. The sampling method combining hybrid reward and TD error also improved the effectiveness of the baseline.

[0107] In Figure 6 and Figure 7 , the quadrotor takes off at a random position, acquires the target position, and is controlled according to the control instructions calculated by the reinforcement learning servo strategy. Figure 6 shows the change of the target position in the field of view. The data of the three series show that the quadrotor can gradually make the target located at the center of the field of view during flight. Figure 7 indicates that the control instructions gradually decrease to a stable value until the target is located at the center of the field of view.

[0108] In Figure 8 , the present invention verified the effectiveness of the action strategy when the target moves. Figure 8 The upper part is the position change of the camera when the target moves, Figure 8 The lower part is the control commands generated by the agent when the target moves. From the trend of periodic change, it can be seen that the action strategy can generate corresponding actions according to the movement of the target.

[0109] In Figure 9 , the present invention visually shows the movement trajectory in three-dimensional space according to the action, and the lines of different colors represent the movement trajectories of different rounds. The quadrotor starts from three different positions, executes the actions provided by the action strategy, and the action strategy takes the coordinates of the target in the field of view of the front camera as input, and finally makes the target located at the center of the field of view. In Figure 10 , the movement trajectories of the quadrotor and the target can be visually seen.

[0110] An embodiment of the present invention also provides a vision servo framework based on reinforcement learning, including: a creation module for creating a simulation environment that describes the relationship between distance and pose; an acquisition module for generating data during the interaction process; a construction module for constructing a neural network to form an Actor and a Critic network. Define an objective function, a novelty metric method, and a hybrid sampling method to optimize the training process during training. During the execution process, the converged Actor network is regarded as a controller, and the observed features are input and the control quantity is output, so as to be used for quadrotor control.

[0111] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention.

Claims

1. A visual servo control method for SAC reinforcement learning based on novelty metric, characterized in that, it includes: Create a virtual environment for the quadrotor to track the target: Construct an Actor-Critic network structure for describing training in SAC reinforcement learning. The Actor-Critic network structure includes a policy network Actor and a value network Critic. The two networks respectively learn the visual servo policy and fit the value function in reinforcement learning; Define the objective function and the extrinsic reward function during the training of the network. The objective function adopts the same function as the reinforcement learning Soft Actor-Critic method, and the extrinsic reward function is defined by the position of the target in the field of view; During the training process, design a method of clustering to measure the novelty degree of the visited state. The novelty of the visited state is measured by the proposed centralized clustering metric as the intrinsic reward; Through the reinforcement learning agent to continuously interact in the virtual environment to obtain training data, including target position, extrinsic reward and intrinsic reward, control amount, and target position at the next moment, and use SAC reinforcement learning to optimize the network; During actual use, first obtain the position information of the target detection result, and then input this information into the Actor network to obtain the control amount, and transmit the control amount to the quadrotor for servo control.

2. The visual servo control method for SAC reinforcement learning based on novelty metric according to claim 1, characterized in that: Construct a target detection model and a quadrotor centroid model in the V-rep virtual environment according to the task background, that is, a virtual tracking scene.

3. The visual servo control method for SAC reinforcement learning based on novelty metric according to claim 1, characterized in that: Both the policy network Actor and the value network Critic are composed of multi-layer neural networks, and the activation function uses LeakyRelu; where the input and output of the Actor network are the target detection result and the control amount respectively, and the input of the Critic network is the target detection result and the control amount, and the output is the fitted Q value.

4. The visual servo control method for SAC reinforcement learning based on novelty metric according to claim 1, characterized in that: The extrinsic reward function is: where (x tar , y tar ) and (x g , y g ) are the coordinates of the target and the quadcopter in the xy-plane respectively, and the corresponding objective function is defined as: Among them, H(π(·|s t )) is the policy entropy representing the uncertainty of the behavioral policy, α is the hyperparameter used to adjust the proportion of the policy entropy, π represents the Actor network policy, represents the expected value of the reward under the policy, R(s t , a t ) represents the reward value at the current moment, s t represents the state in reinforcement learning, i.e., the target position information, a t represents the control quantity, and ρ π represents the state transition probability when the policy is π.

5. The visual servo control method for SAC reinforcement learning based on novelty metric according to claim 4, characterized in that: The method of using a clustering method to measure the novelty degree of the visited state, its objective function is: where s t is the state observed by the agent, center is the defined cluster center, f(s t ; θ) represents the feature output by the network, θ represents the network parameters, r int is defined as the intrinsic reward during the training process.

6. The visual servo control method for SAC reinforcement learning based on novelty metric according to claim 5, characterized in that: A mixed probability sampling method is also designed to obtain the probability of training data, and the probability is defined as: Among them, is a hyperparameter, and the reward given by the environment is the external reward r ext , while the reward given by the novelty network is the internal reward r int , is the temporal difference error of the i-th sample, is the novelty degree of the i-th sample; where ε is a minimum positive number to prevent the probability from being 0; The importance sampling weight is used to correct the backpropagation gradient: where N is the number of samples in the buffer pool, and β is the parameter used to correct the gradient, which is obtained through learning.

7. A visual servo control system for SAC reinforcement learning based on novelty metric for implementing the method described in claim 1, characterized in that it includes: A creation module for creating a simulation environment describing the relationship between distance and pose; An acquisition module for generating data during the interaction process; A building block for building a neural network to form an Actor and a Critic network; defining an objective function, a novelty metric, and a hybrid sampling method to optimize the training process during training; during the execution process, regarding the converged Actor network as a controller, inputting the observation features and outputting the control quantity for quadrotor control.

8. A computer system, characterized in that it includes: one or more processors, and a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to claim 1.

9. A computer-readable storage medium, characterized in that it stores computer-executable instructions that are used to implement the method according to claim 1 when executed.

Citation Information

Patent Citations

  • Power grid real-time adaptive decision-making method based on deep reinforcement learning

    CN114217524A

  • Reward function and vibration suppression reinforcement learning algorithm using same

    CN115327927A