Method and device for personnel skill training based on variable impedance operation difficulty adjustment
By adjusting the damping of the hand controller through fuzzy reinforcement learning, the problem of low efficiency caused by fixed operation difficulty in traditional skill training is solved, and real-time difficulty matching and training quality improvement are achieved.
Patent Information
- Application Number
- CN202311271451.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-09-28
AI Technical Summary
Traditional skills training methods cannot adjust the difficulty of operations in real time, resulting in low training efficiency and poor quality, and are unable to adapt to the randomness of changes in trainees' skills.
A fuzzy reinforcement learning-based approach is adopted to adjust the operation difficulty in real time through a hand controller damping system. The damping coefficient is generated based on the trainee's operation data, and a dynamic matching model between the trainee and the operation difficulty is established. The damping adjustment is optimized using a fuzzy inference system and Q-learning method.
It enables real-time adjustment of operational difficulty without prior assessment of trainees' skills, thereby improving training efficiency and quality and adapting to the needs of trainees with different skill levels.
Smart Images

Figure CN117173963B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a personnel skill training method and device based on variable impedance operation difficulty adjustment. BACKGROUND
[0002] In the process of training the skills of personnel, as the skills grow, the difficulty of the operation task relative to the personnel will change. At this time, for the trainee, if the operation difficulty is too simple, the skill learning efficiency of the trainee will be reduced. If the operation is too difficult, it will also lead to invalid training of the trainee. The reasonable matching of operation difficulty and skill level of the trainee will directly affect the efficiency of skill learning of the trainee. According to the skill level of the trainee, the operation difficulty of the trainee is adjusted, the skill training of the trainee is guided through the difficulty, the cognitive ability of the trainee to the operation skill is improved, so as to enhance the skill learning efficiency of the trainee. The trainee can obtain new skill knowledge from the training and improve his own operation level. In addition, time can also be saved in task planning. For different personnel, the adaptability of the operation difficulty will be different, through the learning-based method, a dynamic model between the skill level of the trainee and the operation difficulty is established, and the operation task difficulty matching that can maximize the learning efficiency is found. The traditional skill training method usually needs to collect the information of the trainee in an offline state to establish the skill model of the trainee and set the training scheme. However, because the skill change of the trainee usually has randomness, the skill change model of the trainee is difficult to accurately establish, and if the training is carried out according to the fixed training scheme, the training efficiency will be very low, and the training quality will also be very poor. SUMMARY
[0003] In order to solve the above technical problems, the present application mainly aims at the problem of low training efficiency caused by fixed operation difficulty, and provides a personnel skill training method and device based on variable impedance operation difficulty adjustment. The method adjusts the size of the damper of the hand controller operated by the trainee according to the skill performance of the trainee in the operation process, so as to change the difficulty of the trainee in operating the hand controller. When the trainee is trained, the skill of the trainee does not need to be evaluated in advance, and only the operation damping coefficient corresponding to the data generated by the trainee in real time operation is needed to generate, and the matching relationship between the trainee and the operation difficulty is established. The damper controller of the hand controller is designed by adopting the fuzzy reinforcement learning method, and the variable admittance control strategy is implemented on the hand controller, so as to effectively solve the problem of task difficulty adjustment in the skill change.
[0004] The first object of the present application is to provide a personnel skill training method based on variable impedance operation difficulty adjustment, which is used for the trainee to hold the hand controller for skill training, comprising:
[0005] Obtaining the position information of the end of the hand controller;
[0006] According to the position information of the hand controller end, a plurality of skill performance characteristics of the trainee and a smoothness of the trainee in the operation process are acquired;
[0007] According to the position information of the hand controller end, a plurality of skill performance characteristics of the trainee and a smoothness of the trainee in the operation process are acquired;
[0008] According to the position information of the hand controller end, a plurality of skill performance characteristics of the trainee and a smoothness of the trainee in the operation process are acquired;
[0009] According to the position information of the hand controller end, a plurality of skill performance characteristics of the trainee and a smoothness of the trainee in the operation process are acquired;
[0010] According to the position information of the hand controller end, a plurality of skill performance characteristics of the trainee and a smoothness of the trainee in the operation process are acquired;
[0011] According to the position information of the hand controller end, a plurality of skill performance characteristics of the trainee and a smoothness of the trainee in the operation process are acquired;
[0012] According to the position information of the hand controller end, a plurality of skill performance characteristics of the trainee and a smoothness of the trainee in the operation process are acquired;
[0013] Preferably, the smoothness of the trainee in the operation process is calculated by calculating the operation smoothness of the trainee under the current damping state, and the smoothness constructs a reward value in the learning process, and the calculation formula is as follows:
[0014]
[0015] In the formula, r(S, U) represents the reward value acquired from the training environment; represents the third derivative value of the displacement value generated by the trainee.
[0016] Preferably, the membership degree corresponding to the skill level of each skill performance characteristic is acquired according to the following steps:
[0017] According to each skill performance characteristic, an evaluation value of the skill level of the trainee corresponding to each skill performance characteristic is acquired;
[0018] According to the membership function, a membership degree corresponding to each skill level evaluation value is acquired, that is, a membership degree corresponding to the skill level of each skill performance characteristic is acquired.
[0019] Preferably, the membership function is a triangular membership function.
[0020] Preferably, when constructing the rule between the skill performance and the hand controller damping, the skill level corresponding to each skill performance feature is different, and the conclusion value corresponding to the current hand controller damping is different, wherein the different conclusion values correspond to an action value;
[0021] The action value represents the weight of the different conclusion values.
[0022] Preferably, the trigger strength is calculated according to the following formula:
[0023]
[0024] In the formula, μ j (S n ) represents the membership degree corresponding to the nth skill level; The trigger strength is represented by μ.
[0025] Preferably, the optimal action value is obtained according to the following steps:
[0026] The reward of the trainee under the damping value generated by the current system is obtained according to the skill feature level of the trainee and the corresponding conclusion value;
[0027] The action value of the next moment is estimated according to the reward value obtained from the current environment, the current action value is updated by calculating the difference between the action value of the next moment and the current value, and the optimal action value is obtained.
[0028] Preferably, the hand controller damping coefficient is calculated according to the following formula:
[0029]
[0030] In the formula, U j (S) represents the hand controller damping coefficient of the hand controller in the S state; The trigger strength of the skill state of the trainee is represented by μ. The conclusion value is represented by S.
[0031] Preferably, the multiple skill performance features include the displacement length of the trainee in the operation process, the smoothness of the trainee in the operation process, the operation time, and the accuracy of the operation.
[0032] The present application provides a personnel skill training device based on variable impedance operation difficulty adjustment, which is used for the skill training of a trainee holding a hand controller, and comprises:
[0033] A data acquisition module is configured to acquire position information of the end of the hand controller.
[0034] The multiple skill performance features of the trainee and the smoothness of the trainee in the operation process are acquired according to the position information of the end of the hand controller.
[0035] a data processing module, configured to obtain a skill level corresponding to each skill performance feature and a membership degree corresponding to the skill level based on a membership function;
[0036] obtain a skill state triggering intensity of the trainee according to the membership degree corresponding to each skill performance feature;
[0037] map a conclusion value corresponding to the system damping according to the skill level corresponding to each skill performance feature by using a fuzzy inference system model, and construct a rule between the skill performance and the hand controller damping;
[0038] based on a Q learning method in reinforcement learning, and obtain feedback information, i.e., smoothness of the trainee in operation, from a training environment, update a current action value according to an estimation of the current action value and a next time action value, and obtain an optimal action value;
[0039] a damping mediation module, configured to obtain an optimal hand controller damping coefficient according to the optimal action value corresponding to the skill triggered in the rule;
[0040] adjust the damping of the hand controller according to the optimal hand controller damping coefficient.
[0041] The present application has at least the following beneficial effects:
[0042] The present application provides a personnel skill training method and device based on variable impedance operation difficulty adjustment. The method can adjust the requirement of operation difficulty change in the personnel training process in real time on the premise of being independent of personnel supervision, by planning the task difficulty based on a fuzzy reinforcement learning method. Meanwhile, the method can also adjust the difficulty of different skill levels of personnel. Adjusting the damping coefficient is significant for personnel operation training. When the hand controller is in a low damping coefficient, the hand controller will be more flexible but the corresponding operation precision is low, and the trainee needs higher skills to complete the operation. A larger damping coefficient will make the hand controller operation relatively dull but more accurate, and the operator needs lower operation ability. In order to make the trainee in a suitable difficulty training state, the operation difficulty is adjusted in real time according to the skill change of the trainee. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 The system structure diagram provided by the present application.
[0044] Figure 2 The shape of the fuzzy membership function in the present application.
[0045] Figure 3 The reward function in the present application.
[0046] Figure 4 The damping coefficient change curve in the present application. DETAILED DESCRIPTION
[0047] In order to illustrate the technical means and effects adopted by the present application to achieve the predetermined inventive objectives, the following will be described in detail in combination with embodiments.
[0048] The present application provides a personnel skill training method based on variable impedance operation difficulty adjustment. The training system of the present application has intelligent learning ability, and the learning is the adjustment behavior of the hand controller damping coefficient according to the skill level of the trainee. The significance of adjusting the damping coefficient for personnel operation training is that when the hand controller is in a low damping coefficient, the hand controller will be more flexible but the corresponding operation precision is low, and the trainee needs higher skill to complete the operation. A larger damping coefficient will make the hand controller operation relatively dull but more accurate, and the operator's operation ability is required to be low. In order to meet the training state of the trainee at the appropriate difficulty, the operation difficulty is adjusted in real time according to the skill change of the trainee.
[0049] In order to realize the control of the operation difficulty, a dynamic model is established according to the hand controller impedance model:
[0050]
[0051] M, D, and K represent inertia characteristics, damping characteristics, and stiffness characteristics, respectively. Wherein f h represents the interaction force between the hand controller and the trainee. x is the position of the hand controller end in the Cartesian coordinate system, is the velocity information of the hand controller end in the Cartesian coordinate system, is the acceleration information of the hand controller end in the Cartesian coordinate system. The subscript d represents the expected value. Due to the operation training, the operation trajectory of the trainee is not fixed, so the corresponding expected value is set to 0. At the same time, since the stiffness characteristics are easy to cause instability in the interaction, the training is affected, and therefore K is taken as 0, so that the trainee can operate freely in the space. The final dynamic model is:
[0052]
[0053] D is the damping coefficient to be adjusted.
[0054] The present application provides a personnel skill training method based on variable impedance operation difficulty adjustment, which is used for the trainee to hold the hand controller for skill training, and mainly trains the trainee's operation control ability of the hand controller. The smoothness during operation is used to reflect this index, and the specific steps include:
[0055] S1, acquiring the position information of the hand controller end;
[0056] According to the position information of the hand controller end, the trainee's multiple skill performance characteristics and the smoothness of the trainee in the operation process are obtained.
[0057] wherein the plurality of skill performance characteristics include: displacement length of the trainee during operation and smoothness of the trainee during operation, and operation time and accuracy of the operation. The length of the trajectory generated by the trainee during operation to complete a task, the control force of the trainee on the operated device, i.e. the smoothness of the operation of the trainee.
[0058] The smoothness of the trainee during operation is calculated by calculating the smoothness of the operation of the trainee under the current damping state, and the formula is as follows:
[0059]
[0060] In the formula, r(S, U) represents the reward value obtained from the training environment, which is determined by the smoothness of the operation of the trainee after the skill state of the trainee and the damping coefficient applied to the hand controller; represents the third derivative value of the displacement value generated by the trainee, i.e. the smoothness of the operation of the trainee.
[0061] In the embodiment, r(S, U) represents the reward obtained by the system from the training environment, which is determined by the smoothness of the operation. The smoothness of the trainee during operation is the cumulative value of the change of acceleration in a period of time, S is the current training state, including the level of each skill characteristic of the trainee; U represents the damping coefficient generated by the system for adjusting the operation damping of the trainee; represents the jerk information, which is used to describe the change degree of the acceleration of the trainee under the current damping coefficient value, and can reflect the smoothness of the operation of the trainee.
[0062] In the embodiment, the position information of the end of the hand controller is the position of the end of the hand controller in the Cartesian coordinate system;
[0063] It should be noted that the end of the operated hand controller is provided with a sensor, and the collected operation performance data is taken as the state input when a certain specific task is executed, i.e. Based on the collected task information, the skill of the trainee will be divided into four levels first. The four levels of the corresponding language variables in the fuzzy system are: poor, general, high level, and skilled. In order to evaluate the good and bad of the trainee during operation, the jerk information of the trainee during training is taken as the reward signal to train the intelligent agent. The jerk information reflects the smoothness of the operation of the trainee during operation. The smoothness of the operation of the trainee under the current damping value state is calculated to construct the instantaneous reward value, and therefore the instantaneous reward is defined as: r(S, U) represents the smoothness of the trainee during operation, and can also represent the instantaneous reward. The target is to minimize the cumulative jerk during the entire training process.
[0064] S2, obtaining the membership degree corresponding to the skill level of each skill performance feature based on the membership function;
[0065] According to the membership degree corresponding to each skill performance feature, the trigger strength of the skill of the trainee is obtained;
[0066] Using the fuzzy inference system model and the skill level corresponding to each skill performance feature, the rule of the conclusion value corresponding to the hand controller damping is constructed according to the action value corresponding to the conclusion value;
[0067] The membership degree corresponding to the level of each skill performance feature is obtained according to the following steps:
[0068] According to each skill performance feature, the evaluation value of the skill of the trainee corresponding to each skill performance feature is obtained;
[0069] Based on the membership function, the membership degree corresponding to each evaluation value and the skill level are obtained, that is, the membership degree corresponding to each skill performance feature and the skill level are obtained.
[0070] The membership function is a triangular membership function.
[0071] When constructing the rule of the conclusion value corresponding to the hand controller damping and the action value corresponding to the conclusion value, the skill level corresponding to each skill performance feature is different, then the conclusion value corresponding to the current hand controller damping is different, wherein different conclusion values correspond to an action value; wherein the action value represents the weight of different conclusion values.
[0072] The trigger strength of the trainee is calculated according to the following formula:
[0073]
[0074] In the formula, μ j (S n ) represents the membership degree corresponding to the nth skill performance feature; represents the skill state trigger strength of the trainee.
[0075] In this embodiment, according to the selected skill performance feature, the length of the trajectory for completing the same task operation in the conventional operation task is regarded as the evaluation operator proficiency feature, for example, in the operation process, the system obtains the end position information at the time of operation. The length of the trajectory of the operator in the operation process is calculated by the following method.
[0076]
[0077] Since it is not possible to pre-quantify the complex and multi-step expected operations, it is not possible to determine an ideal quantitative performance using the traditional absolute evaluation method, and the performance of the trainee can be evaluated according to the performance. Therefore, the skill evaluation method should be independent of the task. The task-independent skill level evaluation is obtained by the following calculation method:
[0078]
[0079] It should be noted that Φ represents the skill identifier, S is for the system state, S can be composed of multiple calculated Φ, or it can be a single one, depending on the skill characteristics of interest.
[0080] The fuzzy reinforcement system structure is shown in Figure 1 The system can embed prior knowledge in reinforcement learning to accelerate the learning process, and can effectively reduce the state space dimension of reinforcement learning and improve the learning efficiency in a fuzzy differentiation manner. In the training system we designed, the fuzzy interaction system converts continuous space data into fuzzy subsets as the input of the fuzzy interaction system. The environmental state is the performance of the trainee in performing the task, and the output action of the fuzzy system is determined by two parts: when there is prior knowledge, the output action is determined by expert knowledge; when there is no expert knowledge, the output action is obtained through reinforcement learning.
[0081] In this embodiment, the rules of the conclusion value corresponding to the hand controller damping and the action value corresponding to the conclusion value are constructed using the fuzzy reasoning system model and the skill level corresponding to each skill performance characteristic; specifically as follows:
[0082] The fuzzy reinforcement learning system can represent the mapping of the state set S = {s i |s i ∈S} to the action set A = {a i |a i ∈A}, where i represents the serial number, i.e. the number of selected characteristics and the number of alternative actions. The value of the Q function is used to evaluate the decision performance to determine the good and bad behavior.
[0083] In the rule part, the prerequisite, i.e. the premise of the rule, is the skill level of the trainee, which can be fixed by experience to differentiate, and the skill level is mapped to the fuzzy subset by the membership function μ(x), and the crisp value is translated into fuzzy language;
[0084] The triangular membership function is selected for representation, as shown in Figure 2 .
[0085] For the input state quantity, i.e. the skill level S n , strong fuzzy differentiation is used to ensure the readability of the rule, i.e. to satisfy:
[0086]
[0087] Define fuzzy language variable L n (v,s l ,s r ), the shape of fuzzy differentiation is determined by the following calculation,
[0088]
[0089] The fuzzy inference system consists of fuzzy antecedents and fuzzy consequents of rules, the number of rules is represented by N, the rule is represented by R(i)(1,..., N), and the fuzzy language variable of skill level is L i .
[0090] The corresponding conclusion part, i.e. the size of the damping coefficient is u, is the value of the center point of the conclusion set, which is a discrete variable. i represents the ith input value, and j represents the jth conclusion. It should be noted that the conclusion set is pre-set.
[0091] The Takagi-Sugeno fuzzy inference system model is used to construct rules, which can be expressed in the following form:
[0092] R(i): if s is L i then u is with
[0093] or u is with
[0094] ...
[0095] or u is with
[0096] By calculating the fuzzy subset language variable of the conclusion part and mapping it to the continuous space, the output value is the clear value of the hand controller damping coefficient.
[0097] Output variable:
[0098] In the formula, U j (S) represents the hand controller damping coefficient of the hand controller in the S state, i.e. the output signal corresponding to the current state of the system, represents the skill state trigger intensity of the trainee, i.e. the trigger intensity of the current observation state; The conclusion value is represented, that is, the corresponding conclusion part is represented. The output value of the system is calculated by formula (2). In a real environment, there are multiple characteristic states S, each S corresponds to a membership function, and these membership functions are combined to be For example, as shown in formula (3).
[0099] For each fuzzy rule, the output of the clear value is determined by the trigger strength, and the calculation mode is as follows: That is, In the formula, Trans represents conversion, which adopts a product form in this embodiment.
[0100] S3, based on the Q learning method in reinforcement learning, the current action value is updated according to the current action value and the next time action value, the trigger strength of the skill state of the trainee and the smoothness of the trainee in the operation process, and the optimal action value is obtained;
[0101] The optimal action value is obtained according to the following steps:
[0102] According to the current action value and the trigger strength of the trainee, the reward of the trainee in the current continuous action is obtained;
[0103] According to the next time action value and the trigger strength of the trainee, the reward of the trainee in the next continuous action is obtained;
[0104] According to the difference between the reward of the trainee in the current continuous action and the reward of the trainee in the next continuous action, and the smoothness of the trainee at the next time, the current action value is updated to obtain the optimal action value.
[0105] In this embodiment, the Q learning method in reinforcement learning is used for optimal damping policy learning, and the ε-greedy search strategy is used to select a(i,j*)(the value of the conclusion set center point) as the conclusion of the ith rule, and the rule is applied to the global behavior q(s,a) value for evaluation to determine the role of the rule in the global.
[0106] Wherein, the ε-greedy search strategy: the ε-greedy search strategy is to ensure that the initial learning is to select an action with a certain randomness.
[0107]
[0108] Q(s,a) is the corresponding action value. The action value q is assigned to each discrete action.
[0109] The global action is the quality function Q(s,a) of the continuous action, which is an estimated value, and the cumulative reward obtained by n actions, and the reward obtained each time is defined as ri .
[0110]
[0111] The final objective function is as (5), the goal is to maximize the value of Q:
[0112]
[0113] The TD-error can be calculated as:
[0114] delta t+1 = r t+1 + gamma * Q(s t+1 , a t+1 ) - Q(s t , a t ) (6)
[0115] Wherein, gamma represents the attenuation factor, indicates the importance between the future estimated value and the current real value, the future estimated value is not important than the current real value in the embodiment, therefore, it is set to 0.9.
[0116] According to the TD-error, the action value can be updated:
[0117]
[0118] Alpha is the learning rate, and is set to 0.01.
[0119] S4, according to the optimal action value corresponding to the conclusion value in the rule and the trigger strength, the optimal hand controller damping coefficient is obtained;
[0120] According to the optimal hand controller damping coefficient, the damping of the hand controller is adjusted; Wherein, the optimal hand controller damping coefficient is calculated according to formula (2).
[0121] The present application provides a kind of personnel skill training device based on variable impedance operation difficulty adjustment, for trainee hand-held hand controller to carry out skill training, comprising:
[0122] Data acquisition module, for obtaining the position information of the end of hand controller;According to the position information of the end of hand controller, obtain the multiple skill performance characteristics of trainee and the smoothness of trainee in operation process;
[0123] The data processing module is configured to obtain the membership degree corresponding to each skill performance feature and the skill level based on the membership function; obtain the skill state triggering intensity of the trainee according to the membership degree corresponding to each skill performance feature; construct the rule of the conclusion value corresponding to the hand controller damping and the action value corresponding to the conclusion value by using the fuzzy inference system model and the skill level corresponding to each skill performance feature; and update the current action value based on the Q learning method in the reinforcement learning, the action value at the next moment, the skill state triggering intensity of the trainee and the smoothness of the trainee in the operation process, and obtain the optimal action value.
[0124] The damping mediation module is configured to obtain the optimal hand controller damping coefficient according to the conclusion value corresponding to the optimal action value in the rule and the triggering intensity; and adjust the damping of the hand controller according to the optimal hand controller damping coefficient.
[0125] Embodiment
[0126] 1Data acquisition and processing
[0127] The operation behavior of the trainee in the training process is collected, the speed information of the trainee in the training process is obtained by the sensor arranged at the end of the hand controller, and a feature state set is formed. After the data is collected and reasonably processed, the data is fuzzified and converted into a fuzzy subset of language variables, which is input into the fuzzy interaction system.
[0128] 2Fuzzy reinforcement learning execution process
[0129] (1) First, initialize the discrete action q value and the global continuous action Q value to 0;
[0130] (2) Calculate
[0131] (3) Select a discrete action based on the ε-greedy mechanism according to the rule;
[0132] (4) Calculate the continuous action U;
[0133] (5) Calculate Q(s t ,a t ) according to the continuous action;
[0134] (6) Execute the continuous action U, i.e. the damping coefficient D; it should be noted that D represents the concept of damping itself, and U is a decision made by the system, which is a control signal for the training system, but is equivalent in value;
[0135] (7) Obtain the reward r t+1 , and observe the new state s t+1 ;
[0136] (8) Estimate new Q value Q(s t+1 ,a t+1 ) according to new state and reward, and calculate time difference error according to formula (6);
[0137] (9) Update q value according to formula (7);
[0138] 3. Hand controller control
[0139] Through each operation, the hand controller guides the damping value through FQL adjustment, selects the action corresponding to the maximum Q value to perform control action selection.
[0140] 4. Difficulty strategy update
[0141] The skill of the trainee is dynamically changed in the training process, and therefore the relative operation difficulty is also changed. When the relative difficulty is changed, the damping parameter is adjusted according to the new state information at this time to generate a new operation difficulty strategy.
[0142] Embodiment: Assuming that a trainee performs operation, the generated speed, acceleration, and interaction force information are input into the training system. The system adjusts the optimal damping parameter through reinforcement learning. The result is shown in the following diagram.
[0143] Figure 3 It is shown that the reward value gradually converges after 1000 training. Figure 4 It is shown that the damping coefficient is changed.
[0144] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A personnel skill training method based on variable impedance operation difficulty adjustment, characterized by, The application relates to a skill training method for a trainee holding a hand controller, which comprises the following steps: acquiring position information of the hand controller end; acquiring multiple skill performance characteristics of the trainee and smoothness of the trainee in the operation process according to the position information of the hand controller end; acquiring a skill level corresponding to each skill performance characteristic and a membership degree corresponding to the skill level based on a membership function; acquiring skill state triggering intensity of the trainee according to the membership degree corresponding to each skill performance characteristic; mapping a conclusion value corresponding to system damping according to the skill level corresponding to each skill performance characteristic by using a fuzzy inference system model, and constructing a rule between skill performance and hand controller damping; acquiring an optimal action value of the skill triggered in the rule based on a Q learning method in reinforcement learning and feedback information of the trainee in the operation process, updating the current action value according to an estimation of the current action value and the action value at the next moment, and acquiring the optimal action value; acquiring an optimal hand controller damping coefficient according to the optimal action value of the skill triggered in the rule; adjusting the damping of the hand controller according to the optimal hand controller damping coefficient; when the skill level corresponding to each skill performance characteristic is different during the construction of the rule between skill performance and hand controller damping, the conclusion value corresponding to the current hand controller damping is different, wherein the different conclusion values correspond to an action value; the action value represents the weight value of the different conclusion values.
2. The personnel skill training method based on the variable impedance operation difficulty adjustment according to claim 1, characterized in that, The smoothness of the trainee in the operation process is calculated by calculating the operation smoothness of the trainee under the current damping state, and the smoothness constructs a reward value in the learning process, and the calculation formula is as follows: In the formula, represents a reward value acquired from the training environment; represents a third-order differential value of the displacement value generated by the trainee.
3. The personnel skill training method based on the variable impedance operation difficulty adjustment according to claim 1, characterized in that, the membership degree corresponding to the skill level of each skill performance characteristic is acquired according to the following steps: acquiring an evaluation value of the skill level corresponding to each skill performance characteristic of the trainee according to each skill performance characteristic; acquiring a membership degree corresponding to each skill level evaluation value based on a membership function, that is, acquiring the membership degree corresponding to the skill level of each skill performance characteristic.
4. The personnel skill training method based on the variable impedance operation difficulty adjustment according to claim 3, characterized in that, The membership function is a triangular membership function.
5. The personnel skill training method based on the variable impedance operation difficulty adjustment according to claim 1, characterized in that, The triggering intensity is calculated according to the following formula: In the formula, represents the membership degree corresponding to the nth skill level; represents the trigger strength.
6. The personnel skill training method based on the variable impedance operation difficulty adjustment according to claim 1, wherein, The optimal action value is acquired according to the following steps: acquiring a reward of the trainee under the damping value generated by the current system according to the skill characteristic level of the trainee and the corresponding conclusion value; estimating the action value at the next moment according to the reward value acquired from the environment at the current moment, updating the current action value by calculating the difference between the action value at the next moment and the current value, and acquiring the optimal action value.
7. The personnel skill training method based on the variable impedance operation difficulty adjustment according to claim 1, characterized in that, The hand controller damping coefficient is calculated according to the following formula: In the formula, represents the hand controller damping coefficient of the hand controller in the S state; represents the skill state trigger intensity of the trainee; represents the conclusion value.
8. The personnel skill training method based on the variable impedance operation difficulty adjustment according to claim 1, characterized in that, The multiple skill performance characteristics include displacement length of the trainee in the operation process, operation time of the trainee in the operation process and operation accuracy of the trainee.
9. A device for training personnel's skills based on impedance variation operation difficulty adjustment, characterized in that, The application relates to a skill training method for a trainee holding a hand controller, which comprises the following steps: a data acquisition module is used for acquiring position information of the hand controller end; acquiring multiple skill performance characteristics of the trainee and smoothness of the trainee in the operation process according to the position information of the hand controller end; a data processing module is used for acquiring a skill level corresponding to each skill performance characteristic and a membership degree corresponding to the skill level based on a membership function; According to the membership degree corresponding to each skill performance characteristic, the skill state trigger intensity of the trainee is obtained; According to the skill level corresponding to each skill performance characteristic, the conclusion value corresponding to the system damping is mapped by using a fuzzy inference system model, so as to construct the rule between the skill performance and the hand controller damping; Based on the Q learning method in the reinforcement learning, the feedback information, i.e. the smoothness of the trainee in the operation, is obtained from the training environment, the current action value is updated according to the estimation of the current action value and the action value at the next moment, and the optimal action value is obtained; The damping adjustment module is used for obtaining the optimal hand controller damping coefficient according to the optimal action value corresponding to the triggered skill in the rule; The damping of the hand controller is adjusted according to the optimal hand controller damping coefficient; When the skill level corresponding to each skill performance characteristic is different during the construction of the rule between the skill performance and the hand controller damping, the conclusion value corresponding to the current hand controller damping is different, wherein the different conclusion values correspond to an action value; the action value represents the weight of the different conclusion values.
Citation Information
Patent Citations
Variable-admittance teleoperation control method with fusion of multi-information
CN105242533A
Training robot, training robot system and training robot control method
JP2002127058A