A method for controlling a robot

Through real-time acquisition of environmental data and dynamic adjustment of security thresholds, the problem of insufficient security when interacting with humans is solved, and a higher security and personalized interactive experience is achieved.

CN119260709BActive Publication Date: 2025-06-13TIETECH SUZHOU CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411357029.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-06-13
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

When existing robot control systems interact with humans, it is difficult to dynamically adjust the distance safety threshold and the applied force safety threshold, resulting in insufficient security in different users and scenarios.

Method used

The environment data is collected in real time through multiple preset sensors, image processing is performed to identify the user's location, the distance between the computer robot and the user and the applied force is real-time, and the security threshold is dynamically adjusted using reinforcement learning and genetic algorithms to trigger an early warning mechanism to ensure safety.

Benefits of technology

It significantly improves the security of the robot and user interaction process, provides a personalized interactive experience, improves user satisfaction and interaction comfort, and ensures that the robot maintains optimal operating status in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119260709B_ABST
    Figure CN119260709B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for controlling a robot, specifically related to the technical field of robot control, including the following steps: real-time collecting robot operating environment data through sensors to identify the position and size of the user interacting with the robot; calculating the distance and contact force between the robot and the user, and adjusting the applied force in real time; triggering an early warning mechanism when the distance or force exceeds the safety range; optimizing the control strategy using reinforcement learning based on historical interaction data, and dynamically adjusting the distance safety threshold and the applied force safety threshold to ensure safety and adaptability during the interaction process; the present invention can not only significantly improve the safety during the interaction between the robot and the user, but also dynamically adjust the control strategy according to the feedback of different users to achieve a personalized interaction experience, and the robot can be optimized according to the needs and safety sensitivities of different users, thereby improving user satisfaction and interaction comfort.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot control, and more specifically, the present invention relates to a method for controlling a robot. Background Art

[0002] With the rapid development of robot technology, robots are increasingly widely used in industrial automation, the service industry, and the medical field. However, existing robot control systems still face a contradiction between safety and flexibility when interacting with humans. Traditional fixed-threshold control methods are difficult to adapt to complex and changeable environments, which may lead to insufficient safety in different user and scenario situations. There is an urgent need in the prior art for an optimized method that can dynamically adjust the distance safety threshold and the applied force safety threshold to ensure the safety of the robot during interaction with the user and be able to make personalized adjustments according to different user and environmental requirements. This requires the introduction of a more flexible and self-adaptive optimization control method to improve the reliability and user experience of the robot system in practical applications. Therefore, a method for controlling a robot is proposed herein. Summary of the Invention

[0003] To achieve the above object, the present invention provides the following technical solutions:

[0004] A method for controlling a robot, comprising the following steps:

[0005] Step 1: Real-time collect environmental data of the robot's operating environment through multiple preset sensors to obtain an operating environment data set;

[0006] Step 2: Based on the operating environment data set, perform image processing to identify the position and size of the user currently interacting with the robot;

[0007] Step 3: Real-time calculate the distance between the robot and the current interacting user, and obtain the acting force during contact in the interaction action and real-time adjust the force applied by the robot;

[0008] Step 4: When the distance between the robot and the current interacting user or the force applied by the robot exceeds the safety range, the warning mechanism is triggered;

[0009] Step 5: According to historical interaction data, use reinforcement learning to optimize the robot control strategy, and obtain the interaction operating state and interaction situation of the robot, and dynamically adjust the distance safety threshold and the applied force safety threshold.

[0010] In a preferred embodiment, performing image processing to identify the position and size of the user currently interacting with the robot means:

[0011] Use the pre-trained convolutional neural network to perform image recognition and classification on the image data in the operating environment dataset to locate the user's position. The output result of the target detection box of the convolutional neural network is: (x, y, w, h); x represents the X-axis coordinate of the center point of the user currently interacting with the robot, y represents the Y-axis coordinate of the center point of the user currently interacting with the robot, w represents the width of the rectangular box established by the user currently interacting with the robot, and h represents the height of the rectangular box established by the user currently interacting with the robot. Then, obtain the depth data z detected by the depth camera, and after summarization, obtain the combined result: (x, y, z, w, h); the depth data z represents the Z-axis coordinate of the center point of the user currently interacting with the robot.

[0012] In a preferred embodiment, the distance between the robot and the current interactive user in real time refers to:

[0013] Obtain the coordinate point of the robot and the combined result (x, y, z, w, h), and then use the Euclidean distance formula to calculate the distance D between the robot and the user.

[0014] In a preferred embodiment, obtaining the acting force during contact in the interaction action and adjusting the force exerted by the robot in real time refers to:

[0015] Obtain the acting force when the robot contacts the user. When the acting force does not exceed the safe range, adjust the acting force in real time. Take the acting force and the contact speed when the robot contacts the user as input variables and convert them into fuzzy values, take the adjustment factor as the output variable, then formulate fuzzy rules, map the multiple input fuzzy values to the output fuzzy values, and synthesize the output fuzzy sets of multiple rules to obtain the final fuzzy output set;

[0016] Use the centroid method for defuzzification. The centroid method formula is expressed as:

[0017] μ sc (x) represents the membership function of the output fuzzy set, x is the numerical value of the output variable, xmin and xmax represent the minimum and maximum values of the output variable x, and μ represents the adjustment factor;

[0018] The expression for adjusting the force exerted by the robot is:

[0019] F0 = Fi * μ; Fi represents the acting force when the robot contacts the user, and F0 represents the force exerted by the robot after adjustment.

[0020] In a preferred embodiment, when the warning mechanism is triggered, a warning signal is issued, or the current action is actively withdrawn, or switched to the safe mode. When the distance between the robot and the current interacting user or the force exerted by the robot exceeds the safe range, the distance between the current robot and the current interacting user and the force exerted by the current robot are obtained and used as the input data for fuzzy inference. The response type entered when the warning mechanism is triggered is used as the output data, and fuzzy inference is used to determine the response type to enter and execute.

[0021] In a preferred embodiment, using reinforcement learning to optimize the robot control strategy based on historical interaction data means that:

[0022] The robot perceives the environmental state s, then selects an action a according to the current policy. Subsequently, the robot executes the action a in the environment. After the robot executes the action, the new state s′ of the environment is obtained and the corresponding reward r is obtained, and the Q value is updated. The expression is:

[0023] α is the learning rate, which controls the balance between new and old information. γ is the discount factor, which determines the importance of future rewards. Q new (s,a) represents the new Q value of taking action a in the environmental state s after this update, and Q old (s,a) represents the old Q value of taking action a in the environmental state s before this update;

[0024] The new state s′ is used as the current state for the update loop until the termination condition is reached. When the Q value function converges, the action with the largest Q value is selected by the robot in each environmental state as the optimal policy.

[0025] In a preferred embodiment, the genetic algorithm is used to dynamically adjust the distance safety threshold and the applied force safety threshold in the robot control strategy optimized by reinforcement learning.

[0026] In a preferred embodiment, the specific steps of the genetic algorithm are as follows:

[0027] Coding and initial population: Randomly generate a group of initial individuals (population), and each individual represents a combination of a distance safety threshold and an applied force safety threshold;

[0028] Fitness evaluation: Use the comfort index to measure the fitness value of each chromosome;

[0029] Selection operation: Use the roulette wheel selection method to screen the offspring as the new parents;

[0030] Crossover operation: Randomly exchange the data in the chromosomes of different parents;

[0031] Mutation operation: Randomly select data from different offspring chromosomes for adjustment;

[0032] Repeat the selection operation, crossover operation, and mutation operation until a preset termination condition is reached, and output the individual with the lowest fitness as the combination of the optimal distance safety threshold and the applied force safety threshold.

[0033] In a preferred embodiment, the comfort index refers to:

[0034] Collect the feedback data of the interacting user to obtain a feedback data set, and then extract the number of times the feedback button is pressed by the interacting user per unit time, the average force, the average depth, and the average change amplitude of the skin conductivity from the feedback data set, and then perform weighted summation to obtain the comfort index.

[0035] Technical effects and advantages of the present invention:

[0036] By collecting environmental data in real time and making dynamic adjustments, the present invention can significantly improve the safety during the interaction between the robot and the user. The triggering conditions of the warning mechanism ensure that when the distance or the applied force exceeds the safe range, the robot can quickly respond to avoid potential harm to the user. By using reinforcement learning and genetic algorithms, the present invention can dynamically adjust the control strategy according to the feedback of different users to achieve a personalized interaction experience. The robot can be optimized according to the needs and safety sensitivities of different users, thereby improving user satisfaction and interaction comfort.

[0037] By using deep learning, fuzzy logic control systems, and genetic algorithms, the present invention enables the robot to adaptively adjust control parameters in complex and changing environments. Whether in a changing environment or facing different users, the robot can maintain the best operating state to ensure the stability and effectiveness of the operation. By combining reinforcement learning and genetic algorithms, the present invention can continuously optimize the robot control strategy. Reinforcement learning is used to learn and improve the operation strategy, while genetic algorithms are used to fine-tune the key safety thresholds to make the control strategy more flexible and precise, ensuring that the robot can operate efficiently while ensuring safety.

[0038] The fitness function designed in the present invention can effectively reduce the discomfort and insecurity that may occur during the interaction between the user. By statistically analyzing the user feedback data and optimizing it, the robot can provide a more smooth and natural interaction experience, which is suitable for application scenarios with high requirements for user experience. The method of the present invention is not only applicable to specific robot control scenarios but also has broad application potential. Whether in the fields of industrial automation, home service, or medical care, this method can provide reliable safety guarantees and personalized interaction experiences to meet the needs of different application scenarios. Description of the Drawings

[0039] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings;

[0040] Figure 1 It is a schematic diagram of a method for controlling a robot in the present invention. Specific embodiments

[0041] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0042] Refer to Figure 1 The following embodiments are obtained:

[0043] Embodiment 1: A method for controlling a robot, comprising the following steps:

[0044] Step 1: Real-time collect environmental data of the robot operating environment through a plurality of preset sensors such as force sensors, proximity sensors, cameras, ultrasonic sensors, etc., to obtain an operating environment data set;

[0045] Step 2: Based on the operating environment data set, perform image processing to identify the position and size of the user currently interacting with the robot; performing image processing to identify the position and size of the user currently interacting with the robot means:

[0046] Use a pre-trained convolutional neural network to perform image recognition and classification on the image data in the operating environment dataset to locate the user's position. The output result of the target detection box of the convolutional neural network is: (x, y, w, h); x represents the X-axis coordinate of the center point of the user currently interacting with the robot, y represents the Y-axis coordinate of the center point of the user currently interacting with the robot, w represents the width of the rectangular box established by the user currently interacting with the robot, and h represents the height of the rectangular box established by the user currently interacting with the robot. These parameters together describe a rectangular box that frames the area of the target object in the image. Through this rectangular box, the robot can identify the two-dimensional plane position of the target object in the image and the area it occupies. In traditional 2D image processing, the image itself has only two dimensions (i.e., height and width), so the detection box output by the CNN also has only four parameters (x, y, w, h), which are specifically used to describe the position information on the two-dimensional plane. Since there is no depth information (i.e., the third dimension z) in the image, the CNN can only give information on the plane. Therefore, in order to obtain the depth information z (i.e., distance) of the target object relative to the robot, additional sensors are usually required, such as: a depth camera that can directly provide the depth information z of the object; a lidar (LIDAR): measures distance information and combines with 2D images to generate 3D coordinates; stereo vision: calculates the depth information through the disparity between two cameras. In the embodiments of the present invention, the depth data z detected by the depth camera is selected and then aggregated to obtain a joint result: (x, y, z, w, h); the depth data z represents the Z-axis coordinate of the center point of the user currently interacting with the robot. Such an operation can enable the robot to expand its sensing range, not only recognize objects on the plane but also know the distance of the objects, forming a complete three-dimensional space perception. At this time, z as the depth information will be included in the output result to form a complete description of the joint result (x, y, z, w, h).

[0047] Step 3: Continuously calculate the distance between the robot and the current interacting user in real time, and obtain the force at the time of contact during the interaction action and adjust the force exerted by the robot in real time; obtain the coordinate points of the robot and the joint result (x, y, z, w, h), and then use the Euclidean distance formula to calculate the distance D between the robot and the user, (x r ,y r ,z r ) represents the coordinate position of the robot, and (x u ,y u ,z u ) represents the coordinate position of the current interacting user. Generally, the greater the distance D between the robot and the user, the safer it is.

[0048] The fuzzy logic control system is a control method based on fuzzy set theory, which is particularly suitable for dealing with systems with uncertainty or ambiguity. In the regulation of the force exerted by the robot, the fuzzy logic control system can determine the appropriate output force according to the sensor input and user feedback. The following details this process and explains how the regulation factor is obtained:

[0049] The fuzzy logic control system determines the output through the following steps:

[0050] Fuzzification: Convert the exact value of the sensor input into a fuzzy value. The input usually involves multiple variables, such as the current contact force detected by the force sensor, the contact speed, the distance between the user and the robot, etc. Fuzzy inference: According to the preset fuzzy rules, combined with the input fuzzy values, calculate the fuzzy output. Defuzzification: Convert the result of the fuzzy inference into a specific numerical output, that is, the output regulation factor.

[0051] Obtain the acting force when the robot contacts the user. When the acting force does not exceed the safe range, adjust the acting force in real time. Take the acting force and the contact speed when the robot contacts the user as input variables and convert them into fuzzy values. Take the regulation factor as the output variable, and then formulate fuzzy rules to map the fuzzy values of multiple inputs to the fuzzy value of the output. Synthesize the output fuzzy sets of multiple rules to obtain the final fuzzy output set;

[0052] In the present invention, for example: Assume that the system has two input variables: contact force E: the current contact force between the user and the robot detected; contact speed V: the current contact speed between the user and the robot. Convert these input variables into fuzzy values. For example, the contact force E can be defined as "low", "medium", "high", and the contact speed V can be defined as "fast", "moderate", "slow". Each fuzzy value is a fuzzy set, usually represented by a membership function (such as a triangular membership function). Based on experience and system requirements, formulate fuzzy rules, for example: Rule 1: If E is "high" and V is "fast", then μ should be "decrease". Rule 2: If E is "low" and V is "slow", then μ should be "maintain". Rule 3: If E is "medium" and V is "moderate", then μ should be "slightly decrease". These rules map the fuzzy values of multiple inputs to the fuzzy value of the output.

[0053] Through the fuzzy inference system, match the input fuzzy values with the preset fuzzy rules and calculate the output fuzzy set. This process includes: calculating the rule activation strength: determining which rules are activated and calculating the activation strength according to the input fuzzy values. Synthesizing the fuzzy output: Synthesize the output fuzzy sets of multiple rules to obtain the final fuzzy output set.

[0054] The defuzzification process converts the result of fuzzy reasoning into a specific numerical output μ. Common defuzzification methods include the centroid method, which calculates the centroid of the fuzzy set to obtain the specific value of μ. The centroid method formula is expressed as:

[0055] μ sc (x) represents the membership function of the output fuzzy set, x is the value of the output variable, xmin and xmax represent the minimum and maximum values ​​of the output variable x, and μ represents the adjustment factor;

[0056] The expression for adjusting the force applied by the robot is: F0 = Fi * μ; Fi represents the force when the robot contacts the user, and F0 represents the force applied by the robot after adjustment. Through the fuzzy logic control system, the adjustment factor μ is obtained based on the input sensor data and preset rules. In actual operation, μ enables the robot to make reasonable force output adjustments in a complex and fuzzy environment, thereby achieving the predetermined task goals and ensuring the safety of the user during the interaction process.

[0057] Step 4. When the distance between the robot and the current interactive user or the force applied by the robot exceeds the safety range, the early warning mechanism is triggered. When the early warning mechanism is triggered, a warning signal is issued or the current action is actively withdrawn or switched to the safety mode. When the distance between the robot and the current interactive user or the force applied by the robot exceeds the safety range, the distance between the current robot and the current interactive user and the force applied by the current robot are obtained and used together as the input data of the fuzzy reasoning. The response type entered when the early warning mechanism is triggered is used as the output data, and the fuzzy reasoning is used to determine the response type that should be entered and execute it.

[0058] The triggering condition of the early warning mechanism is when the distance between the robot and the user is too close or the force applied by the robot exceeds the safety threshold. When triggered, the robot needs to respond immediately to prevent harm to the user. Once the early warning mechanism is triggered, the system will immediately obtain the current distance between the robot and the user and the force applied by the robot. These data are the key inputs for determining how to respond next.

[0059] The obtained distance and force data are used as the input of the fuzzy inference system. The fuzzy inference system converts these precise data into fuzzy variables. For example, the distance can be described as "near" or "far", and the force can be described as "high" or "low". This fuzzification process can better handle problems of uncertainty and ambiguity. The fuzzy inference system determines the response measures according to the preset fuzzy rules. For instance, if it detects that the distance is very close and the applied force is high, the system may consider the current situation very dangerous and thus decide to switch to the safe mode. If it detects that the distance is relatively close but the applied force is small, it may only need to issue a warning signal. If the distance is far and the applied force is low, no special response measures may be taken. According to the output of the fuzzy inference system, the system determines the most suitable response type for the current situation. The possible response types include: issuing a warning signal, actively withdrawing the current action, or switching to the safe mode. Finally, according to the result of the fuzzy inference, the corresponding response measures are executed. When issuing a warning signal, the robot may remind the user through sounds, light signals, etc.; if it chooses to withdraw the action, the robot will pause the current operation and return to a safe position; when switching to the safe mode, the robot will reduce the operation speed or stop all high-risk actions to ensure the safety of the user.

[0060] Step Five: According to the historical interaction data, use reinforcement learning to optimize the robot control strategy, and obtain the interactive operation status and interaction situation of the robot, and dynamically adjust the distance safety threshold and the applied force safety threshold.

[0061] Using reinforcement learning to optimize the robot control strategy according to the historical interaction data means that:

[0062] Reinforcement learning is a learning method based on trial and error. The robot executes actions in the environment and updates its strategy according to the results (rewards or punishments) brought by the actions, and gradually learns how to select the best actions in different states to maximize the long-term benefits.

[0063] The basic components of reinforcement learning: State s: The environmental state where the robot is located, such as the current distance, applied force, speed, etc. Action a: The operations that the robot can execute in the current state, such as moving, turning, grasping, stopping, etc. Reward r: The feedback obtained by the robot from the environment after executing a certain action, which can be positive (reward) or negative (punishment). Policy π(s): The mapping from the state to the action, indicating the probability that the robot selects action a in state s. Function Q(s,a): Evaluates the cumulative rewards that may be obtained in a certain state or state-action pair in the future.

[0064] Initialize the Q-value function Q(s,a), which represents the expected return of taking action a in state s. Usually, the Q-values can be initialized randomly. Set the learning rate α, discount factor γ, and exploration rate ∈.

[0065] The robot perceives the environmental state s (e.g., obtains environmental information through sensors), then selects an action a according to the current policy. A commonly used method is the ε-greedy policy, that is, randomly select an action with a probability of ε, and select the action with the highest current Q value with a probability of 1-ε. Subsequently, the robot executes the action a in the environment. After the robot executes the action, it obtains the new state s′ of the environment and obtains the corresponding reward r, and updates the Q value. The expression is:

[0066] a is the learning rate, which controls the balance between new and old information. γ is the discount factor, which determines the importance of future rewards. Q new (s,a) represents the new Q value of taking action a in the environmental state s after this update. Q old (s,a) represents the old Q value of taking action a in the environmental state s before this update;

[0067] Take the new state s′ as the current state to update the loop, that is, repeat the above steps until the termination condition is reached (such as completing the task or reaching the maximum number of steps). Continuously iterate the above process, and the Q value function can be gradually updated until it converges to the optimal policy. When the Q value function converges, the robot selects the action with the largest Q value in each environmental state as the optimal policy.

[0068] This means that in the state s, the robot will select the action a that can obtain the maximum long-term return. This action refers to the interaction action in the interaction process, which can better interact with the user. The robot can continuously optimize its policy in the process of accumulating experience. Especially in a dynamically changing environment, reinforcement learning can adaptively adjust the policy to cope with new challenges. The new challenges can refer to different interacting users. Different interacting users have different degrees of interest in different interaction actions. Based on reinforcement learning, personalized interactions of different interacting users can be formed, rather than fixed interaction actions, which can stimulate more interaction desires of interacting users and ensure a good interaction experience.

[0069] Use the genetic algorithm to dynamically adjust the distance safety threshold and the applied force safety threshold in the robot control strategy optimized by reinforcement learning. The primary purpose of dynamically adjusting the distance safety threshold and the applied force safety threshold is to ensure the safety during the interaction process. Reinforcement learning has optimized the robot's control strategy in various situations, but in practical applications, individual differences in the environment and users may lead to different safety requirements. The genetic algorithm can further fine-tune these thresholds to adapt to the safety requirements in specific scenarios, thereby minimizing risks.

[0070] The needs and security sensitivities of each user are different. By evaluating and optimizing these thresholds, the genetic algorithm can achieve personalized adjustments. Based on the feedback data of different users, the genetic algorithm can automatically adjust the distance safety threshold and the applied force safety threshold when the robot interacts with a specific user, enabling the robot to better adapt to the personalized needs of each user and thus providing a safer and more comfortable interaction experience.

[0071] Another important purpose of dynamic adjustment is to enhance the adaptability of the robot control system. In complex and changing environments, preset fixed thresholds may not meet the requirements of all scenarios. By continuously evolving and adjusting these thresholds, the genetic algorithm enables the robot to maintain the best operating state in different environments, ensuring both safety. Through optimizing the distance safety threshold and the applied force safety threshold, the robot can achieve a more effective and natural interaction process on the premise of ensuring safety. This optimization makes the robot more flexible and effective in actual operations. In this way, the robot can not only better protect the safety of users, but also provide more customized and flexible operations according to the needs of different users and environments, thus enhancing the overall user experience and the control performance of the robot.

[0072] The specific steps of the genetic algorithm are as follows:

[0073] Coding and initial population: Randomly generate a group of initial individuals (population), and each individual represents a combination of a distance safety threshold and an applied force safety threshold; Chromosome representation: Assuming that the distance safety threshold dnew and the applied force safety threshold Fnew are the optimization objectives, and they are represented by real numbers respectively, then an individual (chromosome) Individual can be represented as: Individual = [dnew, Fnew].

[0074] Fitness evaluation: Use the comfort index to measure the fitness value of each chromosome, define the fitness function to evaluate the quality of each individual (i.e., each group of thresholds). The fitness function should reflect the safety and effectiveness of the robot operation, and the effectiveness refers to meeting the interaction acceptance degree of the interacting user.

[0075] Selection operation: Use the roulette wheel selection method to screen the offspring as the new parents; retain the individual with the lowest fitness in the current population to ensure that the optimal solution is not lost.

[0076] Crossover operation: Randomly exchange data in different parental chromosomes; Mutation operation: Randomly select data in different offspring chromosomes for adjustment; Repeat the selection operation, crossover operation, and mutation operation until a preset termination condition is reached. Common termination conditions include: reaching the preset maximum number of iterations or the change rate of the population fitness being lower than a certain threshold (i.e., the population fitness converges), and output the individual with the lowest fitness as the combination of the optimal distance safety threshold and the applied force safety threshold.

[0077] During the operation of the robot, provide a real-time feedback button or touch interface, and users can press the button at any time to express their discomfort or safety concerns about the current operation. Data processing: Count the number of times and time points when the user presses the button. If the button is pressed frequently or concentrated in a specific operation period, it indicates that there may be comfort problems with this operation. At the same time, monitor the user's physiological parameters through wearable devices. A higher skin conductivity may indicate that the user is in a tense state. Collect the feedback data of the interacting users to obtain a feedback dataset, and then extract the number of times N that the interacting users press the feedback button per unit time, the average force F, the average depth D, and the average value of the change range of skin conductivity ΔG from the feedback dataset. The fitness function should combine these factors to optimize the robot control strategy. A weighted comprehensive fitness function can be designed, where each index is assigned different weights according to its importance. The goal of the fitness function is to minimize the performance of these insecurities and discomforts. Therefore, we hope that the value of the fitness function is as low as possible. The comfort index Fitness refers to: performing a weighted sum of the number of times the interacting users press the feedback button per unit time, the average force, the average depth, and the average value of the change range of skin conductivity extracted from the feedback dataset to obtain the comfort index, Fitness = w1*N + w2*F + w3*S + w4*ΔG; w1, w2, w3, and w4 respectively represent the preset weight coefficients of the number of times N that the interacting users press the feedback button per unit time, the average force F, the average depth S, and the average value of the change range of skin conductivity ΔG, which are used to adjust the influence degree on the comfort index, and w1, w2, w3, and w4 are all positive values.

[0078] The smaller the comfort index is, the higher the acceptance of the current interaction behavior by the interacting user is. Therefore, with the goal of minimizing the fitness function value, the genetic algorithm will adjust the parameters of the robot control strategy to find the best solution that can minimize the user's discomfort and sense of insecurity. Optimization process: In each generation of the genetic algorithm, individuals are selected, crossed, and mutated according to the value of the fitness function, gradually optimizing the control strategy, reducing the frequency, force, speed of the user pressing the button, and the change range of skin conductivity, and ultimately improving the safety and comfort of robot interaction. Through this fitness function, the robot control strategy can gradually reduce the user's negative reactions during the optimization process, enhance the operation experience, and ensure a safe and comfortable interaction experience in different operation scenarios. This design method is particularly suitable for application scenarios where safety is critical or the user experience is of utmost importance.

[0079] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0080] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0081] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0082] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0083] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for controlling a robot, characterized in that: The following steps are involved: Step 1: Collect environmental data of the robot's operating environment in real time through multiple preset sensors to obtain an operating environment data set; Step 2: Based on the operating environment dataset, image processing is performed to identify the position and size of the user currently interacting with the robot; Step 3: Calculate the distance between the robot and the current interactive user in real time, obtain the contact force during the interactive action, and adjust the force applied by the robot in real time; Step 4: When the distance between the robot and the current interactive user or the force applied by the robot exceeds the safety range, the warning mechanism is triggered; Step 5: Based on historical interaction data, use reinforcement learning to optimize the robot control strategy, obtain the robot's interactive operation status and interaction situation, and dynamically adjust the distance safety threshold and the applied force safety threshold; Obtaining the forces at contact during interaction and adjusting the forces applied by the robot in real time means: The force when the robot contacts the user is obtained. When the force does not exceed the safety range, the force is adjusted in real time. The force when the robot contacts the user and the contact speed are used as input variables and converted into fuzzy values. The adjustment factor is used as the output variable. Then, fuzzy rules are formulated to map multiple input fuzzy values ​​to output fuzzy values. The output fuzzy sets of multiple rules are synthesized to obtain the final fuzzy output set. Use the centroid method for defuzzification. The centroid method formula is expressed as: ; represents the membership function of the output fuzzy set, is the value of the output variable, Represents output variable The minimum and maximum values ​​of represents the regulating factor; The expression for regulating the force applied by the robot is: ; represents the force when the robot contacts the user, represents the force applied by the robot after adjustment; When the early warning mechanism is triggered, a warning signal is issued or the current action is actively withdrawn or switched to a safe mode. When the distance between the robot and the current interactive user or the force applied by the robot exceeds the safe range, the distance between the current robot and the current interactive user and the force applied by the current robot are obtained and used as input data for fuzzy reasoning. The response type entered when the early warning mechanism is triggered is used as output data, and fuzzy reasoning is used to determine the response type that should be entered and execute it; The genetic algorithm is used to dynamically adjust the distance safety threshold and the force safety threshold in the robot control strategy optimized by reinforcement learning. The genetic algorithm uses the comfort index to measure the fitness value of each chromosome. The comfort index refers to: The feedback data of interactive users are collected to obtain a feedback data set. The number of times the interactive users press the feedback button per unit time, the average strength, the average depth, and the average change in skin conductivity are extracted from the feedback data set. The weighted sum is then performed to obtain a comfort index.

2. A method for controlling a robot according to claim 1, characterized in that: Image processing to identify the position and size of the user currently interacting with the robot refers to: Use the pre-trained convolutional neural network to perform image recognition and classification on the image data in the operating environment dataset to locate the user's position. The output result of the target detection frame of the convolutional neural network is: ; Indicates the X-axis coordinate of the center point of the user currently interacting with the robot. Indicates the Y-axis coordinate of the center point of the user currently interacting with the robot. Indicates the width of the rectangle created by the user currently interacting with the robot. Indicates the height of the rectangular frame created by the user currently interacting with the robot, and then obtains the depth data detected by the depth camera , and the combined result is: ; Depth data Indicates the Z-axis coordinate of the center point of the user currently interacting with the robot.

3. A method for controlling a robot according to claim 2, characterized in that: Real-time calculation of the distance between the robot and the current interactive user refers to: Get the robot's coordinate points and joint results , and then use the Euclidean distance formula to calculate the distance D between the robot and the user.

4. A method for controlling a robot according to claim 3, characterized in that: Based on historical interaction data, using reinforcement learning to optimize the robot control strategy means: Robot senses the state of the environment , and then choose an action based on the current strategy , and then the robot performs actions in the environment , after the robot performs the action, it obtains the new state of the environment and receive corresponding rewards , update the Q value, the expression is: ; is the learning rate, which controls the balance between new and old information, is the discount factor, which determines the importance of future rewards, Indicates that after this update, the environment status Next Action The new Q value, Indicates the environment status before this update. Next Action The old Q value of The new state The update loop is used as the current state until the termination condition is reached. When the Q-value function converges, the robot chooses the action with the largest Q-value in each environmental state, which is the optimal strategy.

5. A method for controlling a robot according to claim 4, characterized in that: The specific steps of the genetic algorithm are as follows: Coding and initial population: randomly generate a group of initial individuals, each of which represents a combination of a distance safety threshold and an applied force safety threshold; Fitness evaluation: The fitness value of each chromosome is measured using the comfort index; Selection operation: Use the roulette wheel selection method to select offspring as new parents; Crossover operation: randomly exchange data in chromosomes of different parents; Mutation operation: randomly select data of different offspring chromosomes for adjustment; The selection operation, crossover operation and mutation operation are repeated until the preset termination condition is reached, and the individual with the lowest fitness is output as the combination of the optimal distance safety threshold and the applied force safety threshold.

Citation Information

Patent Citations

  • Robot real time control method based on environmental interaction

    CN107292344A