Construction method, control method, system and medium of compliant collision avoidance motion model

By constructing a compliant collision avoidance motion model, the problem of obstacle avoidance for UAVs with arms in complex environments is solved, and high-precision and safe robotic arm motion control is achieved, which is suitable for intelligent obstacle avoidance of UAVs with arms.

CN119681867BActive Publication Date: 2025-10-03STATE GRID SIJI DIGITAL TECH (BEIJING) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411574381.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-10-03
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

The robotic arm of an armed drone has difficulty in accurately and autonomously avoiding obstacles in complex environments, especially when avoiding targets on transmission lines in power grids, which poses problems with motion accuracy and safety.

Method used

A method for constructing a compliant collision avoidance motion model is adopted. By obtaining the robot arm's motion trajectory planning, building a multi-objective reward function and a flexible obstacle avoidance function, and combining the improved velocity potential field algorithm and the proximal strategy optimization algorithm, a compliant collision avoidance motion model is constructed to achieve high-precision and safe obstacle avoidance of the robot arm.

Benefits of technology

It improves the motion accuracy and safety of the UAV robotic arm in complex environments, ensures efficient obstacle avoidance in dynamic environments, and provides intelligent motion control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119681867B_ABST
    Figure CN119681867B_ABST
Patent Text Reader

Abstract

The present invention provides a method for constructing, controlling, system, and medium for a compliant collision avoidance motion model. The method includes: obtaining a motion trajectory plan corresponding to the movement of a robotic arm of an unmanned aerial vehicle (UAV) to a target object on a power transmission line in a power transmission network; constructing a multi-objective reward function corresponding to the motion trajectory plan; wherein the multi-objective reward function includes at least a motion accuracy reward function, a motion smoothness reward function, a motion energy consumption reward function, and an additional reward function; using an improved velocity potential field algorithm to construct a flexible obstacle avoidance function for obstacles within the motion area of ​​the motion trajectory plan; and using a proximal strategy optimization algorithm to construct a compliant collision avoidance motion model corresponding to the movement of the robotic arm to the target object based on the multi-objective reward function and the flexible obstacle avoidance function. The present invention can improve the motion accuracy, safety, and intelligence level of the robotic arm of an UAV.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicles (UAVs), and in particular to a construction method, a control method, a system and a medium for a compliant collision avoidance motion model. Background Art

[0002] With the widespread use of drones for aerial operations, safe and efficient flight is crucial. Existing technologies often struggle to accurately and autonomously avoid obstacles in complex environments. Summary of the Invention

[0003] In order to solve the problems of the prior art, the present invention proposes a method for constructing a compliant collision avoidance motion model, a control method, a system and a medium, aiming to improve the motion accuracy, safety and intelligence level of the robotic arm of an arm-mounted drone.

[0004] The purpose of the present invention is achieved by adopting the following technical solutions:

[0005] In one aspect, the present invention provides a method for constructing a compliant collision avoidance motion model of a robotic arm, the method comprising:

[0006] Obtaining the motion trajectory planning corresponding to the movement of the robotic arm of the UAV to the target object on the power transmission line in the power transmission network;

[0007] Constructing a multi-objective reward function corresponding to the motion trajectory planning; wherein the multi-objective reward function includes at least: a motion accuracy reward function, a motion smoothness reward function, a motion energy consumption reward function, and an additional reward function;

[0008] Using an improved velocity potential field algorithm, a flexible obstacle avoidance function is constructed for obstacles within the motion area where the motion trajectory planning is located;

[0009] A proximal strategy optimization algorithm is adopted to construct a compliant collision avoidance motion model corresponding to the movement of the robotic arm to the target object based on the multi-objective reward function and the flexible obstacle avoidance function.

[0010] Optionally, obtaining a motion trajectory plan corresponding to the movement of a robotic arm of the armed drone to a target object on a power transmission line in the power transmission network includes:

[0011] Obtaining the location information of the target object and the environment information of the drone with an arm;

[0012] Based on the forces and torques applied to the robotic arm during movement, the multi-degree-of-freedom robotic arm kinematic model corresponding to the robotic arm, the position information and the environmental information, the motion trajectory planning corresponding to the movement of the robotic arm to the target object is determined.

[0013] Optionally, the obstacle includes a dynamic obstacle, and the use of the improved velocity potential field algorithm to construct a flexible obstacle avoidance function for obstacles within the motion area where the motion trajectory planning is located includes:

[0014] Using the improved velocity potential field algorithm, construct a first function representing the direction between the end effector of the robotic arm and the target object, a second function representing the running speed of the dynamic obstacle, and a third function representing the tangential velocity between the end effector of the robotic arm and the target object;

[0015] The first function, the second function and the third function are integrated to obtain the flexible obstacle avoidance function.

[0016] Optionally, the proximal strategy optimization algorithm is used to construct a compliant collision avoidance motion model corresponding to the movement of the manipulator to the target object based on the multi-objective reward function and the flexible obstacle avoidance function, including:

[0017] Acquire a training data set; wherein the training data set includes at least: takeoff data, hovering data, movement data, rotation data, grasping data and landing data corresponding to the execution of historical basic tasks by the UAV with an arm;

[0018] Adopting the proximal strategy optimization algorithm, constructing an initial motion model based on the multi-objective reward function and the flexible obstacle avoidance function;

[0019] The training data set is used to perform reinforcement learning training on the initial motion model to obtain the compliant collision avoidance motion model.

[0020] Optionally, the using the training data set to perform reinforcement learning training on the initial motion model to obtain the compliant collision avoidance motion model includes:

[0021] Using the training data set, performing reinforcement learning training on the initial motion model to obtain an intermediate motion model;

[0022] The intermediate motion model is linearized to obtain the compliant collision avoidance motion model.

[0023] Optionally, linearizing the intermediate motion model to obtain the compliant collision avoidance motion model includes:

[0024] performing linearization processing on the intermediate motion model to obtain a motion model to be adjusted;

[0025] Based on a simulation environment having a collision detection and risk assessment module, a performance evaluation is performed on the motion model to be adjusted to obtain an evaluation result;

[0026] Based on the collected historical collision data of the robotic arm and the evaluation results, multiple rounds of iterative training and performance evaluation are performed on the motion model to be adjusted until the compliant collision avoidance motion model is obtained whose performance evaluation indicators meet preset conditions.

[0027] Correspondingly, the present invention provides a system for constructing a compliant collision avoidance motion model of a robotic arm, the system comprising:

[0028] The first acquisition module is used to obtain the motion trajectory planning corresponding to the movement of the robotic arm of the arm-carrying drone to the target object on the power transmission line in the power transmission network;

[0029] A first building module is used to build a multi-objective reward function corresponding to the motion trajectory planning; wherein the multi-objective reward function includes at least: a motion accuracy reward function, a motion smoothness reward function, a motion energy consumption reward function, and an additional reward function;

[0030] The second building module is used to build a flexible obstacle avoidance function for obstacles within the motion area where the motion trajectory planning is located using an improved velocity potential field algorithm;

[0031] A construction module is used to adopt a proximal strategy optimization algorithm to construct a flexible collision avoidance motion model corresponding to the movement of the robotic arm to the target object based on the multi-objective reward function and the flexible obstacle avoidance function.

[0032] In another aspect, the present invention provides a method for controlling an unmanned aerial vehicle with an arm, the method comprising:

[0033] Obtain the current position of the current target on the transmission line in the power grid and the current environment information of the UAV with an arm;

[0034] Inputting the current position and the current environment information into the compliant collision avoidance motion model constructed by any of the above methods to obtain a current motion trajectory;

[0035] Based on the current motion trajectory, the robotic arm of the arm-carrying drone is controlled to move to the current position.

[0036] Correspondingly, the present invention provides a control system for an arm-mounted drone, the system comprising:

[0037] The second acquisition module is used to obtain the current position of the current target object on the transmission line in the transmission network and the current environment information of the UAV with an arm;

[0038] An input module, configured to input the current position and the current environment information into the compliant collision avoidance motion model constructed by the system as described above, to obtain a current motion trajectory;

[0039] A control module is used to control the mechanical arm of the arm-carrying drone to move to the current position based on the current motion trajectory.

[0040] In another aspect, the present application further provides a computing device, comprising: at least one processor and a memory; the memory and the processor are connected via a bus;

[0041] The memory is used to store one or more programs;

[0042] When the one or more programs are executed by the at least one processor, the method for constructing a compliant collision avoidance motion model of a robotic arm as described in any one of the above items and the control method of the arm-equipped drone as described above are implemented.

[0043] On the other hand, the present application also provides a readable storage medium having an execution program stored thereon. When the execution program is executed, it implements the method for constructing a compliant collision avoidance motion model of a robotic arm as described in any of the above items, and the control method of the arm-equipped drone as described above.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] The present invention provides a method for constructing a compliant collision avoidance motion model for a robotic arm. First, a motion trajectory plan corresponding to the motion of the robotic arm of an unmanned aerial vehicle (UAV) to a target object on a transmission line in a power transmission network is obtained; then, a multi-objective reward function corresponding to the motion trajectory plan and a flexible obstacle avoidance function for obstacles within the motion area where the motion trajectory plan is located are constructed; finally, a PPO algorithm is used, combined with the multi-objective reward function and the flexible obstacle avoidance function constructed based on an improved velocity potential field algorithm, to construct a compliant collision avoidance motion model for the robotic arm to move to the target object. In this way, a highly comprehensive compliant collision avoidance motion model can be provided, that is, the compliant collision avoidance motion model can output a motion trajectory with high motion accuracy, safety, and intelligence, thereby ensuring that the compliant collision avoidance motion model is used to output the relevant motion trajectory, and can guide the UAV with arms to move with high precision, safety, and intelligence in complex and dynamic environments.

[0046] The present invention provides a control method for an arm-mounted drone, which uses a highly comprehensive compliant collision avoidance motion model, i.e., a compliant collision avoidance motion model that can improve the motion accuracy, safety, and intelligence level of the arm-mounted drone, to output the corresponding current motion trajectory, so as to guide and improve the motion accuracy, safety, and intelligence level of the robotic arm of the arm-mounted drone based on the current motion trajectory.

[0047] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions provided by the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which:

[0049] Figure 1 A schematic flow chart of a method for constructing a compliant collision avoidance motion model for a robotic arm provided in an embodiment of the present invention;

[0050] Figure 2 A schematic diagram illustrating the relationship between the joints in the robotic arm of the drone with arms provided in an embodiment of the present invention;

[0051] Figure 3 A schematic diagram of a process for constructing a compliant collision avoidance motion model using the construction method provided in an embodiment of the present invention;

[0052] Figure 4 A schematic flow chart of a control method for a drone with arms provided in an embodiment of the present invention;

[0053] Figure 5 A schematic diagram of the composition of a system for constructing a compliant collision avoidance motion model of a robotic arm provided by an embodiment of the present invention;

[0054] Figure 6 A schematic diagram of the control system of a drone with an arm provided in an embodiment of the present invention;

[0055] Figure 7 A schematic diagram of the composition of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0056] The following describes the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art will readily appreciate the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through various other specific embodiments, and the various details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are intended only to illustrate the present invention and are not intended to limit the scope of protection of the present invention.

[0057] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0058] In the following description, the terms "first\second\third" are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.

[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art in the art of the embodiments of the present invention. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the embodiments of the present invention.

[0060] With the widespread use of drones for aerial operations, safe and efficient flight is crucial. Existing technologies often struggle to accurately and autonomously avoid obstacles in complex environments.

[0061] Based on the above problems, the embodiments of the present invention propose a method for constructing a compliant collision avoidance motion model, a control method, a system and a medium, aiming to improve the motion accuracy, safety and intelligence level of the robotic arm of an armed drone.

[0062] Example 1:

[0063] See also Figure 1 The figure is a flow chart of a method for constructing a compliant collision avoidance motion model of a robotic arm provided by the present invention. Figure 1 The steps shown in the following instructions:

[0064] Step 101: Obtain a motion trajectory plan corresponding to the movement of a robotic arm of an armed drone to a target object on a power transmission line in a power transmission network.

[0065] In some embodiments of the present invention, the drone with arms is a drone designed based on a vulture prototype, which can imitate the vulture's aerial hunting. The mechanical arm installed below it, which is similar to an eagle's claw, can grab objects outside the drone's flight path.

[0066] In some embodiments of the present invention, the target object on the transmission line in the transmission network can be: related devices deployed on the transmission line, such as: aviation warning balls, bird-repellent facilities, vibration monitoring devices, etc., or, the target object can be an obstacle on the transmission line, such as: bird nests, branches, hay, kites, balloons, plastic bags, advertising cloth, etc.

[0067] In some embodiments of the present invention, the kinematic model of the robotic arm of the arm-carrying drone, the environmental information of the arm-carrying drone, the forces and torques exerted on the robotic arm during movement, and the position information of the target object can be first obtained, and then the above information obtained can be input into the initial trajectory generation module, or a relevant trajectory generation algorithm can be used to generate a motion trajectory plan corresponding to the arm-carrying drone driving the robotic arm deployed on it to move to the target object.

[0068] In some embodiments of the present invention, the above step 101 may be implemented by the following steps 1011 and 1012 (not shown in the figure):

[0069] Step 1011: Acquire the location information of the target object and the environment information of the drone with arms.

[0070] In some embodiments of the present invention, obtaining the location information of the target object can be achieved through two major steps: the first step is visual perception to detect and identify the target object on the power line; the second step is target positioning, that is, accurately positioning the target object to obtain its location information.

[0071] In some embodiments of the present invention, real-time detection and recognition of objects on power lines can be performed based on a Vision Transformers (ViT) model. Furthermore, simultaneous localization and mapping (SLAM)-based environmental modeling and positioning can be used to further achieve high-precision detection and recognition of objects on power lines.

[0072] Here, the ViT model is used to perform real-time detection and recognition of targets on the transmission line, which can be achieved through the following process:

[0073] In the first step, the target object image obtained by image acquisition can be processed, such as using filters (Gaussian filtering, mean filtering, median filtering, etc.) to reduce noise in the image to improve the quality of the target object image. The target object image can also be enhanced using histogram equalization, contrast enhancement and other technologies to improve the visual effect of the target object image, so as to provide more accurate parameter (image) support for the subsequent accurate identification of the target object in the target object image.

[0074] In the second step, target detection is performed on the image obtained in the first step (the adjusted target image) based on the ViT model to achieve real-time detection and recognition of targets on the power lines. The ViT model can be a Transformer architecture, mainly composed of two parts: a self-attention mechanism and a position encoding. (It should be noted that compared to traditional convolutional neural networks (CNNs), the ViT model can better capture global information and details in the image through self-attention and position encoding, especially showing greater flexibility and performance when processing high-resolution images and complex visual tasks.)

[0075] Here, the Self-Attention Mechanism enables the ViT model to focus on the important parts of the input image data, thereby capturing the global and local information of the image. Its main execution logic is as follows:

[0076] First, input sequence: split the input image (i.e. the adjusted target image) into fixed-size patch vectors (e.g. 16x16 pixels). Here, each patch vector can be flattened into a one-dimensional vector;

[0077] Secondly, linear projection: the flattened multiple patch vectors are projected into a high-dimensional vector space through linear transformation. Its execution logic is similar to mapping local features in the image into a unified representation space.

[0078] Then, calculate the attention weight: determine the dot product between each patch vector and all other patch vectors; where it can be normalized by the Softmax function to obtain the attention weight of each patch vector;

[0079] Finally, weighted summation is performed: Based on the attention weights assigned to each patch vector, all patch vectors are weighted and summed to obtain the final representation for each position (here, each patch vector is enriched with positional information, which can be provided by the positional encoding). The computational complexity of the self-attention mechanism is O(N^2), where N is the length of the input sequence. This mechanism enables the ViT model to capture long-range dependencies and global information in the image.

[0080] Here, because the Transformer architecture does not have inherent position information for processing sequential data, positional encoding is used to supplement this deficiency, enabling the ViT model to understand the positional relationship of elements in the image. The execution logic corresponding to positional encoding is as follows:

[0081] First, the encoding calculation: through the sine and cosine function calculation, a matrix with the same length and dimension as the input sequence mentioned above is generated to obtain the position encoding matrix;

[0082] Then, the encoding is added: the obtained position encoding matrix is ​​added to the input patch vector mentioned above and concatenated to preserve the position information of each patch vector.

[0083] Here, the encoding can also be updated. For example, during the training process of the ViT model, the position encoding can be updated to adapt to different input resolutions and model requirements. Among them, the position encoding can help the ViT model better understand the relative position and order of elements in the adjusted target image, thereby better capturing the structural information in the adjusted target image.

[0084] In some embodiments of the present invention, the training of the ViT model can be achieved by the following steps:

[0085] The first step is pre-training and fine-tuning: The ViT model can be pre-trained on large-scale datasets (such as the large-scale visualization database ImageNet and the ultra-large-scale dataset JFT-300M), and then fine-tuned on smaller datasets. The pre-training phase uses high-resolution images, and the fine-tuning phase keeps the size of the patch vector constant while increasing the length of the input sequence.

[0086] The second step is the classification head: a classification tag is added to the output of the Transformer encoder, and the final classification result is generated after processing by the Multilayer Perceptron (MLP).

[0087] Correspondingly, in some embodiments of the present invention, based on the PointNet++ multimodal neural network, combined with multiple sensor data (such as visual sensors, LiDAR (Light Laser Detection and Ranging), force feedback), the feature extraction and fusion of multimodal data can be achieved through joint training of the network to provide comprehensive environmental information and perception of the target's location information, thereby achieving high-precision target position perception for the target, to ensure accurate positioning of the target's location information in complex and dynamic environments. Among them, the application of the PointNet++ multimodal neural network in target positioning lies in its efficient three-dimensional point cloud feature extraction and hierarchical feature learning capabilities. The following are the technical details of the PointNet++ multimodal neural network for target positioning:

[0088] First, data preprocessing: Before feature extraction, the original point cloud data (i.e., the collected data used to identify the precise positioning of the target object) must first be preprocessed. Among them, point cloud sampling is a key step in preprocessing. The point cloud sampling can be done by using the farthest point sampling (FPS) method to select a set of representative points from the original point cloud data. FPS ensures uniform coverage of the sampling points by ensuring the maximum distance distribution between the sampling points. This can greatly reduce the amount of data while retaining the main geometric features of the point cloud, laying a solid foundation for subsequent feature extraction.

[0089] Secondly, hierarchical feature learning: After data preprocessing is complete, hierarchical feature learning is performed. The first step is grouping and feature extraction. Here, in each layer, the neighborhood points around the sampling point are grouped. Local regions can be formed by setting a fixed radius or a fixed number of neighborhood points. Within these local regions, the PointNet method is used for feature extraction. Here, PointNet uses an MLP to perform point-by-point operations on the features of each point, and then uses a maximum pooling operation to obtain the global features of the local region. This effectively captures the local geometric information of the point cloud, providing the necessary parameter foundation for subsequent global feature construction.

[0090] Next, hierarchical feature transfer: After local feature extraction, the hierarchical feature transfer phase begins. At each layer, features extracted from each local region are aggregated to form a higher-level feature representation. Layer-by-layer transfer and aggregation are crucial in this process. By transferring features layer by layer, new sampling, grouping, and feature extraction operations are continuously performed, gradually improving the expressive power of features and thus constructing a global feature representation. This feature transfer process ensures the fusion of features from local details to global structure, enabling the model to fully understand the overall information of the point cloud.

[0091] Finally, target positioning: After hierarchical feature learning and hierarchical transfer are complete, the target positioning phase begins. It sequentially performs the following steps: feature decoding, which uses deconvolution or interpolation techniques to map high-level features back to the original point cloud space, thereby obtaining a feature representation for each point; and performing position prediction on each point's features through the MLP to obtain the target object's position coordinates. This step remaps the information in the high-dimensional feature space back to three-dimensional space, enabling precise positioning of the target object. By combining feature decoding with position prediction, the PointNet++ multimodal neural network is able to achieve high-precision target positioning in complex three-dimensional environments, ensuring reliability and accuracy in dynamic environments.

[0092] It should be noted that through steps such as data preprocessing, hierarchical feature learning, hierarchical feature transfer, and target positioning, the PointNet++ multimodal neural network can effectively realize the comprehensive feature extraction and fusion of multimodal data, provide high-precision environment and target position perception, and provide reliable technical support for accurately positioning the position information of target objects in complex and dynamic environments.

[0093] In some embodiments of the present invention, the environmental information of the arm-mounted drone can be obtained by collecting environmental data of the arm-mounted drone through sensors or sampling facilities deployed on the arm-mounted drone. The environmental information includes but is not limited to: the height of the arm-mounted drone, ambient temperature, ambient humidity, geographical environment of multiple locations of the arm-mounted drone, wind speed, wind direction and other information.

[0094] Step 1012: Determine the motion trajectory plan corresponding to the movement of the robotic arm to the target object based on the forces and torques applied to the robotic arm during the movement, the multi-degree-of-freedom robotic arm kinematic model corresponding to the robotic arm, the position information, and the environmental information.

[0095] In some embodiments of the present invention, a multi-degree-of-freedom robotic arm kinematic model corresponding to the robotic arm of a drone with an arm can be established using the Modified Denavit-Hartenberg (MDH) parameter method to ensure the continuity of the robotic arm's subsequent motion. The corresponding execution logic may include: first, establishing a reference coordinate system for each joint in the robotic arm of the drone with an arm; then, modeling based on the MDH parameter method and the corresponding reference coordinate system to establish a 6-degree-of-freedom robotic arm kinematic model, thereby ensuring the accurate positioning and coherent motion of each joint in the robotic arm.

[0096] Here, reference coordinate system can be established for each joint. Figure 2 As shown, axis i-1 (joint axis Z i-1 ), axis i (joint axis Z i ) and axis i+1 (joint axis Z i+1 ), where axis i-1 is connected to axis i via connecting rod i-1, and axis i is connected to axis i+1 via connecting rod i. The corresponding execution steps are as follows:

[0097] The first step is to determine the origin O of the connecting rod coordinate system i . That is, find the joint axis Z i and joint axis Z i+1 The common perpendicular line between them, or the joint axis Z i With joint axis Z i+1 The intersection of the joint axis Z i With joint axis Z i+1 The intersection of the common perpendicular line and the joint axis Zi , joint axis Z i+1 The intersection point is taken as the origin O of the link coordinate system {i} i .

[0098] The second step is to stipulate Axis is along the joint axis Z i direction.

[0099] The third step is to stipulate The axis is along the common vertical line. If the joint axis Z i and joint axis Z i+1 If they intersect, then Axis is perpendicular to the joint axis Z i and joint axis Z i+1 The plane where it is located.

[0100] The fourth step is based on the determined Axis and Axis, determined by the right-hand rule Axis, corresponding, can continue to refer to Figure 2 shown.

[0101] Step 5: Repeat the above steps 1 to 4 until a joint reference coordinate system is established for each joint, such as the corresponding Z axis of the joint. i-1 On the top, establish the joint reference coordinate system, whose corresponding origin is O i-1 , and the corresponding axis, Axis and axis.

[0102] Continue to refer Figure 2 As shown, it can be seen that the coordinate system fixed on the link i-1 is through four DH parameters, namely: joint torsion angle α i-1 , connecting rod length a i-1 ( Figure 2 The connecting rod length a is also shown i ), connecting rod offset d i and joint angle θ i , transformed to the coordinate system fixed on the connecting rod i, that is:

[0103] 1.A(a i-1 )=Along Axis, from joint axis Z i-1 Move to joint axis Z i Correspondingly, A(a i )=Along Axis, from joint axis Z i Move to joint axis Z i+1 distance.

[0104] 2.alpha(αi-1 )=Wrap Axis, from joint axis Z i-1 Rotate to joint axis Z i angle.

[0105] 3.D(d i )=Along Axis, from Axis moves to Axis distance.

[0106] 4.theta(θ i )=Along Axis, from Axis moves to The angle of the axis.

[0107] Its corresponding transformation matrix i-1 T i (α i-1 ,a i-1 ,θ i ,d i ), as shown in the following formula (1):

[0108]

[0109] Among them, R X (α i-1 ), D X (a i-1 ), R Z (θ i ) and D Z (d i ) are respectively those mentioned above: joint torsion angle, connecting rod length, joint rotation angle and connecting rod offset.

[0110] In this way, the mathematical model corresponding to the robotic arm of the UAV with arms is established through the MDH parameter method, such as the 6-DOF robotic arm kinematic model, which can further simplify the description of the relationship between the links inside the robotic arm to ensure that each joint in the robotic arm can be accurately positioned and moved.

[0111] In some embodiments of the present invention, the force exerted on the robotic arm during movement can be obtained using a fiber Bragg grating (FBG) sensor. Here, the FBG sensor can be installed at the end or joint of the robotic arm, and the force can be measured by using the change in the optical propagation characteristics of the optical fiber when it is subjected to force. By real-time monitoring and feedback of the force, the safe and precise operation of the robotic arm when performing tasks is ensured, especially in the process of grasping objects using the robotic arm. Among them, the FBG sensor writes a periodic refractive index variation structure in the optical fiber. When light passes through the FBG sensor, light of a specific wavelength is reflected and light of other wavelengths is transmitted. That is, the calculation method of the corresponding Bragg wavelength λB (reflection wavelength) is determined by the following formula (2):

[0112] λB=2neffΛ formula (2);

[0113] Among them, λB is the wavelength of the light wave reflected under specific conditions, neff is the refractive index of the material in the center of the optical fiber, and Λ is the period of the spatial structure in the fiber Bragg grating.

[0114] It should be noted that when the optical fiber is subjected to stress or temperature changes, both Λ and neff will change, resulting in a shift in the reflected wavelength. By measuring the shift ΔλB of λB, the strain on the optical fiber can be calculated. and / or temperature changes The relationship between strain and temperature change is shown in formula (3):

[0115]

[0116] In some embodiments of the present invention, the Newton-Euler method can also be used to construct the dynamic equations. Based on Newton's second law and the Euler equation, the force and torque can be directly calculated. The corresponding formula can be referred to as shown in formula (4):

[0117]

[0118] Where I is the moment of inertia matrix of the particle, ∝ and ω are the angular acceleration and angular velocity of the particle respectively, M is the external torque acting on the particle, F is the force, m is the mass of the object, and a is the acceleration of the object.

[0119] In this way, based on the precisely measured forces and moments at each joint of the robotic arm, the safety and accuracy of the robotic arm of the drone can be ensured in actual operation.

[0120] In this way, first, a 6-DOF robotic arm model is established through the MDH parameter method to ensure the accurate positioning and coherent movement of each joint in the robotic arm of the arm-carrying UAV; secondly, FBG sensors are used as force / torque sensors to monitor the force and torque information borne by the robotic arm in real time to ensure the safety and accuracy of the robotic arm during operation; then, with the help of technologies such as the ViT model, accurate detection and environmental modeling of target objects and / or environmental information are performed, which can improve the recognition accuracy of the arm-carrying UAV for target objects and environmental information, and use the PointNet++ multimodal neural network combined with multiple sensor data to achieve high-precision environmental and target position perception of the target object's position information, thereby providing more accurate parameter support for motion trajectory planning, thereby improving the accuracy of the determined motion trajectory planning.

[0121] Step 102: Build a multi-objective reward function corresponding to the motion trajectory planning.

[0122] The multi-objective reward function includes at least: a motion accuracy reward function, a motion smoothness reward function, a motion energy consumption reward function, and an additional reward function.

[0123] In some embodiments of the present invention, a reinforcement learning framework can be used to convert the control strategy corresponding to the motion trajectory planning obtained in step 101 into a multi-objective optimization strategy. Correspondingly, a multi-objective reward function corresponding to the motion trajectory planning can be constructed; wherein the multi-objective optimization strategy may include: accuracy optimization strategy, smoothness optimization strategy, energy consumption optimization strategy, time optimization strategy, and pulse optimization strategy, etc. Correspondingly, the multi-objective reward function at least includes: a motion accuracy reward function corresponding to the accuracy optimization strategy, a motion smoothness reward function corresponding to the smoothness optimization strategy, a motion energy consumption reward function corresponding to the energy consumption optimization strategy, and other reward functions. In this way, by designing a comprehensive consideration of the multi-objective reward function, the proximal policy optimization (PPO) algorithm is further used to train the relevant model (the model can apply for the network architecture corresponding to the multi-objective reward function here) to obtain an optimal trajectory planning strategy that can balance the various performance indicators of the robotic arm of the arm-carrying drone during flight.

[0124] In some embodiments of the present invention, the multi-objective reward function includes at least a motion accuracy reward function, a motion smoothness reward function, a motion energy consumption reward function, and an additional reward function, which may be:

[0125] 1. Motion accuracy reward function: It is the accuracy cost function f corresponding to the motion trajectory of the robot arm at time t a (θ t ) is converted into a motion accuracy reward function Among them, the accuracy cost function f a (θ t ) can be constructed by the distance between the desired position of the end of the manipulator at time t and the current position. At the same time, the motion accuracy reward function can be referred to as shown in the following formula (5):

[0126]

[0127] Among them, ω a and σ a are the weight coefficients of the motion accuracy reward function respectively.

[0128] 2. Motion smoothness reward function: It is the smoothness function f corresponding to the motion trajectory of the robot arm at time t s (θ t ) is converted to a motion smoothness reward function Among them, the smoothness function f s (θ t ) is a function built to reduce the vibration impact caused by the movement of each joint of the manipulator at time t to achieve smooth motion trajectory. At the same time, the motion smoothness reward function can be referred to as shown in the following formula (6):

[0129]

[0130] Among them, ω s ,λ υ ,λ acc are the weight coefficients of the motion smoothness reward function, and and They respectively refer to the velocity and acceleration of the kth joint in the robotic arm, n is the number of joints in the robotic arm, and k refers to the kth joint in the robotic arm.

[0131] 3. Motion energy reward function: It is the energy cost function f corresponding to the motion trajectory of the robot arm at time t e (θ t ) is converted into a motion energy consumption reward function Among them, the energy consumption cost function f e (θ t ) is used to solve the energy consumption problem related to the movement of the joints in the robotic arm at time t. At the same time, the energy consumption reward function Please refer to the following formula (7):

[0132]

[0133] Among them, ω e represents the weight energy consumption bonus coefficient, Δθ k represents the angle change of the kth joint per unit time, τk represents the torque of the kth joint. Similarly, n is the number of joints in the robot arm, and k refers to the kth joint in the robot arm.

[0134] 4. Additional reward function: It can use the accuracy cost function to build the corresponding additional reward function Here, if the distance between the current endpoint of the robot arm and the target endpoint is less than 0.005, an additional reward of 10 is given, otherwise the corresponding reward is: 10 / (1+10f a (θ t ), which is shown in the following formula (8):

[0135]

[0136] Based on the above description, the complete reward function corresponding to the motion trajectory planning, that is, the multi-objective reward function r can be expressed as shown in the following formula (9):

[0137]

[0138] Step 103: Using an improved velocity potential field algorithm, a flexible obstacle avoidance function is constructed for obstacles within the motion area where the motion trajectory planning is located.

[0139] In some embodiments of the present invention, the Improved Velocity Potential Field (IVPF) algorithm can be used to further develop a flexible obstacle avoidance function for obstacles within the motion area of ​​the motion trajectory planning. Here, obstacles within the motion area of ​​the motion trajectory planning include: static obstacles and dynamic obstacles.

[0140] It should be noted that the detection and identification of obstacles within the motion area where the motion trajectory is planned can be perceived in real time by sensors deployed on the arm-mounted drone (such as lidar, camera), or by referring to the detection and identification of the position information of the target object described above. The present invention does not impose any limitation on this.

[0141] It should be noted that the IVPF algorithm is an improved algorithm designed to address the problems of traditional velocity potential field (VPF) algorithms. Traditional VPF algorithms are prone to getting stuck in local minima, failing to reach the target point, having poor dynamic obstacle avoidance capabilities, and facing safety and efficiency issues. The IVPF algorithm can better handle dynamic obstacles by incorporating directional factors, obstacle speed, and tangential velocity to avoid getting stuck in local minima.

[0142] In some embodiments of the present invention, if the obstacles within the motion area where the motion trajectory is planned include dynamic obstacles, the above step 103 can be implemented by the following steps 1031 and 1032 (not shown in the figure):

[0143] Step 1031: Using the improved velocity potential field algorithm, construct a first function that characterizes the direction between the end effector of the robotic arm and the target object, a second function that characterizes the running speed of the dynamic obstacle, and a third function that characterizes the tangential velocity between the end effector of the robotic arm and the target object.

[0144] In some embodiments of the present invention, the IVPF algorithm can be used to introduce direction factors into the corresponding gravitational field during the movement of the robot arm based on the information of dynamic obstacles, so as to comprehensively consider the direction between the end effector of the robot arm and the target object, thereby constructing the corresponding first function; the speed factor of the dynamic obstacle is introduced into the corresponding repulsive field during the movement of the robot arm to better avoid the dynamic obstacle, thereby constructing the corresponding second function; the tangential speed factor is introduced into the movement of the robot arm to help the robot arm get rid of the trouble of local minimum values, thereby constructing a third function that characterizes the tangential speed between the end effector of the robot arm and the target object.

[0145] Step 1032: Fuse the first function, the second function, and the third function to obtain the flexible obstacle avoidance function.

[0146] In some embodiments of the present invention, the constructed first function, second function, and third function may be integrated to form a flexible obstacle avoidance function for the robotic arm to avoid the dynamic obstacle during movement.

[0147] It should be noted that through experimental comparison (i.e., comparison of specific experimental data), it was determined that the IVPF algorithm showed significant performance improvements in handling different types of obstacles, including better obstacle avoidance and smoother trajectories. In practical applications, the IVPF algorithm is generally used to provide a safe and efficient path planning algorithm for medical / surgical robot arms to ensure the safety of personnel and robots during human-machine collaboration. This algorithm can bring more flexible and safe possibilities to human-machine collaboration in medical / surgical scenarios, opening up new development opportunities. In this invention, the IVPF algorithm is applied to the power transmission grid.

[0148] In this way, the IVPF algorithm is used to achieve flexible obstacle avoidance of the robotic arm of the UAV when it moves to the target object on the power transmission line in the power transmission network, thereby providing technical support for achieving better avoidance effect corresponding to dynamic obstacles and smoother motion trajectory.

[0149] Step 104 : Using a proximal strategy optimization algorithm, based on the multi-objective reward function and the flexible obstacle avoidance function, construct a compliant collision avoidance motion model corresponding to the movement of the robotic arm to the target object.

[0150] In some embodiments of the present invention, a network framework corresponding to the PPO algorithm can be first used to construct an initial motion model corresponding to motion trajectory planning based on a multi-objective reward function and a flexible obstacle avoidance function; and then the initial motion model can be reinforced learning using a relevant training data set to obtain a flexible collision avoidance motion model corresponding to the movement of the robotic arm to the target object (or grasping the target object).

[0151] In some embodiments of the present invention, the above step 104 may be implemented by the following steps 1041 and 1042 (not shown in the figure):

[0152] Step 1041: Obtain a training data set.

[0153] The training data set includes at least the takeoff data, hovering data, movement data, rotation data, grabbing data and landing data corresponding to the execution of historical basic tasks by the UAV with arms.

[0154] In some embodiments of the present invention, the arm-equipped drone can be remotely controlled to demonstrate relevant tasks to collect training data sets. The training data sets include various basic actions corresponding to the arm-equipped drone's execution trajectory, such as movement, rotation, and grasping. These training data sets can provide a foundation for subsequent model training.

[0155] Here, the training dataset contains data corresponding to 100-1000 demonstration actions performed by the drone with arms.

[0156] Step 1042: Using the proximal strategy optimization algorithm, based on the multi-objective reward function and the flexible obstacle avoidance function, construct an initial motion model.

[0157] In some embodiments of the present invention, a network framework corresponding to the PPO algorithm is adopted, and an initial motion model of the robotic arm, that is, a universal motion model of the robotic arm, is constructed based on a multi-objective reward function and a flexible obstacle avoidance function.

[0158] Step 1043: Use the training data set to perform reinforcement learning training on the initial motion model to obtain the compliant collision avoidance motion model.

[0159] In some embodiments of the present invention, the aforementioned training dataset can be used to perform reinforcement learning training on the initial motion model of the robotic arm to achieve fine-tuning and optimization of the initial motion model. This fine-tuning and optimization process significantly improves the performance of the resulting compliant collision avoidance motion model in performing specific tasks (e.g., movement, rotation, grasping, etc.). However, this may also result in a decrease in the performance of the resulting compliant collision avoidance motion model on the original task. To address this issue, the present solution employs the following method to enable seamless integration of new tasks into the new compliant collision avoidance motion model: Based on the fine-tuned and optimized strategy, the arm-carrying drone is typically autonomously deployed and performs tasks in a real environment to collect a large amount of new task data. Here, this new task data can be processed through "post-target relabeling" (here, post-target relabeling technology can relabel the new task data during the autonomous execution process, improving the accuracy and validity of the data) to form a self-improving dataset. The new task data can then be merged with the training data to serve as data for training the next generation of the robotic arm's motion model (i.e., the intermediate motion model obtained during the reinforcement learning training phase). Through this iterative approach, drones can continuously learn new skills and integrate them into their own knowledge system to achieve continuous evolution.

[0160] It should be noted that if an arm-mounted drone wants to achieve autonomous data collection, that is, the "collection of large amounts of new mission data" mentioned above, it needs to solve two key problems: successful detection and environmental reset.

[0161] Success detection can be achieved by training vision-based reward models to detect mission completion success. These reward models can evaluate the quality of mission execution by the arm-mounted drone in real time, providing accurate feedback for autonomous data collection. By training vision-based reward models, the arm-mounted drone can detect mission completion in real time. Furthermore, these reward models utilize data from the arm-mounted drone's cameras and other sensors to determine mission success, providing a basis for subsequent autonomous data collection and policy optimization.

[0162] Environment reset uses a "policy pool" approach to automatically reset the environment. A policy pool is a set of policies with overlapping start and end states that can be reused to reset the environment, reducing manual intervention. This ensures that the drone can quickly switch between tasks and automatically reset the environment after completing a task, preparing for the next round of execution.

[0163] It should be noted that through fine-tuning, autonomous data collection, and model iteration, the UAV can continuously improve its capabilities, transitioning from a small number of demonstrations to high-level autonomous execution. This self-improvement process enables the UAV to continuously adapt to new missions and environmental requirements.

[0164] In some embodiments of the present invention, the PPO algorithm is a policy gradient method that can effectively process continuous action spaces and learn stochastic policies. Its core components include: an actor network (responsible for selecting the optimal action based on the current state), with the actor network's output layer using a Tanh activation function to generate continuous actions; and a critic network (responsible for evaluating the advantage of taking an action in the current state), with the critic network's output layer using a linear activation function to output a single value representing the state value. In other words, the PPO algorithm described above, based on a multi-objective reward function and a flexible obstacle avoidance function, constructs an initial motion model that also includes an actor network and a critic network.

[0165] During the training of the initial motion model, the PPO algorithm alternately updates the Actor network and the Critic network to obtain the final compliant collision avoidance motion model. The training process of the initial motion model first requires sampling a batch of trajectory data from the environmental data of the arm-carrying drone, which includes at least: the state data, action data, and reward data of the arm-carrying drone. Then, the improved policy gradient estimation method (Generalized Advantage Estimation, GAE) can be further used to estimate the advantage value of each state-action pair of the arm-carrying drone. Here, a special objective function LCLIP(θ) can be defined. This objective function LCLIP(θ) can limit the update amplitude of the Actor network by introducing a clipping mechanism to ensure that each training update link of the initial motion model is within the domain of the current Actor network.

[0166] In addition, the parameters of the Actor network can be further optimized using gradient descent to maximize the objective function LCLIP(θ). At the same time, the parameters of the Critic network can be optimized using mean squared error loss to predict more accurate state values.

[0167] In this way, during the training of the initial motion model, the above steps are repeated until the PPO algorithm converges. Here, the PPO algorithm maintains a balance between exploration and exploitation by alternating between updating the Actor network and the Critic network, and ultimately learns a compliant collision avoidance motion model corresponding to a stable and efficient strategy.

[0168] In some embodiments of the present invention, an attenuation (episode) mechanism can be introduced based on the PPO algorithm during the training process of the initial motion model. This attenuation mechanism can automatically set the current attenuation step number to the current step number when the training accuracy reaches the set value, thereby achieving adaptive adjustment in the trajectory planning process and ensuring the trajectory planning performance output by the smooth collision avoidance motion model.

[0169] In this way, a self-improvement method based on iterative learning using small sample data (training dataset) gradually optimizes and iteratively trains the initial motion model constructed using the PPO algorithm based on a multi-objective reward function and a flexible obstacle avoidance function, resulting in a compliant collision avoidance motion model. This enables the compliant collision avoidance motion model to possess stronger autonomous operation capabilities and environmental adaptability. This not only improves the execution efficiency and mission accuracy of the compliant collision avoidance motion model, but also significantly reduces its reliance on human intervention, opening up new possibilities for the subsequent application of arm-mounted drones based on the compliant collision avoidance motion model in complex and dynamic environments.

[0170] In some embodiments of the present invention, step 1043 may also be implemented by the following process:

[0171] First, the training data set is used to perform reinforcement learning training on the initial motion model to obtain an intermediate motion model.

[0172] In some embodiments of the present invention, for the description of obtaining the intermediate motion model, reference may be made to the above description of step 1043 , which will not be repeated here.

[0173] Then, the intermediate motion model is linearized to obtain the compliant collision avoidance motion model.

[0174] In some embodiments of the present invention, in order to ensure the stability of the robotic arm of the drone during obstacle avoidance, it is necessary to linearize the joint angle trajectory within the robotic arm to eliminate discontinuities in its corresponding motion trajectory. Here, the intermediate motion model that needs to output the relevant motion trajectory can be linearized to output a relatively stable and compliant collision avoidance motion model. The linearization process includes but is not limited to the following methods, and can also be a random fusion of the following methods:

[0175] Method 1: trajectory interpolation, which uses the interpolation method to smooth the joint angle trajectory of the robot arm output by the intermediate motion model to ensure the continuity of the motion trajectory of each joint in the robot arm on the time axis;

[0176] Method 2: Velocity and acceleration smoothing. This involves optimizing the changes in velocity and acceleration output by the intermediate motion model to reduce sharp changes in the motion trajectory of each joint within the robotic arm, ensuring smoothness and stability of the robotic arm's motion.

[0177] Method three, path optimization, is to introduce an optimization algorithm into the path planning output by the intermediate motion model to further smooth the changes in joint angles and reduce vibration and impact during movement.

[0178] In this way, the motion trajectory output by the obtained compliant collision avoidance motion model can be reduced in discontinuities, thereby improving the smoothness and stability of the subsequent motion trajectory of the robotic arm output by the compliant collision avoidance motion model.

[0179] Furthermore, the above process of “performing linearization processing on the intermediate motion model to obtain the compliant collision avoidance motion model” can also be implemented by the following steps:

[0180] In the first step, linearization is performed on the intermediate motion model to obtain the motion model to be adjusted.

[0181] Here, reference may also be made to the relevant description of step 1043 above, and the description of adjusting the motion model will not be repeated here.

[0182] The second step is to perform a performance evaluation on the motion model to be adjusted based on a simulation environment having a collision detection and risk assessment module to obtain an evaluation result.

[0183] In some embodiments of the present invention, a simulation environment with a flexible obstacle avoidance control strategy for a robotic arm can be designed based on the deep learning framework PyTorch (PyTorch is an open source deep learning framework for machine learning and deep learning) and the reinforcement learning experimental environment library (Gym). Here, the simulation environment can verify the joint simulation effect of the motion model to be adjusted based on the PPO algorithm and IVPF algorithm mentioned above by constructing a robotic arm kinematic model that is computationally efficient and can stably simulate various collision effects. At the same time, a collision detection and risk assessment module can be further designed in the simulation environment to evaluate the computational efficiency and simulation robustness of the motion model to be adjusted.

[0184] It should be noted that the following aspects may be considered during the construction of the simulation environment with collision detection and risk assessment modules:

[0185] 1. Environment Initialization: A Gym-based simulation environment is created to simulate the robot's motion and obstacle avoidance. This simulation environment includes the kinematic parameters of the six-axis robot arm and sets the constraints for each joint within the arm. Furthermore, the robot's kinematic model is further defined, along with joint constraints and physical parameters, to create a typical live-line working scenario involving wires, the robot arm, and other obstacles. Possible collision obstacles are also configured within the environment, along with the necessary sensors for collision detection.

[0186] Reward function design: To ensure the safety and efficiency of the obstacle avoidance strategy, a reward function can be defined and its associated weights can be set. The reward function primarily considers the following factors: collision detection (rewarding collision avoidance and penalizing collisions), path smoothness (rewarding smooth trajectories and penalizing drastic motion changes), and target reach (rewarding the robot arm for reaching the target point).

[0187] Regarding the implementation of the PPO algorithm, reinforcement learning training based on the PPO algorithm can be performed with relevant hyperparameters, including the learning rate, discount factor, and update frequency. The designed neural network structure can include an input layer, a hidden layer, and an output layer, where the input is the current state and action, and the output is the next action. The specific steps include constructing the neural network structure, setting the PPO algorithm hyperparameters, and implementing the training process, using a simulation environment for data collection and training.

[0188] The third step is to perform multiple rounds of iterative training and performance evaluation on the motion model to be adjusted based on the collected historical collision data of the robotic arm and the evaluation results, until the compliant collision avoidance motion model is obtained whose performance evaluation indicators meet the preset conditions.

[0189] In some embodiments of the present invention, historical collision data of the robotic arm is acquired through historical data collection. Simultaneously, based on this historical collision data and evaluation results, the motion model to be adjusted undergoes multiple rounds of iterative training and performance evaluation until a compliant collision avoidance motion model is obtained whose performance evaluation metrics meet preset conditions. Performance evaluation metrics include at least accuracy, precision, and recall, etc., which are not limited in the present invention.

[0190] It should be noted that for the training and optimization of the motion model to be adjusted, a corresponding deep reinforcement learning network can be constructed. The corresponding input training data is: the current and historical state characteristics of the collision, and the output is the avoidance action of the robot arm. Training is performed with the reward signal of delaying the occurrence of the collision. By using data from the collision database to train the deep reinforcement learning network, and adjusting the network structure and hyperparameters through iterative optimization, the model's detection and avoidance performance is improved. Closed-loop testing is performed in a simulation environment to generate collision situations and observe the detection and avoidance performance of the robot arm. Evaluation indicators such as detection success rate, avoidance reaction time, and trajectory deviation are collected, and the model is further improved based on the test results.

[0191] Furthermore, to implement the collision detection and risk assessment module, a module was designed that can detect the distance between the robot arm and obstacles in real time and assess the collision risk. This module uses distance calculation to calculate the distance between the robot arm end and the obstacle in real time. A collision risk assessment algorithm was also designed to assess the collision risk, taking into account different operational objectives and environmental factors.

[0192] To further improve the generalization performance of the collision avoidance algorithm, a collision monitoring device can be added during the actual operation of the robotic arm. Once a collision is detected, high-speed data acquisition is immediately triggered to obtain various state data at the moment of impact. This data includes multi-source sensor information such as visual images, joint positions, and collision forces. By analyzing the recorded collision data, a collision database is established, and augmented reality technology is used to generate a more diverse range of simulated collision cases.

[0193] During data collection and analysis, the simulation platform was used to collect a large amount of data on collisions encountered by the robotic arm during operation. This data includes the pre-collision state, the dynamic response during the collision, and the post-collision effects. The corresponding steps include collecting collision data (such as the robotic arm's joint position, speed, and collision force), annotating the data (indicating the time, location, and collision object of the collision), and establishing a collision database for classification and retrieval.

[0194] It's important to note that the avoidance algorithm's flexibility and robustness are enhanced through continuous iterative optimization. We built an online learning system framework to continuously collect data and optimize the algorithm during the robot's actual operation. We designed an intelligent collision cost function that fully considers multiple factors, including collision risk and operational impact, to enable the avoidance strategy to output more flexible and smooth avoidance trajectories. We expanded the network structure, added a memory module for historical avoidance results, and introduced a model prediction module to evaluate the long-term effectiveness of current avoidance decisions.

[0195] In some embodiments of the present invention, future work can further enhance the network generalization capabilities of the compliant collision avoidance motion model. Through continuous case learning and algorithm iterative optimization, the intelligence level of the compliant collision avoidance motion model can be continuously improved, enabling the robotic arm to have the ability to autonomously avoid and evade complex real-world environments. Here, actual experimental data shows that the improved compliant collision avoidance motion model significantly improves the robotic arm's adaptability to complex environments and the compliance and efficiency of collision avoidance. Thus, through continuous optimization and iteration, the computational efficiency and robustness of the simulation model can be further enhanced, providing an important reference for the future development of live-line working robot technology.

[0196] In this way, by co-simulating with the PyTorch and Gym simulation environments, a collision detection and risk assessment module was designed to improve the computational efficiency and robustness of the compliant collision avoidance motion model. This model can then output an efficient, safe, and intelligent robotic arm motion strategy suitable for high-precision operations in complex and dynamic environments.

[0197] Based on the above description, a 6-DOF robotic arm model corresponding to the robotic arm of a drone is first established using the MDH parameter method to define and transform the coordinate system and parameters. This ensures the consistency of positioning and motion of each joint within the robotic arm. FBG sensors are then used as force / torque sensors to monitor the forces and torques acting on the robotic arm by measuring changes in the optical fiber reflection wavelength, thereby ensuring safe and precise operation of the robotic arm. Secondly, technologies such as the ViT model are used for real-time detection and recognition of targets and the modeling of the drone's environment. Self-attention mechanisms and position encoding are used to improve the accuracy of target image processing and recognition. Furthermore, the PointNet++ multimodal neural network is further developed, combining vision, LiDAR, and force feedback data for high-precision environmental and target position perception, achieving precise target positioning. Finally, the PPO algorithm is used for robotic arm trajectory planning, and a multi-objective optimization reward function is designed to improve the accuracy, smoothness, and energy efficiency of trajectory planning. Flexible obstacle avoidance is achieved using the IVPF algorithm. Furthermore, through iterative learning with small samples, remote control demonstrations, and autonomous data collection, continuous optimization and environmental adaptation of the UAV with arms can be achieved, thereby improving the autonomous operation capabilities of the UAV with arms. A joint simulation is conducted within the deep learning framework PyTorch and the Gym simulation environment to design a collision detection and risk assessment module. The computational efficiency and simulation robustness of the corresponding compliant collision avoidance motion model are evaluated to optimize the obstacle avoidance strategy. This provides a highly comprehensive compliant collision avoidance motion model to ensure the high precision, safety, and intelligence of the UAV with arms that use the compliant collision avoidance motion model to output relevant motion trajectories in complex and dynamic environments, thereby improving the motion precision, safety, and intelligence of the UAV's robotic arm during movement.

[0198] The present invention provides a method for constructing a compliant collision avoidance motion model for a robotic arm. First, a motion trajectory plan corresponding to the motion of a robotic arm of an unmanned aerial vehicle (UAV) to a target object on a transmission line in a power transmission network is obtained. Then, a multi-objective reward function corresponding to the motion trajectory plan and a flexible obstacle avoidance function for obstacles within the motion area where the motion trajectory plan is located are constructed. Finally, a PPO algorithm is used, combined with the multi-objective reward function and the flexible obstacle avoidance function constructed based on an improved velocity potential field algorithm, to construct a compliant collision avoidance motion model for the robotic arm to the target object. In this way, a highly comprehensive compliant collision avoidance motion model can be provided, that is, the compliant collision avoidance motion model can output a motion trajectory with high motion accuracy, safety, and intelligence, thereby ensuring that the compliant collision avoidance motion model is used to output the relevant motion trajectory, and can guide the UAV with high precision, safety, and intelligence in complex and dynamic environments.

[0199] Correspondingly, such as Figure 3FIG. 1 is a flow chart of constructing a compliant collision avoidance motion model using the construction method provided in an embodiment of the present invention, and the corresponding execution steps are as follows:

[0200] 301. Construct a multi-DOF robotic arm kinematic model based on MDH. Here, a 6-DOF robotic arm kinematic model can be constructed. This 6-DOF robotic arm kinematic model can be used to establish a reference coordinate system for each joint in the robotic arm (determined by determining the origin of the link coordinate system, specifying the directions of the joint axes and the common perpendicular, etc.), and to describe the relationships between the links in the robotic arm using a transformation matrix.

[0201] 302. Strategies for force and torque sensing systems. This can be an FBG sensor installed at the end or joint of a robotic arm, which monitors the force and torque applied to the robotic arm by measuring changes in the reflected wavelength of the optical fiber.

[0202] 303. Visual perception and positioning of targets and / or obstacles. First, real-time detection and recognition of relevant objects can be achieved based on technologies such as the ViT model. Self-attention mechanisms and position encoding can be used to further improve image processing and recognition accuracy. Filtering and image enhancement techniques can be used to improve image quality. Next, the positioning of targets and / or obstacles is achieved based on the PointNet++ multimodal neural network, combining visual, LiDAR, and force feedback data for multimodal data feature extraction and fusion.

[0203] 304. A corresponding model is built for the flexible obstacle avoidance control strategy of a high-precision robotic arm based on the PPO algorithm. That is, the trajectory planning problem is transformed into a multi-objective optimization problem using the reinforcement learning framework. The relevant network is trained by designing the reward function and the PPO algorithm, and the IVPF algorithm is used to achieve flexible obstacle avoidance.

[0204] 305. The constructed model is subjected to iterative learning based on small samples to achieve a self-improvement method. Here, an initial training dataset can be collected through remote control demonstrations of armed drones and the model fine-tuned. Furthermore, the armed drones can be used to perform tasks in real environments and collect new task data, using the "post-hoc target re-labeling" technique to form a self-improvement dataset.

[0205] 306. Simulation Experiment. This involves co-simulating the deep learning framework PyTorch with the Gym simulation environment. A collision detection and risk assessment module is designed to evaluate the model's computational efficiency and simulation robustness. This includes further designing a reward function that takes into account factors such as collision detection, path smoothness, and target reach. A neural network structure is constructed to set hyperparameters for the PPO algorithm, enabling simulation and training.

[0206] in this way, Figure 3Specifically, a method for generating a compliant collision avoidance strategy (compliant collision avoidance motion model) for a robotic arm is provided. The corresponding compliant collision avoidance motion model includes creating joint nodes, collision avoidance detection nodes, collision avoidance strategy nodes, and intelligent operation planning nodes to construct a knowledge graph for compliant collision avoidance and intelligent operation. By acquiring a set of collision avoidance data sets and generating collision avoidance strategy information and intelligent operation planning information, technical support is provided for the subsequent autonomous collision avoidance and efficient operation of the robotic arm in complex environments, thereby improving the safety and operational efficiency of the robotic arm in complex environments.

[0207] Example 2

[0208] The present invention provides a control method for an unmanned aerial vehicle with an arm, see Figure 4 As shown, it is a flow chart of a control method of a drone with arms provided by the present invention; wherein, in combination with Figure 4 As shown, the control method of the drone with arms is described as follows:

[0209] Step 401: Acquire the current position of the current target object on the transmission line in the transmission network and the current environment information of the UAV with an arm.

[0210] Step 402: Input the current position and the current environment information into the compliant collision avoidance motion model constructed by the method described in any of the above embodiments to obtain the current motion trajectory.

[0211] Step 403: Based on the current motion trajectory, control the robotic arm of the UAV with an arm to move to the current position.

[0212] In some embodiments of the present invention, the current environmental information of the arm-carrying drone and the current position of the current target object can be directly input into the compliant collision avoidance motion model constructed in any of the above embodiments. Here, other information such as the forces and torques exerted on the arm-carrying drone during flight can be further input to obtain the current motion trajectory of the arm-carrying drone flying to the current position, so as to control and guide the robotic arm of the arm-carrying drone to move to the current position based on the current motion trajectory to grab the current target object, etc.

[0213] It should be noted that the description of the compliant collision avoidance motion model here can refer to the description of the above-mentioned method for constructing the compliant collision avoidance motion model of the robotic arm, and the present invention will not go into details therein.

[0214] In this way, the present invention provides a control method for an arm-mounted drone, which uses a highly comprehensive compliant collision avoidance motion model, that is, a compliant collision avoidance motion model that can improve the motion accuracy, safety and intelligence level of the arm-mounted drone, to output the corresponding current motion trajectory, so as to guide and improve the motion accuracy, safety and intelligence level of the robotic arm of the arm-mounted drone based on the current motion trajectory.

[0215] Example 3

[0216] The present invention based on the same inventive concept also provides a system 500 for constructing a compliant collision avoidance motion model of a robotic arm, see Figure 5 FIG. 1 is a schematic diagram of a system for constructing a compliant collision avoidance motion model of a robotic arm according to an embodiment of the present invention. The system includes:

[0217] The first acquisition module 501 is used to obtain a motion trajectory plan corresponding to the movement of the robotic arm of the UAV with an arm to the target object on the power transmission line in the power transmission network;

[0218] A first building module 502 is used to build a multi-objective reward function corresponding to the motion trajectory planning; wherein the multi-objective reward function at least includes: a motion accuracy reward function, a motion smoothness reward function, a motion energy consumption reward function, and an additional reward function;

[0219] The second building module 503 is used to build a flexible obstacle avoidance function for obstacles within the motion area where the motion trajectory planning is located using an improved velocity potential field algorithm;

[0220] The construction module 504 is used to use a proximal strategy optimization algorithm to construct a compliant collision avoidance motion model corresponding to the movement of the robotic arm to the target object based on the multi-objective reward function and the flexible obstacle avoidance function.

[0221] In some embodiments of the present invention, the first acquisition module 501 specifically acquires the position information of the target object and the environmental information of the arm-carrying drone; based on the forces and torques applied to the robotic arm during movement, the multi-degree-of-freedom robotic arm kinematic model corresponding to the robotic arm, the position information and the environmental information, determines the motion trajectory planning corresponding to the movement of the robotic arm to the target object.

[0222] In some embodiments of the present invention, the obstacle includes: a dynamic obstacle, a second building module 503, which is specifically used to use the improved velocity potential field algorithm to build a first function representing the direction between the end effector of the robotic arm and the target object, a second function representing the running speed of the dynamic obstacle, and a third function representing the tangential velocity between the end effector of the robotic arm and the target object; and the first function, the second function and the third function are integrated to obtain the flexible obstacle avoidance function.

[0223] In some embodiments of the present invention, the construction module 504 is specifically used to obtain a training data set; wherein, the training data set includes at least: takeoff data, hovering data, movement data, rotation data, grasping data and landing data corresponding to the execution of historical basic tasks by the arm-mounted drone; using the proximal strategy optimization algorithm, based on the multi-objective reward function and the flexible obstacle avoidance function, an initial motion model is constructed; using the training data set, the initial motion model is subjected to reinforcement learning training to obtain the compliant collision avoidance motion model.

[0224] In some embodiments of the present invention, the construction module 504 is specifically used to use the training data set to perform reinforcement learning training on the initial motion model to obtain an intermediate motion model; and linearize the intermediate motion model to obtain the compliant collision avoidance motion model.

[0225] In some embodiments of the present invention, the construction module 504 is specifically used to linearize the intermediate motion model to obtain the motion model to be adjusted; based on a simulation environment with a collision detection and risk assessment module, the performance of the motion model to be adjusted is evaluated to obtain an evaluation result; based on the collected historical collision data of the robotic arm and the evaluation result, the motion model to be adjusted is subjected to multiple rounds of iterative training and performance evaluation until the compliant collision avoidance motion model is obtained whose performance evaluation indicators meet the preset conditions.

[0226] It should be noted that the description of the corresponding embodiment of the system for constructing a compliant collision avoidance motion model for a robotic arm is similar to the description of the aforementioned method embodiment for constructing a compliant collision avoidance motion model for a robotic arm, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the system embodiment of the present invention, please refer to the description of the method embodiment of the present invention for understanding.

[0227] Example 4:

[0228] The present invention based on the same inventive concept also provides a control system 600 for an UAV with an arm, see Figure 6 FIG. 1 is a schematic diagram of a control system for a drone with an arm provided by an embodiment of the present invention, the system comprising:

[0229] The second acquisition module 601 is used to obtain the current position of the current target object on the transmission line in the transmission network and the current environment information of the UAV with an arm;

[0230] An input module 602 is configured to input the current position and the current environment information into the compliant collision avoidance motion model constructed by the system as described in the above embodiment to obtain a current motion trajectory;

[0231] The control module 603 is used to control the robotic arm of the arm-carrying drone to move to the current position based on the current motion trajectory.

[0232] It should be noted that the description of the corresponding embodiment of the control system of the armed drone is similar to the description of the control method embodiment of the armed drone described above, and has similar beneficial effects as the method embodiment. For any technical details not disclosed in the system embodiment of the present invention, please refer to the description of the method embodiment of the present invention for understanding.

[0233] Example 5

[0234] Based on the same inventive concept, Figure 7 As shown, the present invention also provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include at least one processor 701, a memory 702, a transceiver component 703, etc. The processor 701, the memory 702, and the transceiver component 703 are connected via a bus 704. The memory 702 may be used to store an execution program, which may include instructions. The processor 701 is used to execute the instructions stored in the memory. The memory 1002 may also be used to store data, which may be accessed and / or modified during the execution of the instructions.

[0235] The processor may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of the method for constructing the compliant collision avoidance motion model of the robotic arm described in any of the above embodiments, and the steps of the control method of the above-mentioned arm-carrying drone.

[0236] Example 6

[0237] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device readable storage medium (Memory), which is a memory device in an electronic device for storing programs and data. It can be understood that the storage medium here can include both built-in storage media in the electronic device and, of course, extended storage media supported by the electronic device. The storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space. These instructions can be one or more execution programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The processor loads and executes one or more instructions stored in the storage medium to implement the steps of the method for constructing the compliant collision avoidance motion model of the robotic arm described in any of the above embodiments, and the steps of the control method of the above-mentioned arm-carrying drone.

[0238] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0239] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0240] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0241] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0242] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A method for constructing a compliant collision avoidance motion model for a robotic arm, characterized in that: The method comprises: Obtaining the motion trajectory planning corresponding to the movement of the robotic arm of the UAV to the target object on the power transmission line in the power transmission network; Constructing a multi-objective reward function corresponding to the motion trajectory planning; wherein the multi-objective reward function includes at least: a motion accuracy reward function, a motion smoothness reward function, a motion energy consumption reward function, and an additional reward function; Using an improved velocity potential field algorithm, a flexible obstacle avoidance function is constructed for obstacles within the motion area where the motion trajectory planning is located; A proximal strategy optimization algorithm is used to construct a compliant collision avoidance motion model corresponding to the movement of the robotic arm to the target object based on the multi-objective reward function and the flexible obstacle avoidance function; The method adopts a proximal strategy optimization algorithm to construct a compliant collision avoidance motion model corresponding to the movement of the manipulator to the target object based on the multi-objective reward function and the flexible obstacle avoidance function, including: Acquire a training data set; wherein the training data set includes at least: takeoff data, hovering data, movement data, rotation data, grasping data and landing data corresponding to the execution of historical basic tasks by the UAV with an arm; Adopting the proximal strategy optimization algorithm, constructing an initial motion model based on the multi-objective reward function and the flexible obstacle avoidance function; The training data set is used to perform reinforcement learning training on the initial motion model to obtain the compliant collision avoidance motion model.

2. The method according to claim 1, characterized in that The obtaining of the motion trajectory planning corresponding to the movement of the robotic arm of the armed drone to the target object on the power transmission line in the power transmission network includes: Obtaining the location information of the target object and the environment information of the drone with an arm; Based on the forces and torques applied to the robotic arm during movement, the multi-degree-of-freedom robotic arm kinematic model corresponding to the robotic arm, the position information and the environmental information, the motion trajectory planning corresponding to the movement of the robotic arm to the target object is determined.

3. The method according to claim 1 or 2, characterized in that The obstacles include dynamic obstacles. The improved velocity potential field algorithm is used to build a flexible obstacle avoidance function for obstacles within the motion area where the motion trajectory planning is located, including: Using the improved velocity potential field algorithm, construct a first function representing the direction between the end effector of the robotic arm and the target object, a second function representing the running speed of the dynamic obstacle, and a third function representing the tangential velocity between the end effector of the robotic arm and the target object; The first function, the second function and the third function are integrated to obtain the flexible obstacle avoidance function.

4. The method according to claim 1, wherein The method of using the training data set to perform reinforcement learning training on the initial motion model to obtain the compliant collision avoidance motion model includes: Using the training data set, performing reinforcement learning training on the initial motion model to obtain an intermediate motion model; The intermediate motion model is linearized to obtain the compliant collision avoidance motion model.

5. The method according to claim 4, characterized in that The linearizing the intermediate motion model to obtain the compliant collision avoidance motion model includes: performing linearization processing on the intermediate motion model to obtain a motion model to be adjusted; Based on a simulation environment having a collision detection and risk assessment module, a performance evaluation is performed on the motion model to be adjusted to obtain an evaluation result; Based on the collected historical collision data of the robotic arm and the evaluation results, multiple rounds of iterative training and performance evaluation are performed on the motion model to be adjusted until the compliant collision avoidance motion model is obtained whose performance evaluation indicators meet preset conditions.

6. A system for constructing a compliant collision avoidance motion model for a robotic arm, characterized in that: The system comprises: The first acquisition module is used to obtain the motion trajectory planning corresponding to the movement of the robotic arm of the arm-carrying drone to the target object on the power transmission line in the power transmission network; A first building module is used to build a multi-objective reward function corresponding to the motion trajectory planning; wherein the multi-objective reward function includes at least: a motion accuracy reward function, a motion smoothness reward function, a motion energy consumption reward function, and an additional reward function; The second building module is used to build a flexible obstacle avoidance function for obstacles within the motion area where the motion trajectory planning is located using an improved velocity potential field algorithm; A construction module is used to construct a compliant collision avoidance motion model corresponding to the movement of the robotic arm to the target object based on the multi-objective reward function and the flexible obstacle avoidance function by adopting a proximal strategy optimization algorithm; Wherein, the building block is specifically used for: Acquire a training data set; wherein the training data set includes at least: takeoff data, hovering data, movement data, rotation data, grasping data and landing data corresponding to the execution of historical basic tasks by the UAV with an arm; Adopting the proximal strategy optimization algorithm, constructing an initial motion model based on the multi-objective reward function and the flexible obstacle avoidance function; The training data set is used to perform reinforcement learning training on the initial motion model to obtain the compliant collision avoidance motion model.

7. A control method for an arm-mounted drone, characterized in that: The method comprises: Obtain the current position of the current target on the transmission line in the power grid and the current environment information of the UAV with an arm; Inputting the current position and the current environment information into the compliant collision avoidance motion model constructed by the method according to any one of claims 1 to 5 to obtain a current motion trajectory; Based on the current motion trajectory, the robotic arm of the arm-carrying drone is controlled to move to the current position.

8. A control system for an unmanned aerial vehicle with an arm, characterized in that: The system comprises: The second acquisition module is used to obtain the current position of the current target object on the transmission line in the transmission network and the current environment information of the UAV with an arm; an input module, configured to input the current position and the current environment information into the compliant collision avoidance motion model constructed by the system according to claim 6 to obtain a current motion trajectory; A control module is used to control the mechanical arm of the arm-carrying drone to move to the current position based on the current motion trajectory.

9. A readable storage medium, characterized in that: An execution program is stored thereon, and when the execution program is executed, it implements the method for constructing a compliant collision avoidance motion model of a robotic arm as described in any one of claims 1 to 5, and the control method of the arm-equipped drone as described in claim 7.

Citation Information

Patent Citations

  • Barrier-avoiding method for mobile robot based on moving estimation of barrier

    CN101359229A

  • Space manipulator reinforcement learning motion planning method for unfixed obstacles

    CN116619380A