An intelligent control method and system for an operation and maintenance robot based on visual recognition

CN120680525BActive Publication Date: 2026-08-18AOWEI TECH (NANJING) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511060038.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2026-08-18
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

[0005]鉴于现有的视觉的机械手控制方法普遍存在目标识别精度不高、空间位姿计算不准确、动作控制与目标状态匹配性差以及缺乏自适应调节能力的问题,提出了本发明

Benefits of technology

[0053]The beneficial effects of this invention are as follows: By collecting RGB and depth image data from the operation and maintenance area and obtaining standardized image matrices and mapping relationship matrices, standardized processing of multimodal visual information and accurate establishment of spatial mapping relationships are achieved, effectively solving the technical problems of incomplete single RGB image information, missing depth information, and inconsistent processing of image data at different resolutions in traditional methods; by inputting the standardized image matrix into an improved ResNet residual network model to generate a comprehensive feature descriptor, and calculating the spatial position coordinates and attitude angles of the target device based on the mapping relationship matrix, the organic combination of deep learning feature extraction and geometric spatial positioning is achieved, breaking through the technical bottlenecks of limited feature extraction capabilities and insufficient spatial positioning accuracy of traditional visual recognition methods; by combining the current joint angle state of the robotic arm and utilizing the improved Jacobian moment... The inverse kinematics algorithm solves the target angle sequence of each joint and determines the grasping force and motion speed parameters based on the comprehensive feature descriptor matching preset operation mode library. This realizes an intelligent mapping transformation from visual perception to motion control, effectively solving the singularity problem and inaccurate parameter matching problem in the traditional inverse kinematics solution process. By converting the target angle sequence into control commands and sending them to the actuators of each joint of the manipulator, the manipulator is driven to complete motion planning, realizing precise execution from high-level motion planning to low-level servo control. Overall, this invention realizes full-process automation from environmental perception to task execution, effectively solving key technical problems such as strong reliance on manual labor, low operation accuracy, and high safety risks in traditional operation and maintenance operations. It achieves significant technical effects such as greatly improving operation and maintenance efficiency, reducing operation costs, and enhancing operation safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120680525B_ABST
    Figure CN120680525B_ABST
Patent Text Reader

Abstract

The application discloses an operation and maintenance mechanical hand intelligent control method and system based on visual identification, relates to the technical field of intelligent mechanical hand control, and comprises the following steps: collecting RGB image data and depth image data of an operation and maintenance operation area, obtaining a standardized image matrix and a mapping relationship matrix; inputting the standardized image matrix into an improved ResNet residual network model, generating a comprehensive feature descriptor, and calculating the spatial position coordinates and the attitude angle of a target device based on the mapping relationship matrix; based on the current joint angle state of the mechanical hand, solving the target angle sequence of each joint by using an improved Jacobian matrix inverse kinematics algorithm, and determining the grabbing strength parameter and the motion speed parameter based on the comprehensive feature descriptor matching a preset operation mode library; converting the target angle sequence into a control instruction, and sending the control instruction to the joint drivers of the mechanical hand to drive the mechanical hand to complete action planning. The application realizes the full-process automation from environment perception to task execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent robotic arm control technology, and in particular to an intelligent control method and system for maintenance robotic arms based on visual recognition. Background Technology

[0002] With the continuous improvement of industrial automation, robotic arm technology is widely used in manufacturing, inspection, assembly, and maintenance. Traditional robotic arms rely on preset paths or remote manual control to complete tasks, making it difficult to cope with the actual needs of equipment in dynamic environments with uncertain positions and complex posture changes. In recent years, the rapid development of computer vision technology has provided key support for robotic arms to perceive the environment, identify targets, and perform operations, making vision-based robotic arm control a research hotspot. Especially after integrating deep learning and inverse kinematics optimization algorithms, maintenance robotic arms have achieved stronger target recognition and operation accuracy. However, in practical applications, problems still exist such as inaccurate mapping between image information and spatial pose, untimely operation control response, and poor matching between robotic arm motion planning and actual target state, which restrict the stability and efficiency of intelligent maintenance operations.

[0003] CN116038670B discloses a vision-based robotic arm system. Through innovative gripping structure design, it achieves high gripping force output of the gripping block driven by a lever structure, enhancing the structural stability and durability of the gripping components. However, this technology mainly focuses on the mechanical optimization of the end effector gripping mechanism, without addressing how to perform accurate pose perception and path planning based on target recognition results. It lacks a closed-loop feedback mechanism from image perception to control execution, and its adaptability and automation level in dynamic working environments remain limited.

[0004] CN104827483A provides a mobile robotic gripper method combining GPS and binocular vision positioning. It improves the range and accuracy of target positioning through global positioning and local visual precision alignment. However, this solution relies on GPS signals in fixed scenarios and requires the target to possess good visual characteristics, making it unsuitable for complex structures, target occlusion, or space-constrained maintenance environments. Furthermore, this solution does not incorporate deep learning networks for multi-level abstraction of target features and lacks intelligent matching capabilities between target attributes and operational modes. Summary of the Invention

[0005] In view of the common problems of existing vision-based robotic arm control methods, such as low target recognition accuracy, inaccurate spatial pose calculation, poor matching between motion control and target state, and lack of adaptive adjustment capability, this invention is proposed.

[0006] Therefore, the problem to be solved by this invention is how to achieve high-precision identification and spatial perception of maintenance targets by a robotic arm with the support of multi-source visual data, and intelligently generate control commands based on feature matching and kinematic models to achieve efficient, safe and stable grasping and operation of target equipment, thereby improving the autonomy and operational capabilities of the robotic arm in actual maintenance tasks.

[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0008] In a first aspect, embodiments of the present invention provide an intelligent control method for a maintenance robot based on visual recognition, comprising,

[0009] Collect RGB and depth image data of the operation and maintenance area to obtain a standardized image matrix and mapping matrix;

[0010] The standardized image matrix is ​​input into the improved ResNet residual network model to generate a comprehensive feature descriptor, and the spatial position coordinates and attitude angle of the target device are calculated based on the mapping relationship matrix.

[0011] Based on the current joint angle state of the robotic arm, the target angle sequence of each joint is solved using an improved Jacobian matrix inverse kinematics algorithm, and the grasping force parameter and motion speed parameter are determined by matching the comprehensive feature descriptor with a preset operation mode library.

[0012] The target angle sequence is converted into control commands and sent to the actuators of each joint of the robot arm to drive the robot arm to complete the motion plan.

[0013] As a preferred embodiment of the intelligent control method for the maintenance robot based on visual recognition described in this invention, the method includes: converting the target angle sequence into control commands and sending them to the actuators of each joint of the robot to drive the robot to complete the motion planning, including:

[0014] The target angle sequence is divided into discrete joint control commands according to the time step. Each time step corresponds to a set of joint target angle values, thus generating a joint control command sequence.

[0015] Based on the servo control protocol of each joint of the robot, the joint control command sequence is encoded into the control signal of the driver and sent to the driver of each joint of the robot through the real-time communication interface.

[0016] After receiving the control signal, each joint actuator drives the motor to adjust the joint position according to the target angle sequence and feeds back the actual joint angle in real time through the joint encoder.

[0017] The depth sensor continuously collects real-time distance data between the end effector and the target device to calculate the current relative distance.

[0018] The current relative distance is compared with a preset threshold. If the current relative distance is greater than or equal to the preset threshold, the joint control command sequence continues to be executed.

[0019] If the current relative distance is less than a preset threshold, the grasping preparation state is triggered, the motion mode of the end effector is adjusted, and it switches to the precise pose adjustment stage. The preload value of the clamping force of the end effector is adjusted based on the grasping force parameter, and the approach speed of the end effector is limited based on the motion speed parameter. If the end effector makes contact with the target device, the final clamping force is applied according to the grasping force parameter, and the preset operation and maintenance process is executed according to the motion speed parameter until the task is completed.

[0020] As a preferred embodiment of the vision recognition-based intelligent control method for maintenance robots described in this invention, the method for determining the gripping force parameter and the motion speed parameter is as follows:

[0021] Read the current angle sensor data of each joint of the robot arm, obtain the current joint angle state, and generate the current joint angle vector;

[0022] Based on spatial position coordinates and attitude angle parameters, construct the target pose matrix of the target device in the manipulator base coordinate system;

[0023] A forward kinematic transfer function for the manipulator is established. Based on the link length parameters, joint offset angle parameters, and joint limit parameters of each joint of the manipulator, a coordinate transformation chain from the base to the end effector of the manipulator is constructed using the DH parameter method, and a forward kinematic transformation matrix is ​​generated.

[0024] An improved Jacobian matrix is ​​constructed, and based on the target pose matrix and the current joint angle vector, the pose deviation vector between the current pose of the end effector and the target pose is calculated.

[0025] The damped least squares method is used to invert the improved Jacobian matrix, a damping factor is introduced, and the pseudo-inverse matrix of the improved Jacobian matrix is ​​calculated by SVD decomposition to generate a stable Jacobian inverse matrix.

[0026] The pose deviation vector is multiplied by the stable Jacobian inverse matrix to calculate the angle increment that each joint needs to adjust. Combined with joint velocity and acceleration constraints, a smooth transition sequence from the current joint angle to the target joint angle is generated through a trajectory planning algorithm to form the target angle sequence.

[0027] The cosine similarity between the comprehensive feature descriptor and the feature descriptor template of the preset job mode library is calculated to match the optimal job mode and obtain the grasping force parameter and movement speed parameter.

[0028] As a preferred embodiment of the intelligent control method for maintenance robotic arms based on visual recognition described in this invention, wherein: a comprehensive feature descriptor is generated through the improved ResNet residual network model, including:

[0029] The normalized image matrix is ​​loaded as input data into the input layer of the improved ResNet residual network model;

[0030] The standardized image matrix is ​​preliminarily convolved by the shallow feature extraction module, and local detail features of the image are extracted using a small-sized convolution kernel to generate a shallow feature map.

[0031] The shallow feature map is input into a dual-branch feature extraction structure to generate edge feature vectors and texture feature vectors.

[0032] An adaptive feature fusion mechanism is designed to dynamically calculate the optimal fusion weight ratio of the edge feature vector and the texture feature vector based on the material properties and surface complexity of the target device.

[0033] Based on the optimal fusion weight ratio, the edge feature vector and texture feature vector are fused using a linear weighted combination method to generate a comprehensive feature descriptor.

[0034] As a preferred embodiment of the intelligent control method for operation and maintenance robotic arms based on visual recognition described in this invention, the improved ResNet residual network model is based on the traditional ResNet-50 architecture, with the addition of a parallel dual-branch feature extraction structure; the parallel dual-branch feature extraction structure includes an edge detection branch network and a texture analysis branch network; the improved ResNet residual network model includes a shallow feature extraction module, a deep semantic analysis module, and a residual connection module; the edge detection branch network integrates the Obel edge operator and the Laplacian edge operator; the texture analysis branch network is based on the local binary pattern algorithm and the gray-level co-occurrence matrix algorithm.

[0035] As a preferred embodiment of the vision recognition-based intelligent control method for maintenance robots described in this invention, the method for obtaining the spatial position coordinates and attitude angle parameters is as follows:

[0036] A feature template library for the target device is constructed. The comprehensive feature descriptor is matched with the pre-stored device feature templates to calculate similarity, determine the type identifier and geometric parameters of the target device, and generate device identification result data.

[0037] Based on the geometric parameters of the device identification result data, the key feature points of the target device are located in the standardized image matrix, the pixel coordinate values ​​of the key feature points are extracted, and a set of feature point coordinates is generated.

[0038] The mapping relationship matrix is ​​invoked to perform a three-dimensional coordinate transformation on the set of feature point coordinates, converting the two-dimensional pixel coordinates into three-dimensional spatial coordinates in the world coordinate system, thereby generating the spatial position coordinates of the target device.

[0039] Based on the spatial positional relationship of several feature points in the set of feature point coordinates, the principal direction vector and normal vector of the target device are calculated, and the spatial attitude of the target device relative to the base coordinate system of the manipulator is determined by the vector angle calculation formula, thus generating attitude angle parameters.

[0040] As a preferred embodiment of the intelligent control method for the maintenance robot based on visual recognition described in this invention, the method for obtaining the standardized image matrix and the mapping relationship matrix is ​​as follows:

[0041] The RGB image data is subjected to noise filtering. A Gaussian filter is used to eliminate random noise and uneven illumination interference in the RGB image data to obtain filtered RGB image data.

[0042] The filtered RGB image data is converted into grayscale image data, and the grayscale value of each pixel is calculated using a weighted average algorithm.

[0043] Histogram equalization is performed on the grayscale image data to redistribute the pixel grayscale value distribution range of the grayscale image data, enhance image contrast, and generate contrast-enhanced image data.

[0044] Based on preset image size parameters, the contrast-enhanced image data is subjected to size normalization processing to adjust image data of different resolutions to standard size specifications and generate a standardized image matrix.

[0045] Analyze the depth value information of each pixel in the depth image data, and calculate the actual distance data from the binocular vision sensor to the surface of each object in the operation and maintenance area through the triangulation principle;

[0046] Establish the mathematical transformation relationship between the pixel coordinate system and the world coordinate system. Based on the intrinsic and extrinsic parameter matrices of the binocular vision sensor and the actual distance data, calculate the transformation parameters between the pixel coordinates of each pixel and its corresponding actual spatial coordinates, and generate a mapping relationship matrix.

[0047] Secondly, embodiments of the present invention provide an intelligent control system for a maintenance robot based on visual recognition, which includes: an image acquisition and mapping construction module, used to acquire RGB image data and depth image data of the maintenance work area, and obtain a standardized image matrix and a mapping relationship matrix;

[0048] The feature extraction and spatial localization module is used to input the standardized image matrix into the improved ResNet residual network model, generate a comprehensive feature descriptor, and calculate the spatial position coordinates and attitude angle of the target device based on the mapping relationship matrix.

[0049] The joint control parameter calculation module, based on the current joint angle state of the robot, uses an improved Jacobian matrix inverse kinematics algorithm to solve the target angle sequence of each joint, and determines the gripping force parameter and motion speed parameter based on the comprehensive feature descriptor matching the preset operation mode library.

[0050] The execution control and operation feedback module is used to convert the target angle sequence into control commands and send them to the actuators of each joint of the robot to drive the robot to complete the motion plan.

[0051] Thirdly, embodiments of the present invention provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the computer program instructions, when executed by the processor, implement the steps of the intelligent control method for a maintenance robot based on visual recognition as described in the first aspect of the present invention.

[0052] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, they implement the steps of the intelligent control method for a vision-based maintenance robot as described in the first aspect of the present invention.

[0053] The beneficial effects of this invention are as follows: By collecting RGB and depth image data from the operation and maintenance area and obtaining standardized image matrices and mapping relationship matrices, standardized processing of multimodal visual information and accurate establishment of spatial mapping relationships are achieved, effectively solving the technical problems of incomplete single RGB image information, missing depth information, and inconsistent processing of image data at different resolutions in traditional methods; by inputting the standardized image matrix into an improved ResNet residual network model to generate a comprehensive feature descriptor, and calculating the spatial position coordinates and attitude angles of the target device based on the mapping relationship matrix, the organic combination of deep learning feature extraction and geometric spatial positioning is achieved, breaking through the technical bottlenecks of limited feature extraction capabilities and insufficient spatial positioning accuracy of traditional visual recognition methods; by combining the current joint angle state of the robotic arm and utilizing the improved Jacobian moment... The inverse kinematics algorithm solves the target angle sequence of each joint and determines the grasping force and motion speed parameters based on the comprehensive feature descriptor matching preset operation mode library. This realizes an intelligent mapping transformation from visual perception to motion control, effectively solving the singularity problem and inaccurate parameter matching problem in the traditional inverse kinematics solution process. By converting the target angle sequence into control commands and sending them to the actuators of each joint of the manipulator, the manipulator is driven to complete motion planning, realizing precise execution from high-level motion planning to low-level servo control. Overall, this invention realizes full-process automation from environmental perception to task execution, effectively solving key technical problems such as strong reliance on manual labor, low operation accuracy, and high safety risks in traditional operation and maintenance operations. It achieves significant technical effects such as greatly improving operation and maintenance efficiency, reducing operation costs, and enhancing operation safety. Attached Figure Description

[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0055] Figure 1 This is a flowchart of a vision recognition-based intelligent control method for maintenance robots. Detailed Implementation

[0056] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0057] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0058] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0059] As mentioned in the background section, existing vision-based robotic arm control technologies have achieved certain results in structural optimization and basic positioning, but there is still a bottleneck in the lack of efficient mapping between target recognition results and action execution. Especially in multi-task operation and maintenance scenarios, the robotic arm lacks a unified coordination mechanism for parameters such as the operation path, posture adjustment, and force control of the target device, resulting in high grasping errors and low execution efficiency. To address the above problems, a vision-based intelligent control method and system for operation and maintenance robotic arms is proposed.

[0060] Figure 1 This is a flowchart illustrating an intelligent control method for a maintenance robot based on vision recognition, according to an embodiment of the present invention. Figure 1 As shown, a vision-based intelligent control method for maintenance robots includes:

[0061] S1: Collect RGB image data and depth image data of the operation and maintenance work area, and obtain the standardized image matrix and mapping relationship matrix;

[0062] S2: Input the standardized image matrix into the improved ResNet residual network model to generate a comprehensive feature descriptor, and calculate the spatial position coordinates and attitude angle of the target device based on the mapping relationship matrix;

[0063] S3: Based on the current joint angle state of the robotic arm, the target angle sequence of each joint is solved using an improved Jacobian matrix inverse kinematics algorithm, and the grasping force parameters and motion speed parameters are determined by matching the preset operation mode library with the comprehensive feature descriptor.

[0064] S4: Convert the target angle sequence into control commands and send them to the actuators of each joint of the robot arm to drive the robot arm to complete the motion plan.

[0065] In this embodiment of the application, step S1 includes:

[0066] S1.1: Simultaneously acquire RGB image data of the operation and maintenance work area through the left and right image sensors, where the RGB image data includes the pixel values ​​of the red channel, green channel, and blue channel.

[0067] S1.2: Simultaneously activate the structured light projection module to project an coded grating pattern onto the maintenance work area, and receive the reflected light signal through the depth sensor to generate the corresponding depth image data.

[0068] S1.3: Perform noise filtering on the RGB image data. Use a Gaussian filter to eliminate random noise and uneven lighting interference in the RGB image data to obtain filtered RGB image data.

[0069] S1.4: Convert the filtered RGB image data into grayscale image data, and calculate the grayscale value of each pixel using a weighted average algorithm.

[0070] It should be noted that the filtered RGB image data retains the edge features and texture details of the original image; the weighted average algorithm sets different weight coefficients for the red, green, and blue channel pixel values ​​according to the characteristics of human vision.

[0071] S1.5: Perform histogram equalization on the grayscale image data to redistribute the pixel grayscale value distribution range of the grayscale image data, enhance image contrast, and generate contrast-enhanced image data.

[0072] S1.6: Based on preset image size parameters, the contrast-enhanced image data is normalized to adjust the image data of different resolutions to the standard size specifications and generate a standardized image matrix.

[0073] The preferred formula for the standardized image matrix is ​​as follows:

[0074]

[0075] Where, N m×n Let be an m×n standardized image matrix, Ψ(*) be the hyperbolic tangent normalization function, σ be the Gaussian kernel standard deviation, μ be the local mean of the image, and x be the normalized image matrix. k Let K be the pixel neighborhood coordinates, K be the total number of neighboring pixels, ⊙ be the Hadamard product, Φ(*) be the Sigmoid contrast enhancement function, and w be the pixel neighborhood coordinates. c F is the channel weight coefficient. c The data is RGB three-channel filtered data, H(*) is the Hessian matrix edge detection function, and I is the original RGB image.

[0076] It should be noted that contrast-enhanced image data has a more pronounced distinction between light and dark areas; each element of the normalized image matrix corresponds to the pixel grayscale value at a specific location within the maintenance work area; w r =0.299: The contribution weight of the red channel, based on the low sensitivity of the human eye to long wavelengths (700nm); w g =0.587: The green channel has the largest contribution weight because the human eye is most sensitive to mid-wavelength (550nm); w b =0.114: The blue channel has the smallest contribution weight because the human eye is less sensitive to short wavelengths (450nm).

[0077] S1.7: Analyze the depth value information of each pixel in the depth image data, and calculate the actual distance data from the binocular vision sensor to the surface of each object in the operation and maintenance area through the triangulation principle.

[0078] S1.8: Establish the mathematical transformation relationship between the pixel coordinate system and the world coordinate system. Based on the intrinsic and extrinsic parameter matrices of the binocular vision sensor and combined with actual distance data, calculate the transformation parameters between the pixel coordinates of each pixel and its corresponding actual spatial coordinates, and generate a mapping relationship matrix.

[0079] Preferably, the specific formula for the mapping relationship matrix is ​​as follows:

[0080]

[0081] Among them, (f x ,f y (c) is the focal length. x ,c y R is the optical center coordinate, R is the rotation matrix (3×3) describing the camera pose, and t is the translation vector (3×1).

[0082] It should be noted that the actual distance data reflects the three-dimensional spatial geometry of the operation and maintenance area; to verify the accuracy of the mapping relationship matrix, the known actual spatial coordinates are converted into pixel coordinates through a back projection algorithm, and the positional deviation of the calculated results is compared with that of the corresponding pixel points in the standardized image matrix.

[0083] Furthermore, when the deviation value is within the preset error range, the validity of the mapping relationship matrix is ​​confirmed, and the mapping relationship between pixel coordinates and actual spatial coordinates is established.

[0084] For example, the robotic arm system uses Intel RealSense. The D435i binocular camera acquires RGB-D data. When the robotic arm approaches the target device, the left and right image sensors simultaneously capture RGB images with a resolution of 1280×720. At the same time, the structured light projection module emits an 850nm infrared coded grating, and the depth sensor receives the reflected signal and generates the corresponding depth image. Assuming that the RGB value of a certain pixel (320, 240) is (120, 45, 200) and the depth value is 1532mm, the system performs a 5×5 Gaussian filter (σ=1.5) on the RGB image to eliminate noise and converts it into a grayscale image according to the human visual weights (R:0.299, G:0.587, B:0.114). The grayscale value of the RGB value (100, 150, 80) is 134.27. After histogram equalization and size normalization, a 512×512 normalized image matrix is ​​generated, where each pixel value is normalized to the range [-1,1] by the hyperbolic tangent function for easy neural network processing.

[0085] In this embodiment of the application, step S2 includes:

[0086] In an optional implementation, the improved ResNet residual network model is based on the traditional ResNet-50 architecture, with the addition of a parallel dual-branch feature extraction structure. The parallel dual-branch feature extraction structure includes an edge detection branch network and a texture analysis branch network. The improved ResNet residual network model includes a shallow feature extraction module, a deep semantic analysis module, and a residual connection module, which maintains the stability of gradient propagation and prevents feature information loss through a skip connection structure. The edge detection branch network integrates the Obel edge operator and the Laplacian edge operator. The texture analysis branch network is based on the local binary pattern algorithm and the gray-level co-occurrence matrix algorithm. Key feature points include the device center point, corner vertices, and feature marker points. Spatial pose includes the pitch angle around the X-axis, the yaw angle around the Y-axis, and the roll angle around the Z-axis.

[0087] For example, the improved ResNet-50 model employs a dual-branch structure. The edge detection branch uses multi-scale dilated convolution (dilation rate 1 / 2 / 3) to extract the target contour, such as detecting pixels with an edge gradient magnitude exceeding 50 for the hexagonal head of a bolt. The texture analysis branch calculates surface texture using Local Binary Pattern (LBP), such as the regular binary encoding pattern of the threads on a metal bolt (e.g., 10110011). The adaptive feature fusion module dynamically adjusts the weights according to the target material; for example, the edge weight for a metal object is set to 0.7, and the texture weight to 0.3. The fused comprehensive feature descriptor is matched with a preset template library to identify the target as an M10 hexagonal bolt and locate its center point pixel coordinates (400, 300). Combining camera intrinsics (e.g., focal length fx = 900) and the mapping matrix, the pixel coordinates are converted into a three-dimensional position in the world coordinate system (X = 0.5m, Y = 0.2m, Z = 1.5m), and the bolt attitude (pitch angle 10°, yaw angle -5°) is calculated using the spatial vectors of the three feature points.

[0088] S2.1: Load the normalized image matrix as input data into the input layer of the improved ResNet residual network model.

[0089] S2.2: The standardized image matrix is ​​initially convolved by the shallow feature extraction module, and local detail features of the image are extracted using small-sized convolution kernels to generate a shallow feature map.

[0090] S2.3: Input the shallow feature map into the dual-branch feature extraction structure to generate edge feature vectors and texture feature vectors.

[0091] Preferably, the shallow feature map is input into the edge detection branch network, and a multi-scale dilated convolutional layer is used to enhance the detailed features of the target contour. The gradient change rate of adjacent pixels is calculated, gray-scale abrupt change regions are identified, and the boundary contour pixels of the target device are extracted to generate an edge feature vector. At the same time, the shallow feature map is input into the texture analysis branch network, and a local binary mode enhancement convolutional layer is used to strengthen the surface texture features. In combination with an attention mechanism, key texture regions are selected, the texture statistical characteristic parameters of key texture regions are calculated, the material texture pattern of the target device surface is extracted, and a texture feature vector is generated.

[0092] It should be noted that the shallow feature map preserves the fine texture information and edge contour information of the target device surface; the texture statistical characteristic parameters include texture contrast, texture uniformity, texture entropy value and texture correlation.

[0093] S2.4: Design an adaptive feature fusion mechanism that dynamically calculates the optimal fusion weight ratio between the simplified edge feature vector and the standardized texture feature vector based on the material properties and surface complexity of the target device.

[0094] S2.5: Based on the optimal fusion weight ratio, the edge feature vector and texture feature vector are fused through a linear weighted combination method to generate a comprehensive feature descriptor.

[0095] Preferably, the comprehensive feature descriptor comprehensively represents the shape and surface features of the target device.

[0096] S2.6: Construct a feature template library for the target device, perform similarity matching calculations between the comprehensive feature descriptor and the pre-stored device feature templates, determine the type identifier and geometric parameters of the target device, and generate device identification result data.

[0097] S2.7: Based on the geometric parameters of the device recognition result data, locate the key feature points of the target device in the standardized image matrix, extract the pixel coordinate values ​​of the key feature points, and generate a set of feature point coordinates;

[0098] S2.8: Call the mapping relationship matrix to perform a three-dimensional coordinate transformation on the feature point coordinate set, convert the two-dimensional pixel coordinates into three-dimensional spatial coordinates in the world coordinate system, and generate the spatial position coordinates of the target device;

[0099] S2.9: Based on the spatial positional relationship of several feature points in the feature point coordinate set, calculate the principal direction vector and normal vector of the target device, determine the spatial attitude of the target device relative to the robot's base coordinate system through the vector angle calculation formula, and generate attitude angle parameters.

[0100] It should be noted that the equipment identification result data includes equipment category information and size parameter information; the feature point coordinate set describes the spatial distribution of the target equipment in the two-dimensional image plane; the spatial position coordinates determine the accurate location of the target equipment in the operation and maintenance space; and the attitude angle parameter describes the spatial orientation state of the target equipment.

[0101] In this embodiment of the application, step S3 includes:

[0102] S3.1: Read the current angle sensor data of each joint of the robot arm, obtain the current joint angle status, and generate the current joint angle vector.

[0103] It should be noted that the current joint angle status includes the real-time angle values ​​of the robot's base joint, shoulder joint, elbow joint, wrist joint, and end effector joint.

[0104] S3.2: Based on spatial position coordinates and attitude angle parameters, construct the target pose matrix of the target device in the robot arm base coordinate system.

[0105] It should be noted that the target pose matrix includes the spatial position component and the rotational attitude component of the target device.

[0106] S3.3: Establish the forward kinematic transfer function of the manipulator. Based on the link length parameters, joint offset angle parameters and joint limit parameters of each joint of the manipulator, construct the coordinate transformation chain from the base to the end effector of the manipulator through the DH parameter method, and generate the forward kinematic transformation matrix.

[0107] In an optional implementation, constructing the improved Jacobian matrix includes: introducing a singularity avoidance mechanism and a joint constraint mechanism on the basis of the Jacobian matrix; calculating the influence coefficient of each joint angle change on the pose change of the end effector by performing partial differential operations on the positive kinematic transformation matrix; and generating Jacobian matrix elements.

[0108] S3.4: Construct an improved Jacobian matrix, and calculate the pose deviation vector between the current pose of the end effector and the target pose based on the target pose matrix and the current joint angle vector.

[0109] S3.5: The damped least squares method is used to invert the improved Jacobian matrix, a damping factor is introduced, and the pseudo-inverse matrix of the improved Jacobian matrix is ​​calculated by SVD decomposition to generate a stable Jacobian inverse matrix.

[0110] It should be noted that the pose deviation vector includes the position deviation vector and the attitude deviation vector; a damping factor is introduced to prevent numerical instability of the Jacobian matrix near singular points.

[0111] The preferred formula for the stable Jacobian inverse matrix is ​​as follows:

[0112] J * =V(∑ 2 +λ 2 B) -1 ∑U T ;

[0113] Among them, J * To form a stable Jacobian inverse matrix, V and U are orthogonal matrices after SVD decomposition, and λ is the damping factor. Let be the Jacobian matrix of the robotic arm, Σ be the singular value diagonal matrix, and T be the permutation of the orthogonal matrix.

[0114] S3.6: Multiply the pose deviation vector with the stable Jacobian inverse matrix to calculate the angle increment that each joint needs to adjust. Combine the joint velocity limit and acceleration limit, and generate a smooth transition sequence from the current joint angle to the target joint angle through the trajectory planning algorithm to form the target angle sequence.

[0115] S3.7: Perform cosine similarity calculation between the comprehensive feature descriptor and the feature descriptor template of the preset job mode library, match the optimal job mode, and obtain the grasping force parameter and movement speed parameter.

[0116] Preferably, the preset operation mode library includes grasping parameter templates corresponding to different equipment types. The operation mode type that best matches the comprehensive feature descriptor is found through feature similarity calculation, and a matching operation mode identifier is generated. The corresponding grasping force parameter table is queried based on the matching operation mode identifier, and the grasping force parameters adapted to the target equipment are extracted. The grasping force parameter table pre-sets different grasping force values ​​according to the material characteristics, surface roughness, and structural strength of the target equipment. The corresponding motion speed parameter table is determined based on the matching operation mode identifier, and the motion speed parameters suitable for the current operation task are extracted. The motion speed parameter table takes into account factors such as the weight, shape complexity, and safety requirements of the target equipment, and includes speed limit values ​​for each motion stage.

[0117] For example, the current joint angles of the robotic arm are [20°, 30°, 45°, 10°, 0°]. The target pose matrix requires the end effector to reach (0.5m, 0.2m, 1.5m) and maintain a vertical grasping posture. An improved Jacobian matrix introduces a damping factor λ = 0.1 to avoid singularities when joint 2 approaches 90°. The pseudo-inverse matrix is ​​obtained through SVD decomposition, and the sequence of joint angle increments is calculated. The trajectory planning module generates a 5th-order polynomial interpolation curve, limiting the maximum joint acceleration to 30° / s². 2 To ensure smooth movement, the system matches the feature descriptor with the operation mode library to determine the gripping parameters for the metal bolt: gripping force 30N (to prevent slippage) and movement speed 0.5m / s (efficiency priority). If the target is identified as a ceramic insulator, it automatically switches to low force mode (15N) and low speed (0.2m / s).

[0118] In this embodiment of the application, step S4 includes:

[0119] S4.1: The target angle sequence is split into discrete joint control commands according to the time step. Each time step corresponds to a set of joint target angle values, generating a joint control command sequence.

[0120] It should be noted that the joint control command sequence includes the target angle command for the base joint, the target angle command for the shoulder joint, the target angle command for the elbow joint, the target angle command for the wrist joint, and the target angle command for the end effector joint.

[0121] S4.2: Based on the servo control protocol of each joint of the robot, the joint control command sequence is encoded into the control signal of the driver and sent to the driver of each joint of the robot through the real-time communication interface.

[0122] S4.3: After receiving the control signal, each joint actuator drives the motor to adjust the joint position according to the target angle sequence and provides real-time feedback of the actual joint angle through the joint encoder.

[0123] S4.4: Continuously collect real-time distance data between the end effector and the target device using a depth sensor, and calculate the current relative distance.

[0124] Preferably, the laser rangefinder installed on the end effector of the robot arm is activated to emit a laser pulse signal towards the target device. The straight-line distance between the end effector of the robot arm and the surface of the target device is calculated by measuring the round-trip time of the laser pulse, and the current relative distance is generated. A relative distance monitoring cycle is established to continuously acquire the current relative distance at a preset sampling frequency. The relative distance value obtained from each measurement is stored in the distance data buffer to form a continuous distance change sequence.

[0125] Specifically, the current relative distance is compared with a preset threshold. If the current relative distance is greater than or equal to the preset threshold, the joint control command sequence continues to be executed until the robot's end effector approaches the target device. If the current relative distance is less than the preset threshold, the gripping preparation state is triggered, the motion mode of the end effector is adjusted, and it switches to the precise pose adjustment stage. The preload value of the gripping force of the end effector is adjusted based on the gripping force parameter, and the approach speed of the end effector is limited based on the motion speed parameter. If the end effector makes contact with the target device, the final gripping force is applied according to the gripping force parameter, and the preset operation procedure is executed according to the motion speed parameter until the task is completed.

[0126] Preferably, after successfully grasping the target equipment, the robot arm is controlled to perform corresponding operations according to the specific requirements of the maintenance task, including lifting, rotating, moving or placing. The execution status of the maintenance task is monitored, and the position sensor and force sensor of the end effector are used to determine whether the task is completed. When the target equipment reaches the designated position and the contact force is stable, a task completion signal is generated.

[0127] For example, joint control commands are sent in 10ms intervals, such as gradually adjusting from an initial angle [20°, 30°, 45°, 10°, 0°] to a target angle [30°, 45°, 60°, 20°, 0°], with angle increments [1.5°, 2.3°, 3.0°, 1.0°, 0°] sent every 10ms. During the end effector's approach, a laser rangefinder monitors the distance in real time. If the distance drops sharply from 200mm to 50mm, it immediately triggers a deceleration to 10% speed. When the force sensor feedback pressure reaches 5N after contact, the movement stops, and a preset 30N clamping force is applied. If the actual angle of a joint deviates from the target value by more than 2° for 100ms, the system pauses the commands and replans the trajectory to ensure operational safety. After the gripping is completed, the robot moves the bolt to the designated position according to the task requirements, and the force sensor generates a work completion signal after confirming stable placement.

[0128] In summary, this invention achieves standardized processing of multimodal visual information and accurate establishment of spatial mapping relationships by collecting RGB and depth image data from the operation and maintenance area and obtaining standardized image matrices and mapping relationship matrices. This effectively solves the technical problems of incomplete single RGB image information, missing depth information, and inconsistent processing of image data at different resolutions. By inputting the standardized image matrix into an improved ResNet residual network model to generate a comprehensive feature descriptor, and calculating the spatial position coordinates and attitude angles of the target device based on the mapping relationship matrix, it achieves an organic combination of deep learning feature extraction and geometric spatial localization, overcoming the technical bottlenecks of limited feature extraction capabilities and insufficient spatial localization accuracy in traditional visual recognition methods. Furthermore, by combining the current joint angle state of the robotic arm and utilizing an improved Jacobian matrix… The inverse kinematics algorithm solves the target angle sequence of each joint and determines the grasping force and motion speed parameters based on a pre-defined operation mode library matched with a comprehensive feature descriptor. This achieves an intelligent mapping and transformation from visual perception to motion control, effectively solving the singularity problem and inaccurate parameter matching problem in the traditional inverse kinematics solution process. By converting the target angle sequence into control commands and sending them to the actuators of each joint of the manipulator, the manipulator is driven to complete motion planning, achieving precise execution from high-level motion planning to low-level servo control. Overall, this invention achieves full-process automation from environmental perception to task execution, effectively solving key technical problems in traditional operation and maintenance operations such as strong reliance on manual labor, low operational accuracy, and high safety risks. It achieves significant technical effects such as greatly improving operation and maintenance efficiency, reducing operation costs, and enhancing operation safety.

[0129] Furthermore, this embodiment also provides a vision-based intelligent control system for an operation and maintenance robot, including: an image acquisition and mapping construction module, used to acquire RGB image data and depth image data of the operation and maintenance work area, and obtain a standardized image matrix and a mapping relationship matrix; a feature extraction and spatial positioning module, used to input the standardized image matrix into an improved ResNet residual network model, generate a comprehensive feature descriptor, and calculate the spatial position coordinates and attitude angle of the target device based on the mapping relationship matrix; a joint control parameter calculation module, based on the current joint angle state of the robot, using an improved Jacobian matrix inverse kinematics algorithm to solve the target angle sequence of each joint, and based on the comprehensive feature descriptor matching a preset work mode library to determine the gripping force parameter and motion speed parameter; and an execution control and work feedback module, used to convert the target angle sequence into control commands and send them to the joint actuators of the robot to drive the robot to complete the motion planning.

[0130] This embodiment also provides a computer device applicable to the intelligent control method of maintenance robot based on vision recognition, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the intelligent control method of maintenance robot based on vision recognition as proposed in the above embodiment.

[0131] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0132] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for intelligent control of a maintenance robot based on visual recognition, characterized in that: include, Collect RGB and depth image data of the operation and maintenance area to obtain a standardized image matrix and mapping matrix; The standardized image matrix is ​​input into the improved ResNet residual network model to generate a comprehensive feature descriptor, and the spatial position coordinates and attitude angle of the target device are calculated based on the mapping relationship matrix. Based on the current joint angle state of the robotic arm, the target angle sequence of each joint is solved using an improved Jacobian matrix inverse kinematics algorithm, and the grasping force parameter and motion speed parameter are determined by matching the comprehensive feature descriptor with a preset operation mode library. The target angle sequence is converted into control commands and sent to the actuators of each joint of the robot arm to drive the robot arm to complete the motion plan; The target angle sequence is divided into discrete joint control commands according to the time step. Each time step corresponds to a set of joint target angle values, thus generating a joint control command sequence. Based on the servo control protocol of each joint of the robot, the joint control command sequence is encoded into the control signal of the driver and sent to the driver of each joint of the robot through the real-time communication interface. After receiving the control signal, each joint actuator drives the motor to adjust the joint position according to the target angle sequence and feeds back the actual joint angle in real time through the joint encoder. The depth sensor continuously collects real-time distance data between the end effector and the target device to calculate the current relative distance. The current relative distance is compared with a preset threshold. If the current relative distance is greater than or equal to the preset threshold, the joint control command sequence continues to be executed. If the current relative distance is less than a preset threshold, the grasping preparation state is triggered, the motion mode of the end effector is adjusted, and it is switched to the precise pose adjustment stage. The clamping force preload value of the end effector is adjusted based on the grasping force parameter, and the approach speed of the end effector is limited based on the motion speed parameter. If the end effector makes contact with the target device, the final clamping force is applied according to the grasping force parameter, and the preset operation and maintenance process is executed according to the motion speed parameter until the task is completed. The method for determining the grasping force parameter and the movement speed parameter is as follows: Read the current angle sensor data of each joint of the robot arm, obtain the current joint angle state, and generate the current joint angle vector; Based on spatial position coordinates and attitude angle parameters, construct the target pose matrix of the target device in the manipulator base coordinate system; A forward kinematic transfer function for the manipulator is established. Based on the link length parameters, joint offset angle parameters, and joint limit parameters of each joint of the manipulator, a coordinate transformation chain from the base to the end effector of the manipulator is constructed using the DH parameter method, and a forward kinematic transformation matrix is ​​generated. An improved Jacobian matrix is ​​constructed, and based on the target pose matrix and the current joint angle vector, the pose deviation vector between the current pose of the end effector and the target pose is calculated. The damped least squares method is used to invert the improved Jacobian matrix, a damping factor is introduced, and the pseudo-inverse matrix of the improved Jacobian matrix is ​​calculated by SVD decomposition to generate a stable Jacobian inverse matrix. The pose deviation vector is multiplied by the stable Jacobian inverse matrix to calculate the angle increment that each joint needs to adjust. Combined with joint velocity and acceleration constraints, a smooth transition sequence from the current joint angle to the target joint angle is generated through a trajectory planning algorithm to form the target angle sequence. The cosine similarity between the comprehensive feature descriptor and the feature descriptor template of the preset operation mode library is calculated to match the optimal operation mode and obtain the grasping force parameter and the movement speed parameter. Generate a comprehensive feature descriptor using the improved ResNet residual network model, including: The normalized image matrix is ​​loaded as input data into the input layer of the improved ResNet residual network model; The standardized image matrix is ​​preliminarily convolved by the shallow feature extraction module, and local detail features of the image are extracted using a small-sized convolution kernel to generate a shallow feature map. The shallow feature map is input into a dual-branch feature extraction structure to generate edge feature vectors and texture feature vectors. An adaptive feature fusion mechanism is designed to dynamically calculate the optimal fusion weight ratio of the edge feature vector and the texture feature vector based on the material properties and surface complexity of the target device. Based on the optimal fusion weight ratio, the edge feature vector and texture feature vector are fused by a linear weighted combination method to generate a comprehensive feature descriptor. The improved ResNet residual network model is based on the traditional ResNet-50 architecture, with the addition of a parallel dual-branch feature extraction structure. The parallel dual-branch feature extraction structure includes an edge detection branch network and a texture analysis branch network. The improved ResNet residual network model includes a shallow feature extraction module, a deep semantic analysis module, and a residual connection module. The edge detection branch network integrates the Obel edge operator and the Laplacian edge operator. The texture analysis branch network is based on the local binary pattern algorithm and the gray-level co-occurrence matrix algorithm.

2. The intelligent control method for maintenance robotic arms based on visual recognition as described in claim 1, characterized in that: The method for obtaining the spatial position coordinates and attitude angle parameters is as follows: A feature template library for the target device is constructed. The comprehensive feature descriptor is matched with the pre-stored device feature templates to calculate similarity, determine the type identifier and geometric parameters of the target device, and generate device identification result data. Based on the geometric parameters of the device identification result data, the key feature points of the target device are located in the standardized image matrix, the pixel coordinate values ​​of the key feature points are extracted, and a set of feature point coordinates is generated. The mapping relationship matrix is ​​invoked to perform a three-dimensional coordinate transformation on the set of feature point coordinates, converting the two-dimensional pixel coordinates into three-dimensional spatial coordinates in the world coordinate system, thereby generating the spatial position coordinates of the target device. Based on the spatial positional relationship of several feature points in the set of feature point coordinates, the principal direction vector and normal vector of the target device are calculated, and the spatial attitude of the target device relative to the base coordinate system of the manipulator is determined by the vector angle calculation formula, thus generating attitude angle parameters.

3. The intelligent control method for maintenance robotic arms based on visual recognition as described in claim 2, characterized in that: The method for obtaining the standardized image matrix and the mapping relationship matrix is ​​as follows: The RGB image data is subjected to noise filtering. A Gaussian filter is used to eliminate random noise and uneven lighting interference in the RGB image data to obtain filtered RGB image data. The filtered RGB image data is converted into grayscale image data, and the grayscale value of each pixel is calculated using a weighted average algorithm. Histogram equalization is performed on the grayscale image data to redistribute the pixel grayscale value distribution range of the grayscale image data, enhance image contrast, and generate contrast-enhanced image data. Based on preset image size parameters, the contrast-enhanced image data is subjected to size normalization processing to adjust image data of different resolutions to standard size specifications and generate a standardized image matrix. Analyze the depth value information of each pixel in the depth image data, and calculate the actual distance data from the binocular vision sensor to the surface of each object in the operation and maintenance area through the triangulation principle; Establish the mathematical transformation relationship between the pixel coordinate system and the world coordinate system. Based on the intrinsic and extrinsic parameter matrices of the binocular vision sensor and the actual distance data, calculate the transformation parameters between the pixel coordinates of each pixel and its corresponding actual spatial coordinates, and generate a mapping relationship matrix.

4. A vision-based intelligent control system for a maintenance robot, based on the vision-based intelligent control method for a maintenance robot according to any one of claims 1 to 3, characterized in that: include, The image acquisition and mapping module is used to acquire RGB image data and depth image data of the operation and maintenance work area, and obtain a standardized image matrix and mapping relationship matrix. The feature extraction and spatial localization module is used to input the standardized image matrix into the improved ResNet residual network model, generate a comprehensive feature descriptor, and calculate the spatial position coordinates and attitude angle of the target device based on the mapping relationship matrix. The joint control parameter calculation module, based on the current joint angle state of the robot, uses an improved Jacobian matrix inverse kinematics algorithm to solve the target angle sequence of each joint, and determines the gripping force parameter and motion speed parameter based on the comprehensive feature descriptor matching the preset operation mode library. The execution control and operation feedback module is used to convert the target angle sequence into control commands and send them to the actuators of each joint of the robot arm to drive the robot arm to complete the motion plan.

5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the intelligent control method for the maintenance robot based on vision recognition as described in any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the intelligent control method for the maintenance robot based on vision recognition as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Method for grabbing object through mobile manipulator on basis of GPS and binocular vision positioning

    CN104827483A

  • Recognizing and grabbing method of image of mechanical arm part based on Kinect sensor

    CN110480637A

  • Mechanical arm target grabbing method based on deep learning and edge detection

    CN114012722A

  • Manipulator flexible grabbing method capable of automatically switching workpieces

    CN118544365A

  • Mechanical arm grabbing intelligent optimization method and system based on reinforcement learning

    CN120363181A