A method and system for controlling an anti-interference robot arm based on visual servoing
By generating pseudo-abnormal images through random walk algorithm and Transformer network, combined with Gaussian blur and Haar wavelet response, the problems of local abnormal feature simulation and multi-scale interference reconstruction in the anti-interference control of visual servo robotic arm are solved, and the perception robustness and task execution stability of the robotic arm are improved.
Patent Information
- Application Number
- CN202511000147.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-21
AI Technical Summary
Existing anti-interference control methods for robotic arms based on visual servoing lack the simulation of local abnormal features of images, have limited generalization capabilities, and cannot meet the reconstruction requirements of multi-scale interference.
A random walk algorithm is used to generate pseudo-abnormal images, and the Transformer network structure and multi-head self-attention mechanism are combined for image restoration. Feature points are extracted through Gaussian blur map and Haar wavelet response. The least squares method is used to calculate joint angle errors. The control parameters are optimized by combining the PID control algorithm, and the data is stored in a database.
It significantly improves the perception robustness, task execution stability and generalization ability of the robotic arm in disturbed environments, and improves image recognition and control accuracy.
Smart Images

Figure CN120516719B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robotic arm anti-interference control, and in particular to a robotic arm anti-interference control method and system based on visual servoing. Background Art
[0002] With the rapid development of automation, intelligent manufacturing and precision control technology, industrial robots are playing an increasingly critical role in welding, assembly, spraying and handling applications. Among them, the robotic arm, as the core executive component of the industrial robot, has a direct impact on the production efficiency and safety of the overall system due to its control accuracy and robustness. Among the many control methods, the robotic arm control technology based on visual servoing has received extensive research and attention due to its advantages of not requiring high-precision position sensors and having adaptability and flexibility. This technology obtains image information through a camera system and adjusts the robotic arm movement based on image feature feedback, thereby achieving high-precision control.
[0003] However, the existing visual servo-based anti-interference control method for robotic arms lacks the simulation of local abnormal features in the image, resulting in reduced ability to identify and recover potential abnormal structures. It also fails to fully integrate global context and local disturbance features, resulting in limited generalization ability. Secondly, the anti-interference mechanism is mostly limited to image-level filtering or masking, lacks pseudo-anomaly guidance and structure repair mechanisms, and cannot meet the reconstruction needs of multi-scale interference. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method and system for anti-interference control of a robotic arm based on visual servoing, which solves the problems that the existing anti-interference control method of a robotic arm based on visual servoing lacks simulation of local abnormal features of the image, has limited generalization ability, and cannot cope with the reconstruction requirements of multi-scale interference.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for controlling a robotic arm against interference based on visual servoing, which comprises:
[0008] S1: Use an industrial camera to capture the initial image of the robot arm during operation and perform preprocessing;
[0009] S2: After using the random walk algorithm to select a pixel point from the center area of the initial image as the starting point, the step length in the random direction is generated and the next coordinate of the starting point is calculated for clipping. The pixel value perturbation is performed on the obtained walk path set to generate a pseudo-abnormal image. The two-dimensional position encoding is performed based on the pseudo-abnormal image to obtain the embedded sequence vector;
[0010] The wandering path set includes pixel values;
[0011] S3: Based on the embedded sequence vector, the restoration model is used to obtain the restoration image, and the deconvolution operation is combined to generate the reconstructed image;
[0012] S4: Based on the pseudo-anomaly and the reconstructed image, the difference value of the pixel value is calculated to perform regional masking, and the Gaussian blur image is obtained. The second-order derivative and second-order partial derivative are extracted, the Hessian matrix is constructed, the feature points are defined, the Haar wavelet response is calculated, and the descriptor vector is generated for clustering, updating and interpolation to obtain the final repaired image;
[0013] S5: The speed and torque are acquired through the encoder and EtherCAT protocol, and the position is identified using OpenCV. Combined with the repaired image, the visual error is calculated for mapping, which is defined as the joint angle error. The least squares method is used to solve the joint angle error vector, which is then time-aligned and fused to obtain the input vector. The input vector is iteratively optimized using the PID control algorithm, and the control parameters are output and input into the actuator before being stored in the database.
[0014] As a preferred solution of the visual servo-based anti-interference control method for a robotic arm according to the present invention, the method comprises the following steps: selecting a pixel point as a starting point from the central area of an initial image using a random walk algorithm, generating a step length in a random direction, calculating the next step coordinates of the starting point for clipping, obtaining a set of walk paths for pixel value perturbation, and generating a pseudo-abnormal image, which refers to randomly selecting a pixel point as a starting point from the central area of the initial image using a random walk algorithm, and calculating the initial position coordinates of the starting point based on the height and width;
[0015] Use a pseudo-random number generator to generate a step length in a random direction. Combined with the initial position coordinates, calculate the next step coordinates of the starting point. Use the clipping operation to clip the next step coordinates and continue to iterate the random walk algorithm. After the iteration is completed, output the walk path set, extract all the coordinates in the walk path set, and record the pixel value corresponding to each coordinate. Perturb each pixel value and define it as a new pixel value.
[0016] The new pixel coordinate values of the cropping operation are further used to re-crop and then combined to obtain a pseudo abnormal image.
[0017] As a preferred solution of the visual servo-based anti-interference control method for a robotic arm of the present invention, wherein: the method of obtaining a repaired image using a repair model based on an embedded sequence vector and generating a reconstructed image by combining a deconvolution operation refers to constructing a repair model using a Transformer network architecture, including an input layer, 13 Transformer layers, and an output layer;
[0018] The 13 Transformer layers all contain multi-layer perceptrons and multi-head self-attention mechanisms, and the Transformer layers are connected through residual connection layers;
[0019] The output layer includes linear projection and reshape operations;
[0020] Define the mean square error as the loss function, and use the Adam optimizer to iteratively optimize the parameter combination of the repair model. During the iterative process, when the loss value of the loss function no longer decreases significantly, stop the iteration and output the final repair model;
[0021] The embedded sequence vector is input into the final restoration model. The embedded sequence vector is input into the 13-layer Transformer layer through the input layer and subjected to QKV linear transformation to generate query, key, and value vectors. After calculating the attention weights by combining the multi-head self-attention mechanism, a residual connection is performed through the residual connection layer and used as the input of the multi-layer perceptron for dimensionality reduction or dimensionality increase to obtain the reconstructed vector.
[0022] Perform linear projection and reshape operations on the reconstructed vector through the output layer to obtain a spatial feature map, which is defined as the repaired image;
[0023] A deconvolution operation is performed based on the repaired image to obtain a reconstructed image.
[0024] As a preferred solution of the visual servo-based anti-interference control method for a robotic arm according to the present invention, the method comprises the following steps: based on the pseudo-anomaly and the reconstructed image, the difference value of the pixel value is calculated to perform regional masking, a Gaussian blur image is obtained, and the second-order derivative and the second-order partial derivative are extracted to construct a Hessian matrix, and the feature points are defined to extract the pixel values in the reconstructed image and the pseudo-anomaly image, and the absolute value is used to calculate the difference between the pixel values to obtain a residual image;
[0025] Use rules of thumb to set judgment thresholds , combine the difference values in the residual image to perform regional masking, combine all regional mask values greater than the judgment threshold to obtain a mask image, use the Gaussian pyramid to generate a Gaussian blur image at each scale of the mask image, and use the Sobel operator to perform kernel approximation on the Gaussian blur image to obtain the second-order derivative and second-order partial derivative of the Gaussian blur image at each scale. Based on the second-order derivative and second-order partial derivative, construct the Hessian matrix of the Gaussian blur image at each scale, and calculate the determinant value and trace;
[0026] According to the determinant value and trace, calculate the response value of the pixel point in the Gaussian blurred image at each scale;
[0027] The response threshold is set using the domain method, and the maximum response value is selected using the maximization operation. The response value is compared with the response threshold and the maximum response value. When the response value is greater than the response threshold and the maximum response value, it is retained and defined as a feature point. When the response value is less than the response threshold and the maximum response value or less than any one of them, it is eliminated.
[0028] As a preferred solution of the anti-interference control method of the manipulator based on visual servoing of the present invention, wherein: the calculation of the Haar wavelet response, the generation of the descriptor vector for clustering, updating and interpolation, and the acquisition of the final repaired image refer to setting the retained feature point as the center, and setting the neighborhood radius by the scale multiple, using the Euclidean distance, calculating the distance value between the center and all pixels within the neighborhood radius, and using the Gaussian kernel function, combining the residual value and the distance value, calculating the direction of the retained feature point , and then use the weighted summation formula to calculate the weighted response value of the direction;
[0029] According to the weighted response value, the maximum value is selected using the maximization operation and defined as the main direction. The radius area is rotated, normalized and divided according to the main direction, and the Haar wavelet response value of each divided area is calculated;
[0030] Use the Gaussian kernel function combined with the residual value to calculate the weight of each pixel in the divided area;
[0031] Use the weighted summation formula and the weights to calculate the total Haar wavelet response values in the horizontal and vertical gradients, respectively, and obtain the total Haar wavelet response values of the horizontal gradient and the total Haar wavelet response value of the vertical gradient ;
[0032] Using feature splicing technology, the Haar wavelet response values are spliced to obtain a response vector, which is defined as a descriptor vector. The elbow rule is used to set the initial number of clusters, and the K-means algorithm is used to cluster the descriptor vectors in combination with the initial number of clusters. During the clustering process, the cluster center is defined as a structural primitive. After the iteration is completed, the structural primitive set is output, and a sparse dictionary is constructed based on the structural primitive set.
[0033] Use Lasso regression method combined with sparse dictionary for image restoration;
[0034] Define the objective function;
[0035] Use the FISTA algorithm for iterative solution. During the iteration process, use the historical regression analysis method to set the adjustment coefficient and update the current iteration result according to the adjustment coefficient.
[0036] After the iteration is completed, the restored image is output and repaired using the image interpolation method to obtain the final repaired image.
[0037] As a preferred solution of the visual servo-based anti-interference control method of the robotic arm described in the present invention, the fusion refers to obtaining an interaction matrix through the visual servo of the robotic arm to map the visual error, which is defined as the joint angle error. The joint angle error is solved using the least squares method to obtain the joint angle error vector. The joint velocity, joint torque and joint angle error vector are time-aligned, and after setting the weights using empirical rules, weighted fusion is performed using the weight distribution formula to obtain the input vector.
[0038] As a preferred solution of the visual servo-based robotic arm anti-interference control method described in the present invention, the storage through a database refers to storing the final repaired image and control parameters through a database, and after storage, adding different IDs to the two types of data, and classifying and storing the two types of data through a database.
[0039] In a second aspect, the present invention provides a robotic arm anti-interference control system based on visual servoing, comprising:
[0040] The acquisition and processing module is used to collect the initial image of the robot arm during operation through an industrial camera and perform preprocessing;
[0041] The walk module is used to select pixels and generate random direction steps, calculate the next step coordinates for cropping and perturbation, generate pseudo-abnormal images for two-dimensional position encoding, and obtain the embedded sequence vector;
[0042] The reconstruction module is used to obtain the repaired image based on the embedded sequence vector using the repair model and generate the reconstructed image in combination with the deconvolution operation;
[0043] The restoration module is used to calculate the difference value of pixel values for regional masking, obtain the Gaussian blur map, extract the second-order derivative and second-order partial derivative, construct the Hessian matrix, define feature points, calculate the Haar wavelet response, generate the descriptor vector for clustering, updating and interpolation, and obtain the final restoration image;
[0044] The feedback module is used to obtain velocity and torque, identify position, and calculate visual errors based on the repaired image for mapping. This error is defined as the joint angle error and solved using the least squares method to obtain the joint angle error vector. This is then time-aligned and fused to obtain the input vector. The input vector is iteratively optimized using the PID control algorithm, and the control parameters are output and input to the actuator.
[0045] The storage module is used to store data through a database.
[0046] The beneficial effects of the present invention are as follows: by introducing a random walk algorithm, a Transformer network structure combined with a multi-head self-attention mechanism, a Gaussian pyramid and a sparse dictionary, the present invention can significantly improve the perception robustness, task execution stability and generalization ability of the robotic arm in a disturbed environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1 Flowchart of the anti-interference control method of the robotic arm based on visual servoing in Example 1.
[0049] Figure 2 This is a structural diagram of the robotic arm anti-interference control system based on visual servoing in Example 1.
[0050] Figure 3 This is a flowchart for generating a pseudo abnormal image in Example 1. DETAILED DESCRIPTION
[0051] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0052] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0053] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0054] Example 1, reference Figures 1-3 , which is the first embodiment of the present invention, provides a method for controlling a robotic arm against interference based on visual servoing, comprising the following steps:
[0055] S1: Use an industrial camera to capture the initial image of the robot arm during operation and perform preprocessing;
[0056] Specifically, collecting the initial image of the robot arm when it is running by using the industrial camera and performing preprocessing means collecting the initial image of the robot arm when it is running by using the industrial camera, and denoising and normalizing the initial image;
[0057] By collecting the "initial image" in the early stage of the robot arm's movement, the reference state of the working environment can be recorded, which can be used for dynamic change detection and target difference recognition in subsequent image processing. De-noising and normalization lay the foundation for subsequent image recognition, edge extraction and other steps.
[0058] S2: After using the random walk algorithm to select a pixel point from the center area of the initial image as the starting point, the step length in the random direction is generated and the next coordinate of the starting point is calculated for clipping. The pixel value perturbation is performed on the obtained walk path set to generate a pseudo-abnormal image. The two-dimensional position encoding is performed based on the pseudo-abnormal image to obtain the embedded sequence vector;
[0059] Specifically, after selecting a pixel point as the starting point from the center area of the initial image using the random walk algorithm, generating a step length in a random direction and calculating the next step coordinates of the starting point for clipping, the walking path set is obtained for pixel value perturbation to generate a pseudo-abnormal image. This refers to using the random walk algorithm to randomly select a pixel point as the starting point from the center area of the initial image, and calculating the initial position coordinates of the starting point based on the height and width;
[0060] Use a pseudo-random number generator to generate a step length in a random direction, and combine it with the initial position coordinates to calculate the next step coordinates of the starting point. The formula is:
[0061]
[0062]
[0063] Where, Indicates the next step horizontal coordinate of the starting point, Indicates the initial horizontal coordinate of the starting point, represents the horizontal coordinate step size of the random direction, Indicates the next vertical coordinate of the starting point, represents the initial ordinate of the starting point, Indicates the vertical coordinate step size of the random direction;
[0064] Use the clipping operation to clip the next step coordinates and continue to iterate the random walk algorithm. After the iteration is completed, output the walk path set;
[0065] Extract all coordinates in the wandering path set, record the pixel value corresponding to each coordinate, and perturb each pixel value to define it as a new pixel value. The formula is:
[0066]
[0067]
[0068] Where, represents the disturbance value, Represents the recorded pixel value, Represents the new pixel value after perturbation;
[0069] The new pixel coordinate values of the cropping operation are further used to re-crop and then combined to obtain a pseudo abnormal image.
[0070] By selecting a random starting point from the center of the image, the perturbed area is ensured to be in the main visual focus area, thereby improving the effectiveness of the pseudo-anomaly features. At the same time, the central area usually has rich semantic features, which is more conducive to anomaly detection and obtaining the key anomaly distribution pattern. Compared with regular patterns or uniform masks, random walk paths are nonlinear and irregular, and are closer to the spatial distribution of actual industrial defects (such as scratches, corrosion, and pollution). They can construct pseudo-anomaly regions with certain "biological characteristics", improving the authenticity and diversity of samples. While maintaining the original image structure and semantic integrity, this operation finely controls the perturbation amplitude and range and provides "hidden anomaly" samples, which is conducive to the subsequent identification of subtle and marginalized defect areas in complex scenes. During each path generation and post-perturbation boundary cropping process, the pseudo-noise or unreasonable deformation caused by the walk path overflowing the image edge is effectively suppressed, ensuring the structural rationality and visual continuity of the final pseudo-anomaly image, providing a guarantee for image quality control.
[0071] Furthermore, two-dimensional position encoding is performed based on the pseudo-abnormal image to obtain an embedded sequence vector, which means that the initial image and the pseudo-abnormal image are divided using a sliding window technique to obtain sub-regions, flatten each sub-region to obtain the vector corresponding to each sub-region, use linear transformation to project the vector to the embedding dimension, obtain the final embedding vector, and record the dimension;
[0072] Use sine and cosine functions to encode the two-dimensional position of each sub-region, the formula is:
[0073]
[0074]
[0075] Where, represents a natural number, Indicates a round-down operation. represents the embedding dimension, Indicates the Rank The sub-area is in the embedding dimension Positional encoding on represents the sine function, represents the total number of embedding dimensions, represents the cosine function, Represents the modulo operator;
[0076] Using an addition operation, the embedding vector is summed with the two-dimensional positional encoding to form an embedding sequence vector.
[0077] By using a sliding window to divide the image into regions, the local texture, edge, and structural change features can be effectively captured. Especially when there are outliers or non-uniform distributions, sub-region division can enhance the fine-grained description of the image. This operation also retains the spatial continuity of the image, providing ordered input for the subsequent process, and can be standardized into a unified embedding dimension through linear projection (such as weight matrix mapping). The use of sine and cosine functions for position encoding has three major advantages: first, the periodic characteristics can simulate the symmetry and repeatability of local image structures; second, the analytical expression can be generated without parameters, reducing training pressure; third, it is highly scalable and suitable for arbitrary image resolution and embedding dimension. Secondly, by adding the embedding vector to the position encoding vector, an embedding sequence with both "content features" and "spatial position" is constructed. This fusion mechanism significantly improves the integrity of image semantic understanding.
[0078] S3: Based on the embedded sequence vector, the restoration model is used to obtain the restoration image, and the deconvolution operation is combined to generate the reconstructed image;
[0079] Specifically, based on the embedded sequence vector, a restoration model is used to obtain the restoration image, and combined with the deconvolution operation, a reconstructed image is generated. This refers to building a restoration model using a Transformer network architecture, including an input layer, 13 Transformer layers, and an output layer;
[0080] The 13 Transformer layers all contain multi-layer perceptrons and multi-head self-attention mechanisms, and the Transformer layers are connected through residual connection layers;
[0081] The output layer includes linear projection and reshape operations;
[0082] Define the mean square error as the loss function, and use the Adam optimizer to iteratively optimize the parameter combination of the repair model. During the iterative process, when the loss value of the loss function no longer decreases significantly, stop the iteration and output the final repair model;
[0083] The embedded sequence vector is input into the final restoration model. The embedded sequence vector is input into the 13-layer Transformer layer through the input layer and subjected to QKV linear transformation to generate query, key, and value vectors. After calculating the attention weights by combining the multi-head self-attention mechanism, a residual connection is performed through the residual connection layer and used as the input of the multi-layer perceptron for dimensionality reduction or dimensionality increase to obtain the reconstructed vector.
[0084] The attention weight is calculated by combining the multi-head self-attention mechanism. The formula is:
[0085]
[0086] Where, represents the self-attention of query, key, and value vectors, represents the softmax activation function, represents the query vector, represents the key vector, represents a value vector, represents the transpose operation, represents the dimension of the key vector;
[0087] Perform linear projection and reshape operations on the reconstructed vector through the output layer to obtain a spatial feature map, which is defined as the repaired image;
[0088] A deconvolution operation is performed based on the repaired image to obtain a reconstructed image.
[0089] The embedded sequence vector is input into the input layer. After the input layer receives the embedded sequence vector, it is sent to the 13-layer Transformer layer. The 13-layer Transformer layer is the core of the repair model of this application, and each layer contains two key components, namely the multi-head self-attention mechanism and the multi-layer perceptron. When the embedded sequence vector is input into the 13-layer Transformer layer (the 13-layer Transformer layer is serial), the first layer of Transformer layer performs QKV linear transformation to generate query, key, and value vectors. The attention weight is calculated using the multi-head self-attention mechanism in the key component, and then the dimensionality reduction is performed through the multi-layer perceptron. Or after dimensionality increase, the input of this layer is added to the output of this layer through the residual connection layer, and transmitted to the second layer Transformer layer for the same operation (i.e. attention, MLP, residual connection), and passed in sequence until the last layer to generate a reconstructed vector. After that, the reconstructed vector is transmitted to the output layer through the thirteenth layer Transformer layer (including linear projection and reshape operations), so that the output layer transforms and reorganizes the reconstructed vector to generate a spatial feature map, which is defined as the repaired image. For the repaired image, the deconvolution operation (upsampling technology, used to enlarge the image size and try to restore the details of the image in the process) is used to obtain the reconstructed image;
[0090] By adopting a 13-layer Transformer stacking structure, the present invention can deeply obtain high-order correlations between embedded vectors, especially when there are complex occlusions or abnormal areas in the image, by extracting and fusing contextual information layer by layer, the missing image content can be accurately filled. Compared with traditional convolutional networks, Transformer has the ability to capture long-distance dependencies and has excellent performance in structural recovery of the image. It also maps the image sequence vector into three functional vectors through the QKV mechanism and uses them to calculate the attention weight matrix, so that each sub-region vector can select the information source according to its semantic similarity with other regions, simulating the human "attention concentration" process, thereby improving the quality of contextual information completion in abnormal regions. After pseudo-anomaly processing, the abnormal area often shows blurred boundaries and broken textures. The self-attention mechanism can focus on healthy areas with semantically close semantics to the abnormal area, perform information migration and structural reconstruction, and help generate coherent and natural repaired images with good detail preservation. Secondly, retaining the original embedded features during the repair process is the key to ensuring the realism of the repaired image. The residual connection mechanism can effectively suppress the damage to edge details by the intermediate layer and enhance the perception of the original image of the present invention. This operation maps the high-dimensional semantic vector output by the Transformer back to the two-dimensional image space, restores the original size and layout, and embeds context-aware texture information to complete the transition from "semantic reconstruction" to "image reconstruction", ensuring that the repair result is visually consistent with the real image.
[0091] S4: Based on the pseudo-anomaly and the reconstructed image, the difference value of the pixel value is calculated to perform regional masking, and the Gaussian blur image is obtained. The second-order derivative and second-order partial derivative are extracted, the Hessian matrix is constructed, the feature points are defined, the Haar wavelet response is calculated, and the descriptor vector is generated for clustering, updating and interpolation to obtain the final repaired image;
[0092] Specifically, based on the pseudo-anomaly and reconstructed image, the difference value of the pixel value is calculated to perform regional masking, and the Gaussian blur image is obtained. The second-order derivative and second-order partial derivative are extracted to construct the Hessian matrix. The feature points are defined as the pixel values extracted from the reconstructed image and the pseudo-anomaly image, and the absolute value is used to calculate the difference between the pixel values. The formula is:
[0093]
[0094] Where, Represents pixel points The difference value at the position, Pseudo-abnormal image At the pixel The pixel value at position, Represents the reconstructed image At the pixel The pixel value at position, Indicates absolute value;
[0095] Combine all the difference values to get the residual map;
[0096] Use rules of thumb to set judgment thresholds , combined with the difference value in the residual map to perform region masking, the formula is:
[0097]
[0098] Where, Indicates at pixel point The region mask value at position, Indicates the judgment threshold;
[0099] Combine all the area mask values greater than the judgment threshold to obtain a mask map;
[0100] Use the Gaussian pyramid to generate a Gaussian blurred image at each scale of the mask image, and use the Sobel operator to perform kernel approximation on the Gaussian blurred image to obtain the second-order derivative and second-order partial derivative of the Gaussian blurred image at each scale. Based on the second-order derivative and second-order partial derivative, construct the Hessian matrix of the Gaussian blurred image at each scale, and calculate the determinant value and trace. The formula is:
[0101]
[0102]
[0103] Where, represents the determinant value of the Hessian matrix, Indicates that the Gaussian blur image is The second derivative of the direction, Indicates that the Gaussian blur image is The second derivative of the direction, Indicates that the Gaussian blur image is and The second-order partial derivative in the direction, represents the trace of the Hessian matrix;
[0104] According to the determinant value and trace, the response value of the pixel in the Gaussian blur image at each scale is calculated. The formula is:
[0105]
[0106] Where, Indicates at pixel point Position on scale The response value under represents a constant coefficient, which can be set through experiments and related domain knowledge;
[0107] The response threshold is set using the domain method, and the maximum response value is selected using the maximization operation. The response value is compared with the response threshold and the maximum response value. When the response value is greater than the response threshold and the maximum response value, it is retained and defined as a feature point. When the response value is less than the response threshold and the maximum response value or less than any one of them, it is eliminated.
[0108] By comparing the pseudo-abnormal image with the reconstructed image pixel by pixel and using the absolute difference to construct the residual map, it is helpful to accurately locate the inaccurately reconstructed area in the image, which often corresponds to the real abnormality, texture variation or structural fracture. By setting the threshold to screen the significant difference interval and forming a regional mask, computing resources can be concentrated on the area of interest, reducing redundant calculations and improving feature extraction efficiency. By generating Gaussian blur maps at multiple scales, the edge and texture details of the target can be captured at different image resolutions. The Gaussian pyramid provides multi-scale context perception capabilities, making feature extraction more robust to image scaling, rotation and noise. Secondly, the Sobel operator is used for kernel approximation. It seems to be able to significantly reduce computational complexity, ensuring computational efficiency while retaining the trend of derivative changes. The Hessian matrix determinant is used to measure the local curvature intensity of the pixel point, and the trace value is used to eliminate noise responses. In practical applications, feature points often correspond to pixels with large Hessian determinant and moderate trace value. Such points are stable in image alignment, deformation compensation, and feature matching. Finally, through the joint constraint of response threshold and maximum response value, unstable or low-intensity response points can be effectively filtered out, and only the most discriminative feature points are retained. This "local maximum" strategy takes into account global screening and local optimization, which helps to form a sparse but high-quality feature set and improve matching accuracy.
[0109] Furthermore, the Haar wavelet response is calculated, and the descriptor vector is generated for clustering, updating and interpolation to obtain the final repaired image. The retained feature point is set as the center, and the neighborhood radius is set by the scale multiple. The Euclidean distance is used to calculate the distance value between the center and all pixels within the neighborhood radius. The Gaussian kernel function is used to combine the residual value and the distance value to calculate the direction of the retained feature point. , and then use the weighted summation formula to calculate the weighted response value of the direction. The formula is:
[0110]
[0111] Where, Indicates the The retained feature points are in the direction The weighted response value on , represents the neighborhood set, represents the exponential function, represents the neighborhood pixel position, represents the square of the Euclidean distance, represents the standard deviation of the Gaussian kernel, which can be set by empirical rules;
[0112] According to the weighted response value, the maximum value is selected using the maximization operation and defined as the main direction. The radius area is rotated, normalized and divided according to the main direction, and the Haar wavelet response value of each divided area is calculated. The formula is:
[0113]
[0114]
[0115] Where, Indicates at pixel point The Haar wavelet response value of the horizontal gradient at position, Indicates at pixel point The horizontal gradient at position, Indicates at pixel point The Haar wavelet response value of the vertical gradient at position, Indicates at pixel point The vertical gradient at position, Indicates absolute value;
[0116] Use the Gaussian kernel function combined with the residual value to calculate the weight of each pixel in the divided area. The formula is:
[0117]
[0118] Where, Indicates at pixel point The weight at the position, Indicates the maximum value operation;
[0119] Use the weighted summation formula and the weights to calculate the total Haar wavelet response values in the horizontal and vertical gradients, respectively, and obtain the total Haar wavelet response values of the horizontal gradient and the total Haar wavelet response value of the vertical gradient ;
[0120] Using the feature splicing technique, the Haar wavelet response values are spliced to obtain a response vector, which is defined as a descriptor vector;
[0121] The elbow rule is used to set the initial number of clusters, and the K-means algorithm is used in combination with the initial number of clusters to cluster the descriptor vectors. During the clustering process, the cluster center is defined as the structural primitive. After the iteration is completed, the structural primitive set is output, and a sparse dictionary is constructed based on the structural primitive set.
[0122] Use Lasso regression method combined with sparse dictionary for image restoration;
[0123] Define the objective function, the formula is:
[0124]
[0125] Where, represents the sparse coefficient, represents a sparse dictionary, represents the Gaussian blurred image to be restored, represents the L2 norm, represents the sparse regularization value, which can be fitted by the least squares method, represents the L1 norm;
[0126] The FISTA algorithm is used for iterative solution. During the iteration process, the historical regression analysis method is used to set the adjustment coefficient, and the current iteration result is updated according to the adjustment coefficient. The formula is:
[0127]
[0128] Where, Indicates The updated image at each round iteration, Indicates The restored image at the round iteration, represents the adjustment coefficient, Indicates The restored image at the round iteration;
[0129] After the iteration is completed, the restored image is output and repaired using the image interpolation method to obtain the final repaired image.
[0130] The residual value and Euclidean distance in the neighborhood of the feature point are weighted by the Gaussian kernel function to calculate the main direction response, which helps to capture the significant structural direction of the image in the local area. This step ensures that the subsequent descriptor has the ability to be rotated and normalized, improves the matching stability and repair accuracy, and divides the neighborhood into multiple sectors. By calculating the Haar wavelet response values in the horizontal and vertical directions and using the residual weighting strategy to calculate the response intensity, the recognition ability of diagonal edges and crack fine textures is effectively enhanced. The descriptor vector is clustered by the K-means algorithm, and the cluster center is used as the "structural primitive" to construct a sparse dictionary. The dictionary condenses the most representative local structure in the image, which not only reduces redundancy but also improves the expression of image reconstruction. Sparsity and interpretability, and define an objective function with an L1 regularization term, so that the sparse reconstruction process takes into account both image fidelity (L2 norm) and sparsity control (L1 norm). This mechanism can efficiently filter out the interference structure introduced by pseudo-anomalies and achieve accurate recovery of real texture and edge information. In addition, FISTA introduces historical regression analysis and adjustment coefficient mechanism, so that the image estimation of each step is updated in a more optimized direction. This strategy can significantly improve the iteration efficiency and stability, and show faster convergence speed and better reconstruction quality when processing large-scale image restoration tasks. Secondly, the low-resolution areas or missing pixels in the restored image are reconstructed and supplemented through the interpolation algorithm, realizing a complete mapping from sparse structure to high-quality image.
[0131] S5: The speed and torque are acquired through the encoder and EtherCAT protocol, and the position is identified using OpenCV. Combined with the repaired image, the visual error is calculated for mapping, which is defined as the joint angle error. The least squares method is used to solve the joint angle error vector, which is then time-aligned and fused to obtain the input vector. The input vector is iteratively optimized using the PID control algorithm, and the control parameters are output and input into the actuator before being stored in the database.
[0132] Specifically, fusion refers to obtaining the interaction matrix through the visual servoing of the robotic arm to map the visual error, which is defined as the joint angle error. The joint angle error is solved using the least squares method to obtain the joint angle error vector. The joint velocity, joint torque and joint angle error vector are time-aligned, and the weights are set using the empirical rule. Then, weighted fusion is performed using the weight distribution formula to obtain the input vector.
[0133] By using encoders to read the angles of each joint of the manipulator in real time and combining them with the EtherCAT protocol to obtain speed and torque from the driver, this mechanism can fully grasp the dynamic and static states of the manipulator, providing a highly reliable perception foundation for subsequent control. Visual error directly reflects the degree of deviation of the manipulator during task execution, thus forming a real-time adjustment mechanism based on the target image and achieving precise visual guidance of the manipulator's behavior. Secondly, the interaction matrix realizes the physical mapping from visual error (image space) to joint angle error (control space), serving as a mathematical bridge for the fusion of visual and motion information. This mechanism greatly improves the response speed and accuracy of visual servoing, giving the manipulator stronger task adaptability in unstructured environments. The least squares method is used to solve the joint angle error vector, which can effectively suppress the impact of measurement noise on control accuracy. The joint velocity, torque, and angle error vectors are time-aligned and combined with weighting factors to form a unified control input, which can adapt to the different requirements of "dynamic," "force feedback," or "position compensation" weights for different tasks. The fused input vector is input into the PID control module, and the control parameters are continuously adjusted through an iterative method to optimize the response speed and reduce the steady-state error, resulting in high stability and tracking accuracy.
[0134] Furthermore, storing the data in a database means storing the final restored image and the control parameters in the database, adding different IDs to the two types of data after storage, and storing the two types of data in a classified manner in the database.
[0135] By synchronously writing the final repaired image and the control parameters that generated the image into the database and mapping them one by one through IDs, a traceable path between the image and the control logic is established, which facilitates subsequent comparative analysis.
[0136] This embodiment also provides a robotic arm anti-interference control system based on visual servoing, including:
[0137] The acquisition and processing module is used to collect the initial image of the robot arm during operation through an industrial camera and perform preprocessing;
[0138] The walk module is used to select pixels and generate random direction steps, calculate the next step coordinates for cropping and perturbation, generate pseudo-abnormal images for two-dimensional position encoding, and obtain the embedded sequence vector;
[0139] The reconstruction module is used to obtain the repaired image based on the embedded sequence vector using the repair model and generate the reconstructed image in combination with the deconvolution operation;
[0140] The restoration module is used to calculate the difference value of pixel values for regional masking, obtain the Gaussian blur map, extract the second-order derivative and second-order partial derivative, construct the Hessian matrix, define feature points, calculate the Haar wavelet response, generate the descriptor vector for clustering, updating and interpolation, and obtain the final restoration image;
[0141] The feedback module is used to obtain velocity and torque, identify position, and calculate visual errors based on the repaired image for mapping. This error is defined as the joint angle error and solved using the least squares method to obtain the joint angle error vector. This is then time-aligned and fused to obtain the input vector. The input vector is iteratively optimized using the PID control algorithm, and the control parameters are output and input to the actuator.
[0142] The storage module is used to store data through a database.
[0143] This embodiment also provides a computer device, which is suitable for the case of a robotic arm anti-interference control method based on visual servoing, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the robotic arm anti-interference control method based on visual servoing proposed in the above embodiment.
[0144] The computer device may be a terminal, comprising a processor, memory, a communication interface, a display, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. The internal memory provides an environment for the operating system and computer programs stored in the non-volatile storage media. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication. Wireless communication may be achieved via Wi-Fi, a carrier network, NFC (near-field communication), or other technologies. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen overlay on the display, buttons, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse.
[0145] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for implementing the anti-interference control of a robotic arm based on visual servoing as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0146] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for controlling an anti-interference robot arm based on visual servoing, characterized in that: include, S1: Use an industrial camera to capture the initial image of the robot arm during operation and perform preprocessing; S2: After using the random walk algorithm to select a pixel point from the center area of the initial image as the starting point, the step length in the random direction is generated and the next coordinate of the starting point is calculated for clipping. The pixel value perturbation is performed on the obtained walk path set to generate a pseudo-abnormal image. The two-dimensional position encoding is performed based on the pseudo-abnormal image to obtain the embedded sequence vector; The wandering path set includes pixel values; S3: Based on the embedded sequence vector, the restoration model is used to obtain the restoration image, and the deconvolution operation is combined to generate the reconstructed image; S4: Based on the pseudo-anomaly and the reconstructed image, the difference value of the pixel value is calculated to perform regional masking, and the Gaussian blur image is obtained. The second-order derivative and second-order partial derivative are extracted, the Hessian matrix is constructed, the feature points are defined, the Haar wavelet response is calculated, and the descriptor vector is generated for clustering, updating and interpolation to obtain the final repaired image; S5: The speed and torque are acquired through the encoder and EtherCAT protocol, and the position is identified using OpenCV. Combined with the repaired image, the visual error is calculated for mapping, which is defined as the joint angle error. The least squares method is used to solve the joint angle error vector, which is then time-aligned and fused to obtain the input vector. The input vector is iteratively optimized using the PID control algorithm, and the control parameters are output and input into the actuator before being stored in the database.
2. The method for controlling a robotic arm with anti-interference based on visual servoing according to claim 1, wherein: The method comprises the following steps: selecting a pixel point as a starting point from the center area of the initial image using the random walk algorithm, generating a step length in a random direction, calculating the next step coordinates of the starting point for clipping, obtaining a set of walk paths, performing pixel value perturbations, and generating a pseudo-abnormal image. The method comprises the following steps: selecting a pixel point as a starting point from the center area of the initial image using the random walk algorithm, and calculating the initial position coordinates of the starting point based on the height and width; Use a pseudo-random number generator to generate a step length in a random direction. Combined with the initial position coordinates, calculate the next step coordinates of the starting point. Use the clipping operation to clip the next step coordinates and continue to iterate the random walk algorithm. After the iteration is completed, output the walk path set, extract all coordinates in the walk path set, and record the pixel value corresponding to each coordinate. Perturb each pixel value and define it as a new pixel value. The new pixel coordinate values of the cropping operation are further used to re-crop and then combined to obtain a pseudo abnormal image.
3. The method for controlling a robotic arm with anti-interference based on visual servoing according to claim 2, wherein: The method of using a restoration model to obtain a restoration image based on the embedded sequence vector and combining the deconvolution operation to generate a reconstructed image refers to constructing a restoration model using a Transformer network architecture, including an input layer, 13 Transformer layers, and an output layer. The 13 Transformer layers all contain multi-layer perceptrons and multi-head self-attention mechanisms, and the Transformer layers are connected through residual connection layers; The output layer includes linear projection and reshape operations; Define the mean square error as the loss function, and use the Adam optimizer to iteratively optimize the parameter combination of the repair model. During the iterative process, when the loss value of the loss function no longer decreases significantly, stop the iteration and output the final repair model; The embedded sequence vector is input into the final restoration model. The embedded sequence vector is input into the 13-layer Transformer layer through the input layer and subjected to QKV linear transformation to generate query, key, and value vectors. After calculating the attention weights by combining the multi-head self-attention mechanism, a residual connection is performed through the residual connection layer and used as the input of the multi-layer perceptron for dimensionality reduction or dimensionality increase to obtain the reconstructed vector. Perform linear projection and reshape operations on the reconstructed vector through the output layer to obtain a spatial feature map, which is defined as the repaired image; A deconvolution operation is performed based on the repaired image to obtain a reconstructed image.
4. The method for controlling a robotic arm with anti-interference based on visual servoing according to claim 3, wherein: The method comprises the following steps: based on the pseudo-anomaly and the reconstructed image, calculating the difference value of the pixel value for regional masking, obtaining a Gaussian blur image, extracting the second-order derivative and the second-order partial derivative, constructing a Hessian matrix, defining the feature points to extract the pixel values in the reconstructed image and the pseudo-anomaly image, and using the absolute value to calculate the difference value between the pixel values to combine and obtain a residual image; Use rules of thumb to set judgment thresholds , combine the difference values in the residual image to perform regional masking, combine all regional mask values greater than the judgment threshold to obtain a mask image, use the Gaussian pyramid to generate a Gaussian blur image at each scale of the mask image, and use the Sobel operator to perform kernel approximation on the Gaussian blur image to obtain the second-order derivative and second-order partial derivative of the Gaussian blur image at each scale. Based on the second-order derivative and second-order partial derivative, construct the Hessian matrix of the Gaussian blur image at each scale, and calculate the determinant value and trace; According to the determinant value and trace, calculate the response value of the pixel point in the Gaussian blurred image at each scale; The response threshold is set using the domain method, and the maximum response value is selected using the maximization operation. The response value is compared with the response threshold and the maximum response value. When the response value is greater than the response threshold and the maximum response value, it is retained and defined as a feature point. When the response value is less than the response threshold and the maximum response value or less than any one of them, it is eliminated.
5. The method for controlling a robotic arm with anti-interference based on visual servoing according to claim 4, wherein: The method of calculating Haar wavelet response, generating descriptor vectors for clustering, updating and interpolation to obtain the final repaired image means setting the retained feature point as the center, setting the neighborhood radius by the scale multiple, using Euclidean distance, calculating the distance value between the center and all pixels within the neighborhood radius, and using Gaussian kernel function, combining the residual value and the distance value to calculate the direction of the retained feature point. , and then use the weighted summation formula to calculate the weighted response value of the direction; According to the weighted response value, the maximum value is selected using the maximization operation and defined as the main direction. The radius area is rotated, normalized and divided according to the main direction, and the Haar wavelet response value of each divided area is calculated; Use the Gaussian kernel function combined with the residual value to calculate the weight of each pixel in the divided area; Use the weighted summation formula and the weights to calculate the total Haar wavelet response values in the horizontal and vertical gradients, respectively, and obtain the total Haar wavelet response values of the horizontal gradient and the total Haar wavelet response value of the vertical gradient ; Using feature splicing technology, the Haar wavelet response values are spliced to obtain a response vector, which is defined as a descriptor vector. The elbow rule is used to set the initial number of clusters, and the K-means algorithm is used to cluster the descriptor vectors in combination with the initial number of clusters. During the clustering process, the cluster center is defined as a structural primitive. After the iteration is completed, the structural primitive set is output, and a sparse dictionary is constructed based on the structural primitive set. Use Lasso regression method combined with sparse dictionary for image restoration; Define the objective function; Use the FISTA algorithm for iterative solution. During the iteration process, use the historical regression analysis method to set the adjustment coefficient and update the current iteration result according to the adjustment coefficient. After the iteration is completed, the restored image is output and repaired using the image interpolation method to obtain the final repaired image.
6. The method for controlling a robotic arm with anti-interference based on visual servoing according to claim 5, wherein: The fusion refers to obtaining an interaction matrix through the visual servoing of the robotic arm to map the visual error, which is defined as the joint angle error. The joint angle error is solved using the least squares method to obtain the joint angle error vector. The joint velocity, joint torque and joint angle error vector are time-aligned, and after setting the weights using empirical rules, weighted fusion is performed using the weight distribution formula to obtain the input vector.
7. The method for controlling a robotic arm with anti-interference based on visual servoing according to claim 6, wherein: The storing in a database refers to storing the final restoration image and the control parameters in a database, adding different IDs to the two types of data after storage, and classifying and storing the two types of data in the database.
8. A visual servo-based anti-interference control system for a robotic arm, based on the visual servo-based anti-interference control method for a robotic arm according to any one of claims 1 to 7, characterized in that: include, The acquisition and processing module is used to collect the initial image of the robot arm during operation through an industrial camera and perform preprocessing; The walk module is used to select pixels and generate random direction steps, calculate the next step coordinates for cropping and perturbation, generate pseudo-abnormal images for two-dimensional position encoding, and obtain the embedded sequence vector; The reconstruction module is used to obtain the repaired image based on the embedded sequence vector using the repair model and generate the reconstructed image in combination with the deconvolution operation; The restoration module is used to calculate the difference value of pixel values for regional masking, obtain the Gaussian blur map, extract the second-order derivative and second-order partial derivative, construct the Hessian matrix, define feature points, calculate the Haar wavelet response, generate the descriptor vector for clustering, updating and interpolation, and obtain the final restoration image; The feedback module is used to obtain velocity and torque, identify position, and calculate visual errors based on the repaired image for mapping. This error is defined as the joint angle error and solved using the least squares method to obtain the joint angle error vector. This is then time-aligned and fused to obtain the input vector. The input vector is iteratively optimized using the PID control algorithm, and the control parameters are output and input to the actuator. The storage module is used to store data through a database.
Citation Information
Patent Citations
Electric power inspection robot positioning method based on multi-sensor fusion
CN111739063A
Abnormal gait recognition method and system based on space-time attention enhancement graph convolution
CN113887486A