A visual motion control method and device

By acquiring the initial atlas through the camera and performing real-time motion parameter processing and feature matching to generate an environmental model, the problem of low accuracy in robot motion control is solved, and higher-precision motion control and obstacle avoidance are achieved.

CN116664622BActive Publication Date: 2025-09-09SHENZHEN HAOCHUAN AUTOMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310440265.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-23
Publication Date
2025-09-09
Estimated Expiration
2043-04-23

AI Technical Summary

Technical Problem

Existing robot motion control technology has large motion trajectory planning errors when dealing with irregular work, resulting in low motion control accuracy.

Method used

By using cameras fixed on the motion base and the robotic arm to obtain the initial atlas, real-time motion parameter processing and image denoising are performed, posture features are extracted, feature matching and three-dimensional reconstruction are performed, and an environmental model is generated to control the movement of the robotic arm.

Benefits of technology

The accuracy of robot motion control is improved, and the robot arm's ability to understand the environment is enhanced, enabling it to better avoid obstacles and achieve precise assembly movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116664622B_ABST
    Figure CN116664622B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of motion control technology and discloses a visual motion control method, comprising: acquiring a first initial atlas using a first camera, acquiring a second initial atlas using a second camera in real time, performing motion denoising on the second initial atlas to obtain a second motion atlas; extracting a first pose feature from the first initial atlas, extracting a second pose feature from the second motion atlas, performing three-dimensional reconstruction based on the first pose feature and the second pose feature to obtain a three-dimensional model of the environment, performing object segmentation and semantic recognition on the three-dimensional model of the environment to obtain a semantic set of environmental objects; performing motion annotation on the three-dimensional model of the environment using the semantic set of environmental objects, extracting target motion points and target three-dimensional coordinates of the target motion points, and controlling a robotic arm to perform assembly motion based on the target three-dimensional coordinates. The present invention also proposes a visual motion control device. The present invention can improve the accuracy of data risk identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of motion control technology, and in particular to a visual motion control method and device. Background Art

[0002] Robots are the cornerstone of industrial manufacturing. Their research and development, manufacturing, and application are important indicators of a country's scientific and technological innovation and high-end manufacturing level. With the acceleration of the industrialization process, the number of robots in the industry is also increasing day by day. In order to improve the production efficiency of robots, it is necessary to control the motion of robots.

[0003] Most existing robot motion control technologies are based on pre-planned trajectories. The robot is placed at a fixed position on the assembly line and is controlled to move according to a fixed frequency and a pre-planned motion trajectory. In actual applications, when dealing with irregular target work, the motion control method based on pre-planned trajectories requires staff to manually plan the motion trajectories one by one. However, manual measurement of displacement during motion trajectory planning may lead to large errors in actual work, which may result in low accuracy in robot motion control. Summary of the Invention

[0004] The present invention provides a visual motion control method and device, the main purpose of which is to solve the problem of low accuracy when performing robot motion control.

[0005] To achieve the above objectives, the present invention provides a visual motion control method, comprising:

[0006] Using a first camera fixed to a motion base to acquire a first initial atlas of the robotic arm in real time, using a second camera fixed to the robotic arm to acquire a second initial atlas in real time, obtaining real-time motion parameters of the robotic arm, and performing motion denoising on the second initial atlas based on the real-time motion parameters to obtain a second motion atlas;

[0007] De-noising and fusing the first initial atlas into a first pose image, extracting a first pose feature from the first pose image, de-noising and fusing the second motion atlas into a second pose image, extracting a second pose feature from the second pose image, and performing feature matching on the first pose feature and the second pose feature to obtain a feature matching point set, wherein the de-noising and fusing the first initial atlas into the first pose image includes: performing block splitting on the first initial atlas to obtain a first block group set; selecting the first block group in the first block group set one by one as the target first block group, performing wavelet transform on the target first block group to obtain a target first wavelet coefficient group set; selecting the first wavelet coefficient group in the target first wavelet coefficient group set one by one as the target first wavelet coefficient group, and calculating the blurriness corresponding to the target first wavelet coefficient group using the following block blurriness formula:

[0008]

[0009] Among them, S refers to the fuzziness, k refers to the layer number of the first target wavelet coefficient group, K refers to the total number of decomposition layers of the first target wavelet coefficient group, h refers to the row number, N k It refers to the total number of rows after the k-th wavelet decomposition in the first wavelet coefficient group of the target, j refers to the column number, M k Refers to the total number of columns after the k-th wavelet decomposition in the first wavelet coefficient group of the target, w hjk refers to the amplitude value of the wavelet coefficient of the k-th layer, h-th row, and j-th column in the target first wavelet coefficient group, P refers to the total number of pixel rows of the first image block corresponding to the target first wavelet coefficient group, p refers to the row number, Q refers to the total number of pixel columns of the first image block corresponding to the target first wavelet coefficient group, q refers to the column number, and w pqk refers to the amplitude value of the wavelet coefficient in the kth layer, pth row, and qth column of the target first wavelet coefficient group; all ambiguities corresponding to the target first wavelet coefficient group are aggregated into a target ambiguity group, and the ambiguity with the smallest value is screened out from the target ambiguity group as the target ambiguity; the first image block corresponding to the target ambiguity in the target first image block group is used as the target first image block, and all the target first image blocks in the first image block group are spliced ​​into a first pose image;

[0010] Obtaining a first real-time pose of the first camera, calculating a second real-time pose of the second camera based on the first real-time pose, generating camera extrinsics of the second camera based on the second real-time pose, and performing a spatial system transformation on the feature matching point set based on the camera extrinsics to obtain a standard matching point set;

[0011] Performing three-dimensional reconstruction on the first pose image and the second pose image according to the standard matching point set to obtain an environment three-dimensional model, performing object segmentation on the environment three-dimensional model to obtain an environment object model set, performing semantic recognition on the environment object model set to obtain an environment object semantic set, and performing semantic recognition on the environment object model set to obtain an environment object semantic set;

[0012] The environmental three-dimensional model is motion-annotated using the environmental object semantic set to obtain a standard environmental model, a target motion point and the target three-dimensional coordinates of the target motion point are extracted from the standard environmental model, and the robotic arm is controlled to perform assembly motion according to the target three-dimensional coordinates.

[0013] Optionally, performing motion denoising on the second initial atlas according to the real-time motion parameters to obtain a second motion atlas includes:

[0014] establishing a real-time motion model corresponding to the second camera according to the real-time motion parameters;

[0015] Selecting pictures in the second initial atlas one by one as target second initial pictures, obtaining exposure times corresponding to the target second initial pictures, and generating pixel movement trajectories of respective pixels in the target second initial pictures according to the exposure times and the real-time motion model;

[0016] Performing an interpolation operation on all pixel movement trajectories to obtain a picture motion model of the target second initial picture;

[0017] performing deconvolution filtering on the target second initial image according to the image motion model to obtain a target second deconvolution image;

[0018] Grayscale enhancement and image sharpening operations are sequentially performed on the target second deconvolution image to obtain a target second motion image, and all target second motion images are aggregated into a second motion atlas.

[0019] Optionally, establishing a real-time motion model corresponding to the second camera according to the real-time motion parameters includes:

[0020] Obtain the length parameters of each joint of the robotic arm, and use the following forward motion formula to calculate the rotation matrix of the robotic arm based on the length parameters and the angle parameters in the real-time motion parameters:

[0021]

[0022] Where T is the rotation matrix, i is the i-th joint, e is the total number of joints of the manipulator, cos is the cosine function, sin is the sine function, θ iRefers to the offset angle of the i-th joint in the angle parameter, α i Indicates the angle of rotation of the i-th joint in the angle parameter around the forward direction axis of the i-1-th joint coordinate system, d i represents the angle of rotation of the i-th joint in the angle parameter around the vertical axis of the i-1-th joint coordinate system;

[0023] generating a Jacobian matrix of the robotic arm according to the rotation matrix and the angle parameter;

[0024] A real-time motion model of the second camera is generated according to the Jacobian matrix and the real-time motion parameters.

[0025] Optionally, performing deconvolution filtering on the target second initial picture according to the picture motion model to obtain a target second deconvolution picture includes:

[0026] Performing frequency domain conversion on the target second initial image to obtain a target second image frequency domain;

[0027] The target second image frequency domain is deconvolved and filtered using the following deconvolution filtering algorithm and the image motion model to obtain a target second deconvolution frequency domain:

[0028]

[0029] in, It refers to the amplitude of the target second deconvolution frequency domain at the frequency (u, v) at the tth moment, where u refers to the frequency in the horizontal direction, v refers to the frequency in the vertical direction, and t refers to the time of the target second deconvolution frequency domain. * (u, v, t) is the amplitude of the amplitude deconvolution filter at the frequency (u, v) at the tth moment, * is the convolution operator, π refers to pi, σ refers to the standard deviation of the Gaussian distribution, e refers to the Euler number, and x t refers to the position of the image motion model at time t, x0 refers to the initial position of the image motion model, ∈ is a preset constant, and F(u,v,t) refers to the amplitude of the target second image frequency domain at the frequency (u,v) at time t;

[0030] Performing image conversion on the target second deconvolution frequency domain to obtain a target second deconvolution image.

[0031] Optionally, extracting the first pose feature from the first pose image includes:

[0032] Performing multi-layer Gaussian filtering on the first pose image to obtain a first Gaussian image group;

[0033] performing an image difference operation on two adjacent first Gaussian images in the first Gaussian image group one by one to obtain a first differential feature group;

[0034] Performing extreme value filtering on the first differential feature group to obtain a first extreme value feature set;

[0035] Performing feature point positioning on each first extreme value feature in the first extreme value feature set to obtain a first initial feature point set;

[0036] Screening out low-contrast feature points and edge feature points from the first initial feature point set to obtain a first standard feature point set;

[0037] performing direction assignment on each first standard feature point in the first standard feature point set to obtain a first direction feature point set;

[0038] Perform feature description on each first direction feature point in the first direction feature point set one by one to obtain first description feature points, and aggregate all the first description feature points into a first pose feature.

[0039] Optionally, performing direction assignment on each first standard feature point in the first standard feature point set to obtain a first direction feature point set includes:

[0040] Selecting first standard feature points in the first standard feature point set one by one as target first standard feature points, and generating target feature areas of the target first standard feature points;

[0041] Dividing the target feature area into an angle sub-area set, and generating a gradient histogram of the target first standard feature point according to the angle sub-area set;

[0042] Taking the maximum gradient in the gradient histogram as the feature direction of the target first standard feature point to obtain the target first direction feature point;

[0043] All target first direction feature points are gathered into a first direction feature point set.

[0044] Optionally, performing feature matching on the first posture feature and the second posture feature to obtain a feature matching point set includes:

[0045] Selecting first description feature points in the first posture feature one by one as target first description feature points, and calculating feature distances between the target first description feature points and each second description feature point in the second posture feature;

[0046] Gathering second description feature points in the second posture feature whose feature distance is less than a preset first distance threshold into a matching second feature point set;

[0047] Screening out noise feature points from the matching second feature point set to obtain a standard matching second feature point set;

[0048] The feature matching points of the target first description feature points are screened out from the standard matching second feature point set using a nearest neighbor algorithm, and all the feature matching points are aggregated into a feature matching point set.

[0049] Optionally, performing a spatial system transformation on the feature matching point set according to the camera extrinsic parameters to obtain a standard matching point set includes:

[0050] Selecting feature matching points in the feature matching point set one by one as target feature matching points, and performing homogeneous coordinate transformation on the target feature matching points to obtain homogeneous coordinates of the target points;

[0051] Back-projecting the homogeneous coordinates of the target point according to the intrinsic parameter matrix of the second camera to obtain the target camera coordinate point;

[0052] Performing world coordinate projection on the target camera coordinate point according to the camera extrinsic parameters to obtain a target primary matching point;

[0053] Performing an inverse homogeneous transformation on the target primary matching points to obtain standard matching points, and gathering all the standard matching points into a standard matching point set.

[0054] Optionally, performing three-dimensional reconstruction on the first pose image and the second pose image according to the standard matching point set to obtain a three-dimensional model of the environment includes:

[0055] Select the standard matching points in the standard matching point set one by one as the target standard matching point, use the pixel point corresponding to the target standard matching point in the first pose image as the first pose matching point, and use the pixel point corresponding to the target standard matching point in the second pose image as the second pose matching point.

[0056] Performing point cloud conversion on the target standard matching points according to the first pose matching points and the second pose matching points using a triangulation algorithm to obtain a target environment point cloud;

[0057] All target environment point clouds are combined into an environment point cloud set, and the environment point cloud set is subjected to point cloud filtering to obtain a filtered point cloud set;

[0058] The first pose image and the second pose image are used to perform surface reconstruction on the filtered point cloud set to obtain a three-dimensional model of the environment.

[0059] In order to solve the above problems, the present invention further provides a visual motion control device, comprising:

[0060] An atlas acquisition module is configured to acquire a first initial atlas of the robotic arm in real time using a first camera fixed to a motion base, acquire a second initial atlas in real time using a second camera fixed to the robotic arm, obtain real-time motion parameters of the robotic arm, and perform motion denoising on the second initial atlas based on the real-time motion parameters to obtain a second motion atlas;

[0061] A feature matching module is used to remove noise from the first initial atlas and fuse it into a first pose image, extract a first pose feature from the first pose image, remove noise from the second motion atlas and fuse it into a second pose image, extract a second pose feature from the second pose image, and perform feature matching on the first pose feature and the second pose feature to obtain a feature matching point set, wherein the removing noise from the first initial atlas and fusing it into the first pose image includes: performing block splitting on the first initial atlas to obtain a first block group set; selecting the first block group in the first block group set one by one as the target first block group, performing wavelet transform on the target first block group to obtain a target first wavelet coefficient group set; selecting the first wavelet coefficient group in the target first wavelet coefficient group set one by one as the target first wavelet coefficient group, and calculating the blur corresponding to the target first wavelet coefficient group using the following block blur formula:

[0062]

[0063] Among them, S refers to the fuzziness, k refers to the layer number of the first target wavelet coefficient group, K refers to the total number of decomposition layers of the first target wavelet coefficient group, h refers to the row number, N k It refers to the total number of rows after the k-th wavelet decomposition in the first wavelet coefficient group of the target, j refers to the column number, M k Refers to the total number of columns after the k-th wavelet decomposition in the first wavelet coefficient group of the target, w hjk refers to the amplitude value of the wavelet coefficient of the k-th layer, h-th row, and j-th column in the target first wavelet coefficient group, P refers to the total number of pixel rows of the first image block corresponding to the target first wavelet coefficient group, p refers to the row number, Q refers to the total number of pixel columns of the first image block corresponding to the target first wavelet coefficient group, q refers to the column number, and w pqk refers to the amplitude value of the wavelet coefficient in the kth layer, pth row, and qth column of the target first wavelet coefficient group; all ambiguities corresponding to the target first wavelet coefficient group are aggregated into a target ambiguity group, and the ambiguity with the smallest value is screened out from the target ambiguity group as the target ambiguity; the first image block corresponding to the target ambiguity in the target first image block group is used as the target first image block, and all the target first image blocks in the first image block group are spliced ​​into a first pose image;

[0064] a pose matching module, configured to obtain a first real-time pose of the first camera, calculate a second real-time pose of the second camera based on the first real-time pose, generate camera extrinsics of the second camera based on the second real-time pose, and perform spatial system transformation on the feature matching point set based on the camera extrinsics to obtain a standard matching point set;

[0065] a semantic recognition module, configured to perform three-dimensional reconstruction on the first pose image and the second pose image according to the standard matching point set to obtain a three-dimensional model of the environment, perform object segmentation on the three-dimensional model of the environment to obtain a set of environment object models, perform semantic recognition on the set of environment object models to obtain a semantic set of environment objects, and perform semantic recognition on the set of environment object models to obtain a semantic set of environment objects;

[0066] The motion control module is used to use the environmental object semantic set to perform motion annotation on the environmental three-dimensional model to obtain a standard environmental model, extract the target motion point and the target three-dimensional coordinates of the target motion point from the standard environmental model, and control the robot arm to perform assembly motion according to the target three-dimensional coordinates.

[0067] In an embodiment of the present invention, by using a first camera fixed to a motion base to acquire a first initial atlas of a robotic arm in real time, and using a second camera fixed to the robotic arm to acquire a second initial atlas in real time, the cost of three-dimensional modeling equipment can be saved, two environmental images in different orientations can be obtained, thereby facilitating subsequent environmental modeling, and the accuracy of the images can be further improved. By acquiring the real-time motion parameters of the robotic arm and performing motion denoising on the second initial atlas based on the real-time motion parameters to obtain a second motion atlas, motion modeling of the robotic arm end can be achieved, and motion blur in the second initial atlas caused by the robotic arm motion can be eliminated, thereby improving subsequent control accuracy. By denoising and fusing the first initial atlas into a first pose image, image clarity can be improved, image details can be preserved, and the complexity of subsequent calculations can be reduced. By extracting the first pose features from the first pose image, subsequent image feature matching can be facilitated. By performing feature matching on the first pose features and the second pose features to obtain a feature matching point set, a correspondence between each identical object in different images captured by the first camera and the second camera can be established, thereby facilitating the subsequent establishment of a three-dimensional environment model.

[0068] By obtaining a first real-time pose of the first camera, calculating a second real-time pose of the second camera based on the first real-time pose, generating camera extrinsics of the second camera based on the second real-time pose, and performing spatial transformation on the feature matching point set based on the camera extrinsics to obtain a standard matching point set, each feature matching point can be converted into a three-dimensional spatial coordinate, facilitating the subsequent establishment of a three-dimensional model of the environment. By performing three-dimensional reconstruction on the first pose image and the second pose image based on the standard matching point set to obtain a three-dimensional environment model, a model of the environment surrounding the robotic arm can be obtained. By performing object segmentation on the three-dimensional environment model to obtain an environmental object model set, and performing semantic recognition on the environmental object model set to obtain an environmental object semantic set, the robotic arm can understand the position and spatial structure of each object in the environment, thereby facilitating motion control. By using the environmental object semantic set to perform motion annotation on the three-dimensional environment model to obtain a standard environment model, target motion points and target three-dimensional coordinates of the target motion points are extracted from the standard environment model. The robotic arm is controlled to perform assembly motion based on the target three-dimensional coordinates, which can improve the robotic arm's ability to avoid obstacles and improve the accuracy of the robotic arm's motion control. Therefore, the visual motion control method and device proposed in the present invention can solve the problem of low accuracy when performing robot motion control. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 A schematic flow chart of a visual motion control method provided by one embodiment of the present invention;

[0070] Figure 2 A schematic diagram of a process for performing motion noise removal according to an embodiment of the present invention;

[0071] Figure 3 A schematic diagram of a process for generating a standard matching point set according to an embodiment of the present invention;

[0072] Figure 4 A functional module diagram of a visual motion control device provided by an embodiment of the present invention;

[0073] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0074] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0075] The embodiment of the present application provides a visual motion control method. The execution subject of the visual motion control method includes but is not limited to at least one of the electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the visual motion control method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0076] Reference Figure 1 FIG. 1 is a flow chart of a visual motion control method according to an embodiment of the present invention. In this embodiment, the visual motion control method includes:

[0077] S1. Use a first camera fixed on a motion base to obtain a first initial atlas of a robotic arm in real time, use a second camera fixed on the robotic arm to obtain a second initial atlas in real time, obtain real-time motion parameters of the robotic arm, and perform motion denoising on the second initial atlas based on the real-time motion parameters to obtain a second motion atlas.

[0078] In an embodiment of the present invention, the motion base is a fixed base of the robotic arm, and the motion base is connected to a motion controller. The motion controller can be an XPLC500 controller or an HMC912BE-2 controller. The motion controller is used to obtain images captured by the first camera and the second camera, generate a three-dimensional environment model, and control the movement of the robotic arm.

[0079] In detail, the first camera and the second camera are high-definition cameras with visible light. The models of the first camera and the second camera can be MV-GE1600C-T camera, OMT-918C camera or XGA-130VM-T camera. The second camera is fixedly installed at the end of the robotic arm. The robotic arm refers to a high-precision, multi-input and multi-output, highly nonlinear, and strongly coupled multi-axis robot. Due to its unique operational flexibility, it has been widely used in industrial assembly, safety and explosion-proof fields, etc., wherein the first camera and the second camera have been calibrated with internal parameters in advance and camera distortion correction has been performed.

[0080] Specifically, the first initial atlas is an atlas consisting of pictures of the robotic arm and the environment surrounding the robotic arm, and the second initial atlas is an atlas consisting of pictures of the working target of the robotic arm and the surrounding environment.

[0081] In detail, the use of the first camera fixed on the motion base to obtain the first initial atlas of the robotic arm in real time means using the first camera fixed on the motion base to continuously shoot the direction of the robotic arm multiple times to obtain a first initial atlas composed of multiple first initial pictures, and the use of the second camera fixed on the robotic arm to obtain the second initial atlas in real time means using the second camera fixed on the front end of the robotic arm to continuously shoot the direction of the end of the robotic arm multiple times to obtain a second initial atlas composed of multiple second initial pictures.

[0082] In detail, the real-time motion parameters refer to parameters such as the angle parameters, velocity parameters and acceleration parameters of the end of the robotic arm. The angle parameters refer to the offset angle of each joint of the robotic arm. The velocity parameters refer to the joint angular velocity and Cartesian coordinate velocity of each joint of the robotic arm, etc., which reflect the speed information of the robotic arm. The acceleration parameters refer to the joint angular acceleration and Cartesian coordinate acceleration of each joint of the robotic arm, etc., which reflect the acceleration information of the robotic arm.

[0083] Specifically, the real-time motion parameters of the robotic arm can be acquired using velocity sensors and acceleration sensors on various joints of the robotic arm.

[0084] In the embodiment of the present invention, referring to Figure 2 As shown, the performing motion denoising on the second initial atlas according to the real-time motion parameters to obtain a second motion atlas includes:

[0085] S21, establishing a real-time motion model corresponding to the second camera according to the real-time motion parameters;

[0086] S22: Selecting pictures in the second initial atlas one by one as target second initial pictures, obtaining exposure times corresponding to the target second initial pictures, and generating pixel movement trajectories of respective pixels in the target second initial pictures according to the exposure times and the real-time motion model;

[0087] S23, performing an interpolation operation on all pixel movement trajectories to obtain a picture motion model of the target second initial picture;

[0088] S24. Perform deconvolution filtering on the target second initial image according to the image motion model to obtain a target second deconvolution image;

[0089] S25. Perform grayscale enhancement and image sharpening operations on the target second deconvolution image in sequence to obtain a target second motion image, and collect all the target second motion images into a second motion atlas.

[0090] Specifically, the real-time motion model refers to a motion model of the second camera during the shooting process, that is, a relationship model between the position and time of the second camera.

[0091] In detail, establishing a real-time motion model corresponding to the second camera according to the real-time motion parameters includes:

[0092] Obtain the length parameters of each joint of the robotic arm, and use the following forward motion formula to calculate the rotation matrix of the robotic arm based on the length parameters and the angle parameters in the real-time motion parameters:

[0093]

[0094] Where T is the rotation matrix, i is the i-th joint, e is the total number of joints of the manipulator, cos is the cosine function, sin is the sine function, θ i Refers to the offset angle of the i-th joint in the angle parameter, α i Indicates the angle of rotation of the i-th joint in the angle parameter around the forward direction axis of the i-1-th joint coordinate system, d i represents the angle of rotation of the i-th joint in the angle parameter around the vertical axis of the i-1-th joint coordinate system;

[0095] generating a Jacobian matrix of the robotic arm according to the rotation matrix and the angle parameter;

[0096] A real-time motion model of the second camera is generated according to the Jacobian matrix and the real-time motion parameters.

[0097] In detail, by using the forward kinematics formula to calculate the rotation matrix of the robotic arm according to the length parameter and the angle parameter in the real-time motion parameter, the position parameters of the second camera at the end position of the robotic arm can be determined by the angle offset of each robotic arm joint, thereby improving the calculation accuracy of the subsequent real-time motion model.

[0098] In an embodiment of the present invention, the Jacobian formula can be used to generate the Jacobian matrix of the robotic arm based on the rotation matrix and the angle parameters. The Jacobian matrix is ​​a very important concept in vector calculus, which describes the partial derivative of each output component of a vector function with respect to its input component. In robotics and kinematics, the Jacobian matrix is ​​used to describe the relationship between the movement of the robot end effector, such as the gripper of the robotic arm, and the joint movement, reflecting the mapping relationship between the joint space, joint angle and the position and direction of the robot end effector.

[0099] Specifically, generating the real-time motion model of the second camera according to the Jacobian matrix and the real-time motion parameters means calculating the real-time speed of the second camera according to the Jacobian matrix and the speed parameters in the real-time motion parameters, calculating the real-time acceleration of the second camera according to the Jacobian matrix and the acceleration parameters in the real-time motion parameters, and generating the real-time motion model according to the real-time motion speed, the real-time acceleration and the rotation matrix.

[0100] In detail, the exposure time refers to the time the shutter needs to be open when light is projected onto the photosensitive surface of the photosensitive material of the second camera during the process of capturing the second initial image of the target.

[0101] Specifically, the pixel movement trajectory of each pixel in the second initial image of the target is generated according to the exposure time and the real-time motion model, including: calculating the shooting displacement of the second camera according to the exposure time and the real-time motion model; selecting pixel points in the second initial image of the target as target pixel points one by one, and taking the focal length corresponding to the target pixel point as the target focal length; calculating the pixel movement trajectory of the target pixel point according to the target focal length and the shooting displacement, wherein the pixel movement trajectory of the target pixel point calculated according to the target focal length and the shooting displacement includes calculating the horizontal movement trajectory and the vertical movement trajectory of the target pixel point. For example, the calculation of the horizontal movement trajectory of the target pixel point includes multiplying the focal length of the target pixel point on the horizontal axis by the difference between the horizontal axis coordinate of the second camera and the horizontal axis coordinate of the center point of the pixel coordinate system, and dividing it by the vertical axis coordinate of the second camera.

[0102] Specifically, an interpolation algorithm such as a nearest neighbor interpolation method or a bilinear interpolation method may be used to perform interpolation calculation on all pixel movement trajectories to obtain a picture motion model of the target second initial picture.

[0103] In detail, performing deconvolution filtering on the target second initial picture according to the picture motion model to obtain a target second deconvolution picture includes:

[0104] Performing frequency domain conversion on the target second initial image to obtain a target second image frequency domain;

[0105] The target second image frequency domain is deconvolved and filtered using the following deconvolution filtering algorithm and the image motion model to obtain a target second deconvolution frequency domain:

[0106]

[0107] in, It refers to the amplitude of the target second deconvolution frequency domain at the frequency (u, v) at the tth moment, where u refers to the frequency in the horizontal direction, v refers to the frequency in the vertical direction, and t refers to the time of the target second deconvolution frequency domain. * (u, v, t) is the amplitude of the amplitude deconvolution filter at the frequency (u, v) at the tth moment, * is the convolution operator, π refers to pi, σ refers to the standard deviation of the Gaussian distribution, e refers to the Euler number, and x t refers to the position of the image motion model at time t, x0 refers to the initial position of the image motion model, ∈ is a preset constant, and F(u,v,t) refers to the amplitude of the target second image frequency domain at the frequency (u,v) at time t;

[0108] Performing image conversion on the target second deconvolution frequency domain to obtain a target second deconvolution image.

[0109] In detail, the target second initial image can be transformed into the frequency domain using the fast Fourier transform formula to obtain the target second image frequency domain, and the target second deconvolution frequency domain can be transformed into the image using the inverse Fourier transform formula to obtain the target second deconvolution image. The deconvolution filter can be a deconvreg filter, a Wiener filter or a Lucy-Richardson filter.

[0110] In an embodiment of the present invention, by using the deconvolution filtering algorithm and the image motion model to perform deconvolution filtering on the target second image frequency domain, a target second deconvolution frequency domain is obtained. The mathematical relationship between the image motion model and the Gaussian spread function can be used to eliminate motion blur of the image, thereby improving the clarity of the image.

[0111] Specifically, a grayscale histogram method may be used to perform a grayscale enhancement operation on the target second deconvolution image, and a filtering algorithm such as Canny filtering, Laplace filtering or Sobel filtering may be used to perform an image sharpening operation on the target second deconvolution image to obtain a target second motion image.

[0112] In an embodiment of the present invention, by using a first camera fixed on a motion base to obtain a first initial atlas of a robotic arm in real time, and using a second camera fixed on the robotic arm to obtain a second initial atlas in real time, the cost of three-dimensional modeling equipment can be saved, and environmental images in two different orientations can be obtained, thereby facilitating subsequent environmental modeling and further improving the accuracy of the images. By obtaining the real-time motion parameters of the robotic arm and performing motion denoising on the second initial atlas according to the real-time motion parameters to obtain a second motion atlas, motion modeling of the end of the robotic arm can be realized, and motion blur of the second initial atlas caused by the motion of the robotic arm can be eliminated, thereby improving the subsequent control accuracy.

[0113] S2. De-noise the first initial atlas and fuse it into a first pose image, extract the first pose feature from the first pose image, de-noise the second motion atlas and fuse it into a second pose image, extract the second pose feature from the second pose image, perform feature matching on the first pose feature and the second pose feature to obtain a feature matching point set.

[0114] In the embodiment of the present invention, the first pose image is the clearest image obtained by fusing the first initial atlas.

[0115] In an embodiment of the present invention, the step of removing noise from the first initial atlas and fusing it into the first pose image includes:

[0116] Performing tile splitting on the first initial atlas to obtain a first tile group set;

[0117] selecting first image block groups in the first image block group set one by one as target first image block groups, performing wavelet transform on the target first image block groups to obtain target first wavelet coefficient groups;

[0118] The first wavelet coefficient groups in the target first wavelet coefficient group set are selected one by one as the target first wavelet coefficient group, and the blurriness corresponding to the target first wavelet coefficient group is calculated using the following block blurriness formula:

[0119]

[0120] Among them, S refers to the fuzziness, k refers to the layer number of the first target wavelet coefficient group, K refers to the total number of decomposition layers of the first target wavelet coefficient group, h refers to the row number, N k It refers to the total number of rows after the k-th wavelet decomposition in the first wavelet coefficient group of the target, j refers to the column number, M k Refers to the total number of columns after the k-th wavelet decomposition in the first wavelet coefficient group of the target, w hjkrefers to the amplitude value of the wavelet coefficient of the k-th layer, h-th row, and j-th column in the target first wavelet coefficient group, P refers to the total number of pixel rows of the first image block corresponding to the target first wavelet coefficient group, p refers to the row number, Q refers to the total number of pixel columns of the first image block corresponding to the target first wavelet coefficient group, q refers to the column number, and w pqk refers to the amplitude value of the wavelet coefficient in the k-th layer, p-th row, and q-th column in the target first wavelet coefficient group;

[0121] Gathering all ambiguities corresponding to the target first wavelet coefficient group into a target ambiguity group, and selecting the ambiguity with the smallest value from the target ambiguity group as the target ambiguity;

[0122] The first picture block corresponding to the target blur in the target first picture block group is used as the target first picture block, and all the target first picture blocks in the first picture block group are spliced ​​into a first pose image.

[0123] Specifically, the splitting of the first initial atlas into blocks to obtain the first block group set means selecting the first initial pictures in the first initial atlas one by one as the target first initial pictures, performing multi-level quadtree blocking, ternary tree blocking and binary tree blocking on the target first initial pictures in turn to obtain the first block group, and aggregating all the first block groups into the first block group set.

[0124] In detail, performing wavelet transform on the target first picture block group to obtain a target first wavelet coefficient group set refers to selecting the first picture blocks in the target first picture block group one by one as the target first picture block, continuously filtering the target first picture block using filters such as Haar wavelet filters, Daubechies wavelet filters and Coiflets wavelet filters, and performing wavelet transform on multiple subbands obtained by filtering using wavelet transform algorithms such as two-dimensional discrete wavelet transform or two-dimensional continuous wavelet transform to obtain a first wavelet coefficient group, and gathering all the first wavelet coefficient groups into a target first wavelet coefficient group set.

[0125] In the embodiment of the present invention, by using the block blur formula to calculate the blur corresponding to the target first wavelet coefficient group, the blur of the image can be determined according to the grayscale frequency domain of the wavelet transformed image, thereby improving the recognition accuracy of the image blur.

[0126] In detail, the method of denoising and fusing the second motion atlas into a second pose image and extracting the second pose features from the second pose image is consistent with the method of denoising and fusing the first initial atlas into a first pose image and extracting the first pose features from the first pose image described in the above steps, and will not be repeated here.

[0127] Specifically, extracting the first pose feature from the first pose image includes:

[0128] Performing multi-layer Gaussian filtering on the first pose image to obtain a first Gaussian image group;

[0129] performing an image difference operation on two adjacent first Gaussian images in the first Gaussian image group one by one to obtain a first differential feature group;

[0130] Performing extreme value filtering on the first differential feature group to obtain a first extreme value feature set;

[0131] Performing feature point positioning on each first extreme value feature in the first extreme value feature set to obtain a first initial feature point set;

[0132] Screening out low-contrast feature points and edge feature points from the first initial feature point set to obtain a first standard feature point set;

[0133] performing direction assignment on each first standard feature point in the first standard feature point set to obtain a first direction feature point set;

[0134] Perform feature description on each first direction feature point in the first direction feature point set one by one to obtain first description feature points, and aggregate all the first description feature points into a first pose feature.

[0135] In detail, performing multi-layer Gaussian filtering on the first pose image to obtain a first Gaussian image group means performing multiple Gaussian filtering on the first pose image to generate a series of first Gaussian images with different scales and resolutions, wherein the resolution of the first Gaussian images of each layer is lower than that of the previous layer, but the first Gaussian images in the same layer have the same scale.

[0136] Specifically, the first differential feature group can be subjected to primary extreme value filtering using a maximum filter to obtain a first primary extreme value feature set, and the first primary extreme value features in the first primary extreme value feature set that are less than a preset feature threshold are filtered out to obtain the first extreme value feature set.

[0137] In detail, the principal curvature ratio of the Hessian matrix or the local interpolation algorithm can be used to locate the feature points of each first extreme feature in the first extreme feature set to obtain a first initial feature point set. The low-contrast feature point refers to the first initial feature point in the first initial feature point set whose average grayscale difference of the pixel points around the feature point is less than a preset grayscale threshold. The edge feature point refers to the first initial feature point in the first initial feature point set whose gradient direction of the pixel points around the feature point is very different from the main gradient direction around the feature point.

[0138] In the embodiment of the present invention, the step of assigning directions to each first standard feature point in the first standard feature point set to obtain the first directional feature point set includes:

[0139] Selecting first standard feature points in the first standard feature point set one by one as target first standard feature points, and generating target feature areas of the target first standard feature points;

[0140] Dividing the target feature area into an angle sub-area set, and generating a gradient histogram of the target first standard feature point according to the angle sub-area set;

[0141] Taking the maximum gradient in the gradient histogram as the feature direction of the target first standard feature point to obtain the target first direction feature point;

[0142] All target first direction feature points are gathered into a first direction feature point set.

[0143] Specifically, the target feature area is a circular area with the target first standard feature point as the center point, the angle sub-area set is a set composed of multiple angle sub-areas, and the angle range of each angle sub-area is 30 degrees to 45 degrees, and generating the gradient histogram of the target first standard feature point based on the angle sub-area set refers to calculating the gradient amplitude, gradient direction and accumulated gradient amplitude of the pixels in each angle sub-area in the angle sub-area set to obtain a gradient histogram.

[0144] Specifically, the feature description of each first direction feature point in the first direction feature point set one by one to obtain the first description feature point means describing the first direction feature point by the pixel values ​​and directions around each first direction feature point to obtain the feature vector of the first direction feature point, and using the feature vector as the descriptor of the first direction feature point to obtain the first description feature point.

[0145] In detail, performing feature matching on the first posture feature and the second posture feature to obtain a feature matching point set includes:

[0146] Selecting first description feature points in the first posture feature one by one as target first description feature points, and calculating feature distances between the target first description feature points and each second description feature point in the second posture feature;

[0147] Gathering second description feature points in the second posture feature whose feature distance is less than a preset first distance threshold into a matching second feature point set;

[0148] Screening out noise feature points from the matching second feature point set to obtain a standard matching second feature point set;

[0149] The feature matching points of the target first description feature points are screened out from the standard matching second feature point set using a nearest neighbor algorithm, and all the feature matching points are aggregated into a feature matching point set.

[0150] In detail, the cosine distance algorithm or the Euclidean distance algorithm can be used to calculate the feature distance between the first description feature point of the target and each second description feature point in the second posture feature, and the RANSAC algorithm or the threshold-based screening algorithm can be used to filter out noise feature points from the matching second feature point set to obtain a standard matching second feature point set.

[0151] In an embodiment of the present invention, by denoising and fusing the first initial atlas into a first pose image, the image clarity can be improved, the image details can be retained, and the complexity of subsequent calculations can be reduced. By extracting the first pose feature from the first pose image, the subsequent image feature matching can be facilitated. By performing feature matching on the first pose feature and the second pose feature to obtain a feature matching point set, the correspondence between each identical object in different images taken by the first camera and the second camera can be constructed, thereby facilitating the subsequent establishment of a three-dimensional environment model.

[0152] S3. Obtain a first real-time pose of the first camera, calculate a second real-time pose of the second camera based on the first real-time pose, generate camera extrinsics of the second camera based on the second real-time pose, and perform spatial system transformation on the feature matching point set based on the camera extrinsics to obtain a standard matching point set.

[0153] In an embodiment of the present invention, the first real-time pose refers to the position and orientation of the first camera, and the first real-time pose is composed of the rotation matrix and translation vector of the first camera. The second real-time pose refers to the position and orientation of the second camera, and the second real-time pose is composed of the rotation matrix and translation vector of the second camera. The camera extrinsic parameter refers to the projection matrix of the position and orientation of the camera, including the camera's intrinsic parameter matrix and the camera's pose matrix. The first real-time pose of the first camera can be obtained by the gravity sensor and angle sensor on the first camera.

[0154] In the embodiment of the present invention, the method for obtaining the second real-time pose of the second camera calculated based on the first real-time pose is consistent with the method for establishing the real-time motion model corresponding to the second camera based on the real-time motion parameters in the above step S1, and will not be repeated here.

[0155] In an embodiment of the present invention, generating the camera extrinsic parameters of the second camera according to the second real-time pose includes: obtaining the camera intrinsic parameters of the second camera, generating an intrinsic parameter matrix according to the camera intrinsic parameters, extracting a pose matrix from the second real-time pose, and multiplying the intrinsic parameter matrix by the pose matrix to obtain the camera extrinsic parameters of the second camera.

[0156] For details, refer to Figure 3 As shown, the feature matching point set is transformed into a spatial system according to the camera extrinsic parameters to obtain a standard matching point set, including:

[0157] S31, selecting feature matching points in the feature matching point set one by one as target feature matching points, performing homogeneous coordinate transformation on the target feature matching points, and obtaining homogeneous coordinates of the target points;

[0158] S32, back-projecting the homogeneous coordinates of the target point according to the intrinsic parameter matrix of the second camera to obtain the target camera coordinate point;

[0159] S33, performing world coordinate projection on the target camera coordinate point according to the camera extrinsic parameter to obtain a target primary matching point;

[0160] S34 , performing an inverse homogeneous transformation on the target primary matching points to obtain standard matching points, and gathering all the standard matching points into a standard matching point set.

[0161] Specifically, performing homogeneous coordinate conversion on the target feature matching point to obtain the homogeneous coordinate of the target point means adding 1 to the end of the coordinate of the target feature matching point to obtain a three-dimensional coordinate, which facilitates matrix space operations.

[0162] In an embodiment of the present invention, by obtaining the first real-time pose of the first camera, calculating the second real-time pose of the second camera based on the first real-time pose, generating the camera extrinsic parameters of the second camera based on the second real-time pose, and performing spatial system transformation on the feature matching point set based on the camera extrinsic parameters to obtain a standard matching point set, each feature matching point can be transformed into a three-dimensional spatial coordinate, thereby facilitating the subsequent establishment of a three-dimensional model of the environment.

[0163] S4. Perform three-dimensional reconstruction on the first pose image and the second pose image according to the standard matching point set to obtain an environment three-dimensional model, perform object segmentation on the environment three-dimensional model to obtain an environment object model set, and perform semantic recognition on the environment object model set to obtain an environment object semantic set.

[0164] In an embodiment of the present invention, the three-dimensional environmental model refers to a three-dimensional model of the space surrounding the robotic arm. By generating the three-dimensional environmental model, the robotic arm can easily identify spatial information and thus perform precise motion control.

[0165] In an embodiment of the present invention, performing three-dimensional reconstruction on the first pose image and the second pose image according to the standard matching point set to obtain a three-dimensional model of the environment includes:

[0166] Select the standard matching points in the standard matching point set one by one as the target standard matching point, use the pixel point corresponding to the target standard matching point in the first pose image as the first pose matching point, and use the pixel point corresponding to the target standard matching point in the second pose image as the second pose matching point.

[0167] Performing point cloud conversion on the target standard matching points according to the first pose matching points and the second pose matching points using a triangulation algorithm to obtain a target environment point cloud;

[0168] All target environment point clouds are combined into an environment point cloud set, and the environment point cloud set is subjected to point cloud filtering to obtain a filtered point cloud set;

[0169] The first pose image and the second pose image are used to perform surface reconstruction on the filtered point cloud set to obtain a three-dimensional model of the environment.

[0170] Specifically, the surface of the filtered point cloud set can be reconstructed using reconstruction algorithms such as Poisson reconstruction, Marching Cubes reconstruction, and Ball Pivoting reconstruction to obtain a three-dimensional model of the environment; the object segmentation of the three-dimensional model of the environment can be performed using a region growing algorithm or a density clustering algorithm to obtain a set of environmental object models.

[0171] In detail, the semantic recognition of the environmental object model set to obtain the environmental object semantic set includes: selecting environmental object models in the environmental object model set one by one as target environmental object models, performing multi-dimensional feature extraction on the target environmental objects to obtain a target object feature group; performing semantic recognition on the target environmental objects according to the target object feature group to obtain environmental object semantics, and aggregating all environmental object semantics into an environmental object semantic set.

[0172] Specifically, the multi-dimensional feature extraction of the target environment object to obtain the target object feature group refers to extracting the color features, shape features, texture features and other features of the target environment object in sequence, and aggregating them into the target object feature group, wherein the color features, shape features, texture features and other features of the target environment object can be extracted using feature extraction algorithms such as Principal Component Analysis (PCA), Fast Point Feature Histograms (FPFH) and Signature of Histograms of Orientations (SHOT).

[0173] In detail, a trained support vector machine (SVM), convolutional neural network (CNN) or k-nearest neighbors (KNN) neural network or machine learning algorithm can be used to perform semantic recognition of the target environmental object based on the target object feature group to obtain environmental object semantics, and all environmental object semantics are aggregated into an environmental object semantic set.

[0174] In an embodiment of the present invention, by performing three-dimensional reconstruction on the first pose image and the second pose image according to the standard matching point set, a three-dimensional model of the environment is obtained, and a model of the environment surrounding the robotic arm can be obtained. By performing object segmentation on the three-dimensional model of the environment, an environmental object model set is obtained, and semantic recognition is performed on the environmental object model set to obtain an environmental object semantic set. This allows the robotic arm to understand the position and spatial composition of each object in the environment, thereby facilitating motion control.

[0175] S5. Use the environmental object semantic set to perform motion annotation on the environmental three-dimensional model to obtain a standard environmental model, extract the target motion point and the target three-dimensional coordinates of the target motion point from the standard environmental model, and control the robot arm to perform assembly movement according to the target three-dimensional coordinates.

[0176] Specifically, using the environmental object semantic set to perform motion annotation on the environmental three-dimensional model to obtain a standard environmental model means annotating each environmental object model in the environmental three-dimensional model into motion control semantics. For example, when the environmental object semantics of the environmental object is a screw to be assembled, it is marked as a target motion point; when the environmental object semantics of the environmental object is a book, it is marked as an obstacle.

[0177] Specifically, the target motion point refers to the coordinates of the object where the robot arm needs to perform working motion, such as the three-dimensional spatial coordinates of the screws to be assembled. Controlling the robot arm to perform assembly motion according to the target three-dimensional coordinates refers to controlling the robot arm to move to the position of the target three-dimensional coordinates for motion.

[0178] In detail, the relative displacement can be calculated based on the target three-dimensional coordinates and the three-dimensional coordinates of the end of the robotic arm, and the motion path of the robotic arm can be obtained using a path analysis algorithm such as the A algorithm or the ant colony algorithm. The inverse kinematics algorithm is used to calculate the rotation direction and angle of each joint corresponding to the motion path, thereby controlling the robotic arm to perform assembly movement.

[0179] In an embodiment of the present invention, the environment three-dimensional model is motion-annotated by utilizing the environment object semantic set to obtain a standard environment model, and the target motion point and the target three-dimensional coordinates of the target motion point are extracted from the standard environment model. The robot arm is controlled to perform assembly movement according to the target three-dimensional coordinates, thereby improving the ability of the robot arm to avoid obstacles and improving the accuracy of the motion control of the robot arm.

[0180] In an embodiment of the present invention, by using a first camera fixed to a motion base to acquire a first initial atlas of a robotic arm in real time, and using a second camera fixed to the robotic arm to acquire a second initial atlas in real time, the cost of three-dimensional modeling equipment can be saved, two environmental images in different orientations can be obtained, thereby facilitating subsequent environmental modeling, and the accuracy of the images can be further improved. By acquiring the real-time motion parameters of the robotic arm and performing motion denoising on the second initial atlas based on the real-time motion parameters to obtain a second motion atlas, motion modeling of the robotic arm end can be achieved, and motion blur in the second initial atlas caused by the robotic arm motion can be eliminated, thereby improving subsequent control accuracy. By denoising and fusing the first initial atlas into a first pose image, image clarity can be improved, image details can be preserved, and the complexity of subsequent calculations can be reduced. By extracting the first pose features from the first pose image, subsequent image feature matching can be facilitated. By performing feature matching on the first pose features and the second pose features to obtain a feature matching point set, a correspondence between each identical object in different images captured by the first camera and the second camera can be established, thereby facilitating the subsequent establishment of a three-dimensional environment model.

[0181] By obtaining a first real-time pose of the first camera, calculating a second real-time pose of the second camera based on the first real-time pose, generating camera extrinsics of the second camera based on the second real-time pose, and performing spatial transformation on the feature matching point set based on the camera extrinsics to obtain a standard matching point set, each feature matching point can be converted into a three-dimensional spatial coordinate, facilitating the subsequent establishment of a three-dimensional model of the environment. By performing three-dimensional reconstruction on the first pose image and the second pose image based on the standard matching point set to obtain a three-dimensional environment model, a model of the environment surrounding the robotic arm can be obtained. By performing object segmentation on the three-dimensional environment model to obtain an environmental object model set, and performing semantic recognition on the environmental object model set to obtain an environmental object semantic set, the robotic arm can understand the position and spatial structure of each object in the environment, thereby facilitating motion control. By using the environmental object semantic set to perform motion annotation on the three-dimensional environment model to obtain a standard environment model, target motion points and target three-dimensional coordinates of the target motion points are extracted from the standard environment model. The robotic arm is controlled to perform assembly motion based on the target three-dimensional coordinates, which can improve the robotic arm's ability to avoid obstacles and improve the accuracy of the robotic arm's motion control. Therefore, the visual motion control method proposed in the present invention can solve the problem of low accuracy when performing robot motion control.

[0182] like Figure 4 , which is a functional module diagram of a visual motion control device provided by an embodiment of the present invention.

[0183] The visual motion control device 100 described in the present invention can be installed in an electronic device. Depending on the functionality to be implemented, the visual motion control device 100 may include an atlas acquisition module 101, a feature matching module 102, a pose matching module 103, a semantic recognition module 104, and a motion control module 105. A module, also referred to as a unit, is a series of computer program segments that can be executed by an electronic device processor and perform a fixed function, and is stored in the electronic device's memory.

[0184] In this embodiment, the functions of each module / unit are as follows:

[0185] The atlas acquisition module 101 is configured to acquire a first initial atlas of the robotic arm in real time using a first camera fixed to a motion base, acquire a second initial atlas in real time using a second camera fixed to the robotic arm, obtain real-time motion parameters of the robotic arm, and perform motion denoising on the second initial atlas based on the real-time motion parameters to obtain a second motion atlas;

[0186] The feature matching module 102 is used to remove noise from the first initial atlas and fuse it into a first pose image, extract a first pose feature from the first pose image, remove noise from the second motion atlas and fuse it into a second pose image, extract a second pose feature from the second pose image, perform feature matching on the first pose feature and the second pose feature to obtain a feature matching point set, wherein the removing noise from the first initial atlas and fusing it into the first pose image includes: performing block splitting on the first initial atlas to obtain a first block group set; selecting the first block group in the first block group set one by one as the target first block group, performing wavelet transform on the target first block group to obtain a target first wavelet coefficient group set; selecting the first wavelet coefficient group in the target first wavelet coefficient group set one by one as the target first wavelet coefficient group, and calculating the fuzziness corresponding to the target first wavelet coefficient group using the following block fuzziness formula:

[0187]

[0188] Among them, S refers to the fuzziness, k refers to the layer number of the first target wavelet coefficient group, K refers to the total number of decomposition layers of the first target wavelet coefficient group, h refers to the row number, N k It refers to the total number of rows after the k-th wavelet decomposition in the first wavelet coefficient group of the target, j refers to the column number, M k Refers to the total number of columns after the k-th wavelet decomposition in the first wavelet coefficient group of the target, w hjk refers to the amplitude value of the wavelet coefficient of the k-th layer, h-th row, and j-th column in the target first wavelet coefficient group, P refers to the total number of pixel rows of the first image block corresponding to the target first wavelet coefficient group, p refers to the row number, Q refers to the total number of pixel columns of the first image block corresponding to the target first wavelet coefficient group, q refers to the column number, and w pqk refers to the amplitude value of the wavelet coefficient in the kth layer, pth row, and qth column of the target first wavelet coefficient group; all ambiguities corresponding to the target first wavelet coefficient group are aggregated into a target ambiguity group, and the ambiguity with the smallest value is screened out from the target ambiguity group as the target ambiguity; the first image block corresponding to the target ambiguity in the target first image block group is used as the target first image block, and all the target first image blocks in the first image block group are spliced ​​into a first pose image;

[0189] The pose matching module 103 is configured to obtain a first real-time pose of the first camera, calculate a second real-time pose of the second camera based on the first real-time pose, generate camera extrinsics of the second camera based on the second real-time pose, and perform spatial system transformation on the feature matching point set based on the camera extrinsics to obtain a standard matching point set;

[0190] The semantic recognition module 104 is configured to perform three-dimensional reconstruction on the first pose image and the second pose image according to the standard matching point set to obtain a three-dimensional environment model, perform object segmentation on the three-dimensional environment model to obtain an environment object model set, perform semantic recognition on the environment object model set to obtain an environment object semantic set, and perform semantic recognition on the environment object model set to obtain an environment object semantic set;

[0191] The motion control module 105 is used to use the environmental object semantic set to perform motion annotation on the environmental three-dimensional model to obtain a standard environmental model, extract the target motion point and the target three-dimensional coordinates of the target motion point from the standard environmental model, and control the robot arm to perform assembly motion according to the target three-dimensional coordinates.

[0192] In detail, each module in the visual motion control device 100 of the embodiment of the present invention is used in the same manner as above. Figures 1 to 3 The visual motion control method described in the present invention has the same technical means and can produce the same technical effects, so I will not go into details here.

[0193] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical function division, and other division methods may be used in actual implementation.

[0194] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the purpose of the solution of this embodiment according to actual needs.

[0195] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.

[0196] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0197] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.

[0198] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0199] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in a system embodiment may also be implemented by a single unit or device through software or hardware. Terms such as first and second are used to indicate names and do not imply any particular order.

[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A visual motion control method, characterized in that: The method comprises: S1: using a first camera fixed on a motion base to acquire a first initial atlas of a robotic arm in real time, using a second camera fixed on the robotic arm to acquire a second initial atlas in real time, obtaining real-time motion parameters of the robotic arm, and performing motion denoising on the second initial atlas based on the real-time motion parameters to obtain a second motion atlas; S2: Denoising and fusing the first initial atlas into a first pose image, extracting a first pose feature from the first pose image, denoising and fusing the second motion atlas into a second pose image, extracting a second pose feature from the second pose image, and performing feature matching on the first pose feature and the second pose feature to obtain a feature matching point set, wherein denoising and fusing the first initial atlas into the first pose image includes: S21: performing tile splitting on the first initial atlas to obtain a first tile group set; S22: selecting first image block groups from the first image block group set one by one as target first image block groups, performing wavelet transform on the target first image block groups to obtain target first wavelet coefficient sets; S23: Select the first wavelet coefficient groups in the target first wavelet coefficient group set one by one as the target first wavelet coefficient group, and calculate the fuzziness corresponding to the target first wavelet coefficient group using the following block fuzziness formula: Among them, S refers to the fuzziness, k refers to the layer number of the first target wavelet coefficient group, K refers to the total number of decomposition layers of the first target wavelet coefficient group, h refers to the row number, N k It refers to the total number of rows after the k-th wavelet decomposition in the first wavelet coefficient group of the target, j refers to the column number, M k Refers to the total number of columns after the k-th wavelet decomposition in the first wavelet coefficient group of the target, w hjk refers to the amplitude value of the wavelet coefficient of the k-th layer, h-th row, and j-th column in the target first wavelet coefficient group, P refers to the total number of pixel rows of the first image block corresponding to the target first wavelet coefficient group, p refers to the row number, Q refers to the total number of pixel columns of the first image block corresponding to the target first wavelet coefficient group, q refers to the column number, and w pqk refers to the amplitude value of the wavelet coefficient in the k-th layer, p-th row, and q-th column in the target first wavelet coefficient group; S24: Gathering all ambiguities corresponding to the target first wavelet coefficient group into a target ambiguity group, and selecting the ambiguity with the smallest value from the target ambiguity group as the target ambiguity; S25: taking the first image block corresponding to the target blur in the target first image block group as the target first image block, and splicing all the target first image blocks in the first image block group into a first pose image; S3: Acquire a first real-time pose of the first camera, calculate a second real-time pose of the second camera based on the first real-time pose, generate camera extrinsics of the second camera based on the second real-time pose, and perform spatial system transformation on the feature matching point set based on the camera extrinsics to obtain a standard matching point set; S4: performing three-dimensional reconstruction on the first pose image and the second pose image according to the standard matching point set to obtain a three-dimensional environment model, performing object segmentation on the three-dimensional environment model to obtain an environment object model set, performing semantic recognition on the environment object model set to obtain an environment object semantic set, and performing semantic recognition on the environment object model set to obtain an environment object semantic set; S5: Use the environmental object semantic set to perform motion annotation on the environmental three-dimensional model to obtain a standard environmental model, extract the target motion point and the target three-dimensional coordinates of the target motion point from the standard environmental model, and control the robot arm to perform assembly movement according to the target three-dimensional coordinates.

2. The visual motion control method according to claim 1, wherein: The performing motion denoising on the second initial atlas according to the real-time motion parameters to obtain a second motion atlas includes: establishing a real-time motion model corresponding to the second camera according to the real-time motion parameters; Selecting pictures in the second initial atlas one by one as target second initial pictures, obtaining exposure times corresponding to the target second initial pictures, and generating pixel movement trajectories of respective pixels in the target second initial pictures according to the exposure times and the real-time motion model; Performing an interpolation operation on all pixel movement trajectories to obtain a picture motion model of the target second initial picture; performing deconvolution filtering on the target second initial image according to the image motion model to obtain a target second deconvolution image; Grayscale enhancement and image sharpening operations are sequentially performed on the target second deconvolution image to obtain a target second motion image, and all target second motion images are aggregated into a second motion atlas.

3. The visual motion control method according to claim 2, wherein: The step of establishing a real-time motion model corresponding to the second camera according to the real-time motion parameters includes: Obtain the length parameters of each joint of the robotic arm, and use the following forward motion formula to calculate the rotation matrix of the robotic arm based on the length parameters and the angle parameters in the real-time motion parameters: Where T is the rotation matrix, i is the i-th joint, e is the total number of joints of the robot arm, cos is the cosine function, sin is the sine function, θ i Refers to the offset angle of the i-th joint in the angle parameter, α i Indicates the angle of rotation of the i-th joint in the angle parameter around the forward direction axis of the i-1-th joint coordinate system, d i represents the angle of rotation of the i-th joint in the angle parameter around the vertical axis of the i-1-th joint coordinate system; generating a Jacobian matrix of the robotic arm according to the rotation matrix and the angle parameter; A real-time motion model of the second camera is generated according to the Jacobian matrix and the real-time motion parameters.

4. The visual motion control method according to claim 2, wherein: The performing deconvolution filtering on the target second initial picture according to the picture motion model to obtain a target second deconvolution picture includes: Performing frequency domain conversion on the target second initial image to obtain a target second image frequency domain; The target second image frequency domain is deconvolved and filtered using the following deconvolution filtering algorithm and the image motion model to obtain a target second deconvolution frequency domain: in, It refers to the amplitude of the target second deconvolution frequency domain at the frequency (u, v) at the tth moment, where u refers to the frequency in the horizontal direction, v refers to the frequency in the vertical direction, and t refers to the time of the target second deconvolution frequency domain. * (u, v, t) is the amplitude of the amplitude deconvolution filter at the frequency (u, v) at the tth moment, * is the convolution operator, π refers to pi, σ refers to the standard deviation of the Gaussian distribution, e refers to the Euler number, and x t refers to the position of the image motion model at time t, x0 refers to the initial position of the image motion model, ∈ is a preset constant, and F(u,v,t) refers to the amplitude of the target second image frequency domain at the frequency (u,v) at time t; Performing image conversion on the target second deconvolution frequency domain to obtain a target second deconvolution image.

5. The visual motion control method according to claim 1, wherein: The extracting the first pose feature from the first pose image includes: Performing multi-layer Gaussian filtering on the first pose image to obtain a first Gaussian image group; performing an image difference operation on two adjacent first Gaussian images in the first Gaussian image group one by one to obtain a first differential feature group; Performing extreme value filtering on the first differential feature group to obtain a first extreme value feature set; Performing feature point positioning on each first extreme value feature in the first extreme value feature set to obtain a first initial feature point set; Screening out low-contrast feature points and edge feature points from the first initial feature point set to obtain a first standard feature point set; performing direction assignment on each first standard feature point in the first standard feature point set to obtain a first direction feature point set; Perform feature description on each first direction feature point in the first direction feature point set one by one to obtain first description feature points, and aggregate all the first description feature points into a first pose feature.

6. The visual motion control method according to claim 5, wherein: The step of assigning directions to each first standard feature point in the first standard feature point set to obtain a first directional feature point set includes: Selecting first standard feature points in the first standard feature point set one by one as target first standard feature points, and generating target feature areas of the target first standard feature points; Dividing the target feature area into an angle sub-area set, and generating a gradient histogram of the target first standard feature point according to the angle sub-area set; Taking the maximum gradient in the gradient histogram as the feature direction of the target first standard feature point to obtain the target first direction feature point; All target first direction feature points are gathered into a first direction feature point set.

7. The visual motion control method according to claim 1, wherein: The performing feature matching on the first posture feature and the second posture feature to obtain a feature matching point set includes: Selecting first description feature points in the first posture feature one by one as target first description feature points, and calculating feature distances between the target first description feature points and each second description feature point in the second posture feature; Gathering second description feature points in the second posture feature whose feature distance is less than a preset first distance threshold into a matching second feature point set; Screening out noise feature points from the matching second feature point set to obtain a standard matching second feature point set; The feature matching points of the target first description feature points are screened out from the standard matching second feature point set using a nearest neighbor algorithm, and all the feature matching points are aggregated into a feature matching point set.

8. The visual motion control method according to claim 1, wherein: The step of performing a spatial transformation on the feature matching point set according to the camera extrinsic parameters to obtain a standard matching point set includes: Selecting feature matching points in the feature matching point set one by one as target feature matching points, and performing homogeneous coordinate transformation on the target feature matching points to obtain homogeneous coordinates of the target points; Back-projecting the homogeneous coordinates of the target point according to the intrinsic parameter matrix of the second camera to obtain the target camera coordinate point; Performing world coordinate projection on the target camera coordinate point according to the camera extrinsic parameters to obtain a target primary matching point; Performing an inverse homogeneous transformation on the target primary matching points to obtain standard matching points, and gathering all the standard matching points into a standard matching point set.

9. The visual motion control method according to claim 1, wherein: The three-dimensional reconstruction of the first pose image and the second pose image according to the standard matching point set to obtain a three-dimensional model of the environment includes: Select the standard matching points in the standard matching point set one by one as the target standard matching point, use the pixel point corresponding to the target standard matching point in the first pose image as the first pose matching point, and use the pixel point corresponding to the target standard matching point in the second pose image as the second pose matching point. Performing point cloud conversion on the target standard matching points according to the first pose matching points and the second pose matching points using a triangulation algorithm to obtain a target environment point cloud; All target environment point clouds are combined into an environment point cloud set, and the environment point cloud set is subjected to point cloud filtering to obtain a filtered point cloud set; The first pose image and the second pose image are used to perform surface reconstruction on the filtered point cloud set to obtain a three-dimensional model of the environment.

10. A visual motion control device, characterized in that: The device comprises: An atlas acquisition module is configured to acquire a first initial atlas of the robotic arm in real time using a first camera fixed to a motion base, acquire a second initial atlas in real time using a second camera fixed to the robotic arm, obtain real-time motion parameters of the robotic arm, and perform motion denoising on the second initial atlas based on the real-time motion parameters to obtain a second motion atlas; A feature matching module is used to remove noise from the first initial atlas and fuse it into a first pose image, extract a first pose feature from the first pose image, remove noise from the second motion atlas and fuse it into a second pose image, extract a second pose feature from the second pose image, and perform feature matching on the first pose feature and the second pose feature to obtain a feature matching point set, wherein the removing noise from the first initial atlas and fusing it into the first pose image includes: performing block splitting on the first initial atlas to obtain a first block group set; selecting the first block group in the first block group set one by one as the target first block group, performing wavelet transform on the target first block group to obtain a target first wavelet coefficient group set; selecting the first wavelet coefficient group in the target first wavelet coefficient group set one by one as the target first wavelet coefficient group, and calculating the blur corresponding to the target first wavelet coefficient group using the following block blur formula: Among them, S refers to the fuzziness, k refers to the layer number of the first target wavelet coefficient group, K refers to the total number of decomposition layers of the first target wavelet coefficient group, h refers to the row number, N k It refers to the total number of rows after the k-th wavelet decomposition in the first wavelet coefficient group of the target, j refers to the column number, M k Refers to the total number of columns after the k-th wavelet decomposition in the first wavelet coefficient group of the target, w hjk refers to the amplitude value of the wavelet coefficient of the k-th layer, h-th row, and j-th column in the target first wavelet coefficient group, P refers to the total number of pixel rows of the first image block corresponding to the target first wavelet coefficient group, p refers to the row number, Q refers to the total number of pixel columns of the first image block corresponding to the target first wavelet coefficient group, q refers to the column number, and w pqk refers to the amplitude value of the wavelet coefficient in the kth layer, pth row, and qth column of the target first wavelet coefficient group; all ambiguities corresponding to the target first wavelet coefficient group are aggregated into a target ambiguity group, and the ambiguity with the smallest value is screened out from the target ambiguity group as the target ambiguity; the first image block corresponding to the target ambiguity in the target first image block group is used as the target first image block, and all the target first image blocks in the first image block group are spliced ​​into a first pose image; a pose matching module, configured to obtain a first real-time pose of the first camera, calculate a second real-time pose of the second camera based on the first real-time pose, generate camera extrinsics of the second camera based on the second real-time pose, and perform spatial system transformation on the feature matching point set based on the camera extrinsics to obtain a standard matching point set; a semantic recognition module, configured to perform three-dimensional reconstruction on the first pose image and the second pose image according to the standard matching point set to obtain a three-dimensional model of the environment, perform object segmentation on the three-dimensional model of the environment to obtain a set of environment object models, perform semantic recognition on the set of environment object models to obtain a semantic set of environment objects, and perform semantic recognition on the set of environment object models to obtain a semantic set of environment objects; The motion control module is used to use the environmental object semantic set to perform motion annotation on the environmental three-dimensional model to obtain a standard environmental model, extract the target motion point and the target three-dimensional coordinates of the target motion point from the standard environmental model, and control the robot arm to perform assembly motion according to the target three-dimensional coordinates.

Citation Information

Patent Citations

  • Three-dimensional measurement method, device and equipment based on binocular imaging and storage medium

    CN115880448A

  • Mechanical arm autonomous mobile grabbing method under complex illumination conditions based on visual-tactile fusion

    WO2023056670A1