Mechanical arm control method and system based on voice control and multi-source sensing

Through the method based on voice control and multi-source sensing, the complex and cumbersome problem of robot control in the existing technology is solved, simple operation and efficient control of the robot arm are realized, technical threshold is lowered, and the use of the machine is promoted by the people.

CN120080316APending Publication Date: 2025-06-03SOUTH CHINA NORMAL UNIV

Patent Information

Application Number
CN202510203046.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing robot control methods require manual programming, which is complex and cumbersome, with high technical thresholds, making it difficult to achieve the use of machines by the people.

Method used

The robot arm control method based on voice manipulation and multi-source sensing is adopted to generate target object keywords and action keywords through speech recognition, and combine the depth camera and visual haptic sensor to realize the automatic control of the robot arm.

Benefits of technology

It realizes simple operation of the robotic arm, lowers the technical threshold of users, improves control efficiency and accuracy, and promotes the use of machines by the people.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120080316A_ABST
    Figure CN120080316A_ABST
Patent Text Reader

Abstract

The invention relates to the field of robot control, in particular to a mechanical arm control method and system based on voice control and multi-source sensing, and the method comprises the following steps: S1, recognizing a voice operation instruction, and generating a target object keyword and an action keyword; s2, performing analysis and target identification on the environment image to obtain depth information of each pixel point in the image and an identification frame of a target object corresponding to the target object keyword; s3, depth information of each pixel point in the recognition frame of the target object is extracted, and the position of the target object relative to the mechanical arm and the clamping posture of the mechanical arm are calculated; and S4, according to the position of the target object relative to the mechanical arm, the mechanical arm is controlled to move to the target position and clamp the target object in the clamping posture, and then the clamping force is calculated and adjusted according to the touch image obtained by the visual touch sensor. And S5, judging an action needing to be executed on the target object, and calling the custom performance function to implement a custom action on the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot control, and particularly to a manipulator control method and system based on voice control and multi-source sensing. Background Art

[0002] With the rapid development of industrial automation and intelligent robot technology, how to control robots more conveniently has become an important requirement. In daily use or industrial production, robots usually need to receive a series of instruction codes to perform a series of actions to complete a specified task.

[0003] An existing method for controlling a robot is achieved by directly programming the robot. This control method requires manual pre-acquisition of the current position and target position of the robot, and then programming the movement of the robot according to the position information, and executing the specified actions through the instruction codes set manually. This manual programming control method is complex and cumbersome, requires a long time to calibrate and program the robot, and has high requirements for the professional technical level of the operator, with a high technical threshold, which is not conducive to the civilian use of robots. Summary of the Invention

[0004] Based on this, the purpose of the present invention is to overcome the defects or deficiencies of the prior art, and provide a manipulator control method and system based on voice control and multi-source sensing.

[0005] A manipulator control method based on voice control and multi-source sensing includes the following steps:

[0006] S1: Identify and parse voice operation instructions to generate a target object keyword and an action keyword;

[0007] S2: Control a depth camera to capture an environmental image;

[0008] Parse the environmental image to obtain the depth information of each pixel point in the image; and

[0009] Perform target recognition on the environmental image according to the target object keyword to obtain the recognition frame of the target object corresponding to the target object keyword;

[0010] S3: Extract the depth information of each pixel point in the recognition frame of the target object, and then calculate the position of the target object relative to the manipulator and the clamping posture of the manipulator;

[0011] S4: Control the manipulator to move to the target position and clamp the target object in the clamping posture according to the position of the target object relative to the manipulator and the clamping posture of the manipulator, and then calculate and adjust the clamping force according to the tactile image obtained by the visual tactile sensor.

[0012] Through voice control and multi-source sensing of vision and hearing, the robotic arm can understand the instructions given by the user, and can analyze the visual images captured by the depth camera to obtain the position information of the target object. Thus, the robotic arm can be moved to the target position and calculate the corresponding gripping posture when the robotic arm grips the target object. Then, through step S4, the gripping force when gripping the object is identified and the magnitude of the force is controlled, enabling the robotic arm to successfully pick up the object, which has the advantages of simple operation and low technical threshold for users.

[0013] Furthermore, it further includes step S5: judging the action to be performed on the target object according to the action keywords, and calling the custom function from the custom function library to perform custom actions on the target object. The custom function can be customized by the user or installed from the existing function library. By calling the custom function, after the robotic arm successfully grips the object, different operations can be performed on the object to meet different user needs.

[0014] Furthermore, in the said step S3:

[0015] Calculating the position of the target object relative to the robotic arm includes the steps: calculating the three-dimensional coordinate position information of the target object in the depth camera coordinate system according to the depth information of each pixel point in the recognition frame of the target object; then converting the three-dimensional coordinate position information in the depth camera coordinate system to the three-dimensional coordinate system of the robotic arm;

[0016] Calculating the gripping posture of the robotic arm includes the steps;

[0017] Inputting the recognition frame of the target object into the SAM 2 model as the mask of the target object to obtain the mask area of the target object;

[0018] Obtaining the coordinate point data set of the mask area of the target object in the depth camera through the depth camera and converting it into the three-dimensional coordinate point cloud data set of the robotic arm;

[0019] Calculating the main axis direction and the center point of the target object through the three-dimensional coordinate point cloud data set of the target object of the robotic arm, and taking the center point of the target object perpendicular to the main axis direction as the gripping point of the robotic arm.

[0020] Transforming the coordinate points of the object obtained in the depth camera into the three-dimensional coordinates of the robotic arm, so that the robotic arm can understand the current position of the target object relative to the robotic arm, facilitating the movement of the robotic arm to the target position. At the same time, calculating the main axis direction and the center point of the target object through the coordinate point cloud of the target object, enabling the robotic arm to obtain the gripping point of the target object, so as to successfully grip the object.

[0021] Further, in step S4, the magnitude of the clamping force is calculated in real time according to the tactile image obtained by the visual tactile sensor. The specific steps are as follows:

[0022] Preprocess the tactile image and convert it into a format suitable for neural network processing;

[0023] Convert the preprocessed tactile image into three feature maps of different sizes;

[0024] Fuse the three feature maps of different sizes to extract richer feature information and obtain the feature map after feature fusion;

[0025] Analyze the feature map after feature fusion to obtain the normal force information and tangential force information when the robotic arm clamps the object.

[0026] The loss function adopted by the neural network model is the binary cross-entropy loss function and the smooth L1 loss function. Its overall expression is:

[0027]

[0028] Among them, BCE represents the binary cross-entropy loss, which is used to calculate the loss between the positive sample probability and the real situation in the segmented area; SmoothL1 represents the smooth L1 loss, which is used to reduce the sensitivity to outliers when calculating the clamping force; B is the total number of samples in the batch; cls pre ,cls gt represents the predicted value and the real value of the object category; f pre ,f gt represents the predicted value and the real value of the normal force; t pre ,t gt represents the predicted value and the real value of the tangential force; box pre ,box gt represents the coordinates of the predicted bounding box and the real bounding box; obj pre ,obj gt represents the predicted value and the real value of the target existence probability; IoU represents the intersection over union, which is used to evaluate the overlapping degree of the predicted box and the real box;

[0029] Further, when controlling the robotic arm to clamp the target object in step S4, first clamp the target object with a relatively small force to make the target object move a certain distance away from its support to obtain the friction force time series data after the target object completely leaves its support during the moving process; calculate the variance and the average value of the friction force time series data. If the ratio of the friction force variance to the average value of the friction force is greater than a certain value, it is determined that the clamping force is too small, and the robotic arm is controlled to increase the clamping force and clamp again; if the ratio of the friction force variance to the average value of the friction force is less than a certain value, it is determined that the clamping is successful; among them, the friction force is the tangential force obtained by the visual tactile sensor.

[0030] Further, in step S2, a pre-trained YOLO-World model is used to perform object recognition on the environmental image, and a CLIP model is used to optimize the object keywords input into the YOLO-World model.

[0031] Most of the objects existing in nature can be recognized by the pre-trained object recognition model, and their recognition frames are marked in the image. At the same time, in order to improve its understanding ability in specific scenarios, the CLIP model is used to optimize its input, so as to optimize the fusion ability of image features and text features in specific scenarios without changing the data set. Obtaining the depth information of the image enables the image captured by the depth camera to have not only two-dimensional information in the plane coordinate system, but also depth information, which is convenient for positioning the target object in the scene.

[0032] Further, step S1 specifically includes the steps:

[0033] S11: Recognize the voice operation instruction received by the sound sensor and generate a text statement;

[0034] S12: Perform semantic recognition on the text statement to generate object keywords and action keywords.

[0035] Further,

[0036] After step S12, the following steps are also executed: Generate a response statement for the start of the task and convert it into voice for playback;

[0037] When performing object recognition on the environmental image according to the object keywords in step S2, if the target object is not recognized, the following steps are executed: Generate a response statement for recognition failure and convert it into voice for playback;

[0038] After step S5, the following steps are also executed: Generate a response statement for successful task execution and convert it into voice for playback.

[0039] After the user issues an instruction to the robotic arm, the robotic arm will reply to the user to confirm receiving the instruction and start executing the task. When performing object recognition, if the target object is not found, a response of recognition failure will be made to the user. When the task is completed, a response of task completion will be made to the user to improve the user experience.

[0040] A control device for a robotic arm based on voice control and multi-source sensing, including a voice processor, an environmental image processor, a robotic arm motion calculator, a gripping force calculator, a custom function library, and a custom action executor;

[0041] The voice processor is used to recognize and parse voice operation instructions and generate object keywords and action keywords;

[0042] An environmental image processor for controlling a depth camera to capture an environmental image;

[0043] Parse the environmental image to obtain the depth information of each pixel point in the image; and

[0044] Perform object recognition on the environmental image according to the target object keyword to obtain the recognition frame of the target object corresponding to the target object keyword;

[0045] A robotic arm motion calculator: used to extract the depth information of each pixel point in the recognition frame of the target object, and then calculate the position of the target object relative to the robotic arm and the gripping posture of the robotic arm;

[0046] A gripping force calculator for controlling the robotic arm to move to the target position and grip the target object in the gripping posture according to the position of the target object relative to the robotic arm and the gripping posture of the robotic arm, and then calculate and adjust the magnitude of the gripping force in real time according to the tactile image obtained by the visual-tactile sensor;

[0047] A custom function library: used to store custom functions for controlling the robotic arm to execute actions;

[0048] A custom action executor for judging the action to be performed on the target object according to the action keyword, and calling the custom function from the custom function library to perform a custom action on the target object.

[0049] A robotic arm control system based on voice control and multi-source sensing, including a robotic arm, a visual-tactile sensor arranged at the gripping end of the robotic arm, a depth camera arranged on the robotic arm, a sound sensor, and a control device; the sound sensor receives the instruction of voice control; the depth camera captures an environmental image; the visual-tactile sensor contacts the target object and generates a tactile image when the robotic arm grabs the target object; the control device receives the instruction of voice control obtained by the sound sensor, the environmental image captured by the depth camera, and the tactile image of the visual-tactile sensor, executes the above-mentioned robotic arm control method based on voice control and multi-source sensing, and controls the robotic arm to move and grip the object.

[0050] Further, the visual-tactile sensor includes a housing, a flexible film, a transparent acrylic plate, a camera, an LED light, and a light homogenizing plate. The housing is a semi-enclosed box body with opposite openings. The transparent acrylic plate is covered on one of the openings of the housing. The flexible film is attached to the transparent acrylic plate and can contact the target object. The camera and the LED light are arranged on the other opening of the housing, and the light homogenizing plate is arranged inside the housing. The light emitted by the LED light passes through the light homogenizing plate and the transparent acrylic plate and irradiates on the flexible film. When the flexible film contacts an object, it deforms and reflects back light with gradient changes, and the camera can capture the light reflected back by the flexible film;

[0051] The flexible film is a self-prepared product. It is obtained by mixing liquid A (vinyl-terminated polydimethylsiloxane) and liquid B (hydrogen-terminated polydimethylsiloxane) of polydimethylsiloxane (PDMS) in a ratio of 10:1, putting it into a vacuum plasma surface treatment instrument to pump to vacuum and laying it flat on a mold, and then placing it in an incubator for 24 hours. And a metal is plated on the surface of the flexible film through a magnetron sputtering instrument to better collect tactile images.

[0052] Further, a mechanical clamp for clamping an object is further arranged at the end of the robotic arm. The mechanical clamp includes a housing, a driving servo, a link transmission mechanism, and a clamping end. One end of the housing is connected to the end of the robotic arm, and the other end is provided with the clamping end. A hollow accommodating part is formed inside the housing for placing the control device. The driving servo and the link transmission mechanism are arranged on the housing. The link transmission mechanism includes a first link, a second link, and a third link. The first link is fixed on the driving servo and can be driven by the driving servo to rotate. Two ends of the first link are respectively hinged to one ends of the second link and the third link. The other ends of the second link and the third link are respectively hinged to two clamping arms of the clamping end. The driving servo can drive the link transmission mechanism to drive the two clamping arms of the clamping end to close or separate. The visual-tactile sensor is arranged at the clamping end of the mechanical clamp and can contact the object.

[0053] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned robotic arm control method based on voice control and multi-source sensing.

[0054] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Description of the Drawings

[0055] Figure 1 It is a structural diagram of the mechanical clamp;

[0056] Figure 2 It is a structural diagram of a visual-tactile sensor;

[0057] Figure 3 It is a schematic structural diagram of a control device;

[0058] Figure 4 It is a flowchart of a manipulator control method based on voice control and multi-source sensing;

[0059] Figure 5 It is a broken line graph of the friction force change in case of grasping failure and successful grasping;

[0060] Figure 6 It is a flowchart of a manipulator control method based on voice control and multi-source sensing in another embodiment. Detailed implementation manners

[0061] The manipulator control system based on voice control and multi-source sensing of the present invention includes a manipulator, a sound sensor, a depth camera, a visual-tactile sensor and a control device. The end of the manipulator includes a mechanical gripper for gripping an object. The sound sensor and the depth camera are arranged on the manipulator, and the visual-tactile sensor is arranged at the gripping end of the mechanical gripper. The manipulator can move freely. The sound sensor can receive voice operation instructions input by a user. The depth camera can capture environmental images of the outside world. The visual-tactile sensor can capture tactile images when the manipulator grips an object. The control device can control the manipulator to perform corresponding actions according to the voice operation instructions input by the user, the environmental images captured by the depth camera, and the tactile images captured by the visual-tactile sensor. The manipulator control system based on voice control and multi-source sensing of the present invention has the advantages of simple operation and low technical threshold.

[0062] Please refer to Figure 1 , the mechanical gripper includes a housing 310, a driving servo 320, a link transmission mechanism 330 and a gripping end 340. One end of the housing 310 is connected to the end of the manipulator, and the other end is provided with the gripping end 340. An empty accommodating part is formed in the upper part of the housing 310 for placing the control device. The driving servo 320 and the link transmission mechanism 330 are arranged on the housing 310. The driving servo 320 can drive the link transmission mechanism 330 to drive the two gripping arms of the gripping end 340 to close or separate. The visual-tactile sensor is arranged on the gripping end 340 of the mechanical gripper and can come into contact with an object.

[0063] Specifically, the link transmission mechanism 330 includes a first link 332, a second link 334, and a third link 336. The first link 332 is fixed to the driving servo 320 and is driven by the driving servo 320 to rotate. Two ends of the first link 332 are respectively hinged to one ends of the second link 334 and the third link 336. The other ends of the second link 334 and the third link 336 are respectively hinged to two clamping arms of the clamping end 340. When the driving servo 320 rotates clockwise, it drives the first link 332 to rotate clockwise, so that the second link 334 and the third link 336 hinged to the first link 332 drive the two clamping arms of the clamping end 340 to close. Similarly, when the driving servo 320 rotates counterclockwise, it will drive the two clamping arms of the clamping end 340 to separate.

[0064] Please refer to Figure 2 , the visual tactile sensor includes a housing 201, a flexible film 202, a transparent acrylic plate 203, a camera 204, an LED lamp 205, and a light homogenizing plate 206. The housing 201 is a semi-enclosed box body with opposite openings. The transparent acrylic plate 203 is covered on one of the openings of the housing 201. The flexible film 202 is attached to the transparent acrylic plate 203 and can contact the target object. The camera 204 and the LED lamp 205 are arranged on the other opening of the housing 201. The light homogenizing plate 206 is arranged inside the housing 201. The light emitted by the LED lamp 205 passes through the light homogenizing plate 206 and the transparent acrylic plate 203 and irradiates on the flexible film 202. When the flexible film 202 contacts an object and deforms, it will reflect back light with a gradient change, and the camera 204 can capture the light reflected back by the flexible film 202. In the present invention, the flexible film 202 is a self-prepared product. First, the A liquid (vinyl-terminated polydimethylsiloxane) and the B liquid (hydrogen-terminated polydimethylsiloxane) of polydimethylsiloxane (PDMS) are mixed in a ratio of 10:1 and pumped to a vacuum in a vacuum plasma surface treatment instrument, then laid flat on a mold and placed in an incubator for 24 hours to obtain the flexible film. Finally, a metal is plated on the surface of the flexible film by a magnetron sputtering instrument to better collect tactile images.

[0065] Specifically, please refer to Figure 3 and Figure 4 where Figure 3 is a schematic structural diagram of the control device, Figure 4 is a flowchart of a robotic arm control method based on voice control and multi-source sensing. The control device includes a voice processor, an environmental image processor, a robotic arm motion calculator, a grasping force calculator, a custom database, and a custom action executor.

[0066] When the operator wants to control the robotic arm, a control command can be sent to the system by voice. The voice processor is used to execute step S1: identify and parse the voice operation command, and generate a target object keyword and an action keyword.

[0067] The voice processor includes a voice recognition module 112 and a semantic recognition module 114, and step S1 includes steps S11 and S12.

[0068] The voice recognition module 112 is used to execute step S11: recognize the voice operation command received by the sound sensor and generate a text statement. Specifically, the voice signal is recognized by the whisper-large-v3-turbo model. By calling the interface provided by the model, the received voice can be converted into a text statement.

[0069] The semantic recognition module 114 is used to execute step S12: perform semantic recognition on the text statement and generate a target object keyword and an action keyword.

[0070] Specifically, semantic recognition is performed through the multi-modal large language model Qwen2-VL 72B Instruct model. A call request is sent to the API of the Qwen2-VL 72B Instruct model, and the text statement obtained through the voice recognition model in step S1 is used as a parameter to input into the multi-modal large language model. The multi-modal large language model can automatically extract the target object keyword and the action keyword in the text statement, and judge the action that the operator needs to perform on the target object in the statement.

[0071] Preferably, in order to enable the large language model to have a certain memory function, the context generated by the conversation is stored, and the context is passed in together each time it is called.

[0072] When an object needs to be clamped, the environmental image processor is used to execute step S2: control the depth camera to capture an environmental image; parse the environmental image to obtain the depth information of each pixel point in the image; and perform target recognition on the environmental image according to the target object keyword to obtain the recognition frame of the target object corresponding to the target object keyword.

[0073] Specifically, the environmental image processor includes a target recognition module 122 and a depth calculation module 124, and step S2 further includes step S2A and step S2B.

[0074] The target recognition module 122 is used to execute step S2A: perform target recognition on the environmental image and judge whether the target object is recognized; if the target object is not recognized, the action is stopped; if the target object is recognized, the recognition frame of the target object is drawn in the environmental image.

[0075] Specifically, the YOLO-World model is used to perform object recognition on the environmental images captured by the depth camera. This model can detect any object in the image in real time, so as to mark the objects corresponding to the keywords in the image in the form of recognition frames. And the object recognition model is a pre-trained model, which can itself recognize most of the objects existing in nature and can be used without further training. At the same time, when the object recognition model recognizes the image, since there may be multiple objects meeting the requirements in the image, only the object with the highest confidence level considered by the object recognition model is used as the target object corresponding to the keyword and a recognition frame is generated.

[0076] Furthermore, in this embodiment, in order to improve the understanding ability of the object recognition model in a specific scenario, a method for tuning through the CLIP model is provided. The CLIP model is a model that can map images and texts to the same feature space, so that without changing the data set, the fusion ability of image features and text features in a specific scenario can be optimized only by tuning the input keywords. By performing semantic analysis on the objects in a specific scenario in advance and extracting their semantic feature vectors, the recognition success rate in the specific scenario can be improved. After loading the English text library of items in a specific application scenario, the category names of the keywords are converted into tokens. Tokens are the basic units in the processing of the CLIP model for texts; the CLIP model is used to generate text embeddings. Embeddings are a way of representing texts as vectors in a low-dimensional vector space, and these vectors contain the semantic information of the texts. Then the embeddings are converted into a numpy array and saved in a file in the.npy format. The npy array file is pre-loaded in YOLO-World, so that the model can utilize these supplementary semantic feature vectors during operation, thereby optimizing the fusion of image features and text features in a specific scenario and improving the recognition ability of objects in the specific scenario.

[0077] The depth calculation module 124 is used to execute step S2B: parse the environmental image to obtain the depth information of each pixel point in the image. Specifically, the depth camera simulates the visual principle of human binoculars, uses two or more cameras to capture the same scene from different angles to obtain multiple images. Since the imaging of the same object by cameras at different positions has a parallax, by calculating this parallax through the triangulation principle, the depth information of the pixel points can be determined.

[0078] Furthermore, the robotic arm motion calculator is used to execute step S3: extract the depth information of each pixel point in the recognition frame of the target object, and then calculate the position of the target object relative to the robotic arm and the gripping posture of the robotic arm.

[0079] The manipulator motion calculator includes a target position calculation module 132 and a gripping posture calculation module 134, and step S3 further includes steps S3A and S3B.

[0080] The target position calculation module 132 is used to execute step S3A: According to the depth information of each pixel point in the recognition frame of the target object, calculate the position information of the target object and convert it to the three-dimensional coordinate system of the manipulator. Specifically, in order to convert the coordinate system of the target object in the depth image to the three-dimensional coordinate system of the manipulator, it is necessary to perform hand-eye calibration on the depth camera and the manipulator using a black and white grid calibration board, including the following steps:

[0081] SA: Take at least three RGB images of the black and white grid calibration board at different positions through the depth camera, select a certain grid point of the calibration board, and obtain the depth distance information depth of the corresponding grid on the depth image, so as to obtain the coordinate point (pix_x, pix_y, depth) of the grid point in the depth camera.

[0082] SB: Convert the coordinate point (pix_x, pix_y, depth) in the depth camera to a three-dimensional coordinate point (cam_x, cam_y, cam_z) with the depth camera as the origin of the three-dimensional coordinate system. Specifically, first obtain the internal parameter matrix K of the depth camera. The internal parameter matrix K is the parameter of the depth camera itself and can be obtained from the camera's instruction manual. Through the internal parameter matrix K and the coordinate point (pix_x, pix_y, depth) in the depth camera, the three-dimensional coordinate point (cam_x, cam_y, cam_z) with the depth camera as the origin of the three-dimensional coordinate system can be calculated, and its calculation expression is as follows:

[0083]

[0084] SC: Move the execution end of the manipulator to the selected grid point and obtain the manipulator coordinate point (arm_x, arm_y, arm_z) corresponding to each grid point.

[0085] SD: Obtain the transformation matrix T for converting the three-dimensional coordinates of the depth camera to the three-dimensional coordinates of the manipulator. Specifically, the camera coordinate point A is converted to the manipulator coordinate point B, which should satisfy B = T * A. After obtaining no less than three groups of A and the corresponding B, the least squares method can be used to solve for T.

[0086] The gripping posture calculation module 134 is used to execute step S3B: According to the depth information of each pixel point in the recognition frame of the target object, calculate the gripping posture of the manipulator.

[0087] Specifically, when the robotic arm moves to the position of the target object, for objects such as wrenches and pencils, successful grasping can only be achieved when grasping perpendicular to the spindle direction. Therefore, it is necessary to estimate the pose of the grasped object in order to grasp the object normally. In this embodiment, the SAM2 model is used to calculate the grasping pose of the robotic arm, which includes the following steps:

[0088] S3B1: Input the recognition frame of the target object into the SAM 2 model as the mask of the target object to obtain the mask region of the target object.

[0089] S3B2: Obtain the coordinate point dataset of the mask region of the target object in the depth camera through the depth camera, and convert it into the 3D coordinate point cloud dataset of the robotic arm.

[0090] S3B3: Calculate the spindle direction and the center point of the target object through the 3D coordinate point cloud dataset of the target object on the robotic arm, and use the center point of the target object perpendicular to the spindle direction as the grasping point of the robotic arm.

[0091] Specifically, substitute the 3D coordinate point cloud dataset of the robotic arm into the following formula to obtain the covariance matrix, where C is the covariance matrix, n is the number of points in the point cloud, p i is the coordinate (x, y, z) of the i-th point, μ is the average coordinate of the point cloud (μ x , μ y , μ z ), and T represents the transpose. Perform eigenvalue decomposition on the covariance matrix C to obtain three eigenvalues (λ 1 , λ 2 , λ 3 ) and the corresponding eigenvectors (v 1 , v 2 , v 3 ). The eigenvalue represents the variance magnitude along the direction of the corresponding eigenvector, and the eigenvector represents the direction of the spindle axis. The eigenvector corresponding to the largest eigenvalue is the spindle direction of the point cloud. This vector points in the direction where the object extends the longest. For example, if λ 1 > λ 2 > λ 3 , then v 1 is the spindle direction of the target object. Control the robotic arm to grasp the object in a direction perpendicular to the spindle axis. At the same time, calculate the center point of the object through the 3D coordinate point cloud dataset of the robotic arm, and set the grasping point as the center point of the object.

[0092] Furthermore, the grasping force calculator 140 is used to execute step S4: Control the robotic arm to move to the target position and grasp the target object in this grasping pose according to the position of the target object relative to the robotic arm and the grasping pose of the robotic arm, and then calculate and adjust the magnitude of the grasping force according to the tactile image obtained by the visual tactile sensor;

[0093] When picking up different objects, due to the different materials and surface properties of the objects, it is necessary to ensure that the clamping force of the robotic arm can pick up the objects smoothly. When the robotic arm picks up an object, a clamping force perpendicular to the clamping plane and a frictional force parallel to the clamping plane will be generated. If the frictional force is too small, the object will slip off.

[0094] Specifically, a neural network model is used to identify the tactile image received by the visual tactile sensor and convert it into the magnitude of force. It includes an input layer, a feature extraction layer, a feature fusion layer, and a decoding layer. The input layer is used to preprocess the tactile image to convert it into a format suitable for neural network processing. The backbone network of the feature extraction layer is CSPnet, which can convert the preprocessed tactile image into feature maps of three different sizes: 80×80, 40×40, and 20×20, so as to perform block processing on 8400 regions of the tactile image. The feature fusion layer fuses the three feature maps of different sizes through a feature pyramid network to extract richer feature information and obtain a feature map after feature fusion. The decoding layer sends the feature map after feature fusion into the decoder for decoding. The decoder includes the normal force information (clamping force) and tangential force information (frictional force) when the robotic arm picks up an object.

[0095] The loss functions adopted by the neural network model are the binary cross-entropy loss function and the smooth L1 loss function, and its overall expression is:

[0096]

[0097] where BCE represents the binary cross-entropy loss, which is used to calculate the loss between the positive sample probability and the actual situation in 8400 regions; SmoothL1 represents the smooth L1 loss, which is used to reduce the sensitivity to outliers when calculating the clamping force; B is the total number of samples in the batch; cls pre ,cls gt represents the predicted value and the actual value of the object category; f pre ,f gt represents the predicted value and the actual value of the normal force; t pre ,t gt represents the predicted value and the actual value of the tangential force; box pre ,box gt represents the coordinates of the predicted bounding box and the actual bounding box; obj pre ,obj gt represents the predicted value and the actual value of the target existence probability; IoU represents the intersection over union, which is used to evaluate the overlapping degree of the predicted box and the actual box.

[0098] A sliding detection algorithm is used to control the clamping force when the robotic arm clamps an object to ensure that the robotic arm can successfully clamp the object. Please refer to Figure 4 When the robotic arm clamps an object, if the clamping force is too small, during the upward movement, the frictional force is not sufficient to balance the gravity. After the mechanical clamp moves upward, the object will slip due to insufficient frictional force, and the frictional force will continuously fluctuate when the object is about to slip; if the object can be successfully clamped, the frictional force will gradually increase as the object is moved upward, and finally remain equal to the gravity value of the object. The sliding detection algorithm first clamps the object with a relatively small force, and then controls the robotic arm to move vertically upward by 1 cm; during the movement, the magnitude of the frictional force between the robotic arm and the object is obtained in real time through the visual-tactile sensor, and the frictional force time series data is obtained; and the variance of the frictional force time series data after the object completely leaves its support is calculated. Since for different objects, the difference in their own gravity and the surface material will result in different clamping forces required for successful clamping. For objects with small gravity, slight changes in the frictional force may lead to clamping failure, while for objects with large gravity, clamping failure may only occur when the frictional force changes too much. Therefore, it is not possible to judge whether the object will slip only based on the variance of the frictional force. The gravity of the object also needs to be considered. Therefore, in this embodiment, it is set that the ratio of the frictional force variance to the average value of the frictional force (approximately equal to the gravity of the object) is greater than a certain value (0.1 in this embodiment), which means that the clamping force is insufficient and the object may slip, and then the clamping force is increased in the next pre-clamping for re-clamping; the ratio of the frictional force variance to the average value of the frictional force is less than a certain value, indicating successful clamping.

[0099] The custom function library stores custom functions for controlling the actions of the robotic arm, and the custom functions can be defined by the user or installed from other databases.

[0100] The custom action executor 150 is used to execute step S5: determine the action to be performed on the target object according to the action keyword, and call the custom function from the custom function library to perform the custom action on the target object.

[0101] Specifically, in this embodiment, the voice operation instruction input by the user in step S1 is: "Please gently screw the screw into the nut", then the voice processor extracts keywords: "screw", "nut", "gently", etc. from it, and calls the custom functions in the custom function library according to these keywords: control_arm_grab and screw_control.

[0102] The `control_arm_grab` function is responsible for controlling the robotic arm to pick up an object in the correct gripping posture and move it to the target position. Its input parameters include `start_object` and `end_object`. `start_object` represents the initial position where the robotic arm picks up the object (i.e., the position of the screw), and `end_object` represents the target position (i.e., the position of the nut). The parameters `start_object` and `end_object` are obtained through the step S3A, and the gripping posture of the robotic arm is calculated through the step S3B. The gripping force of the robotic arm is controlled through the step S4.

[0103] The `screw_control` function is responsible for controlling the robotic arm to perform the action of screwing a screw and controlling the force according to the force feedback. Its input parameters include `target_object` and `force`. `target_object` represents the target object (i.e., the nut), and `force` represents the tightening force. The `screw_control` function drives the robotic arm to rotate around the central axis of the nut to screw the screw into the nut. At the same time, the tangential force (torque) of the grip is detected in real time by the visual-tactile sensor according to the step S4. Once the torque is greater than the force represented by "gently", the twisting stops, indicating that the current screw has been tightened.

[0104] Preferably, when the voice command input by the user in the step S1 is: "Please gently screw all the screws into the corresponding nuts", the voice processor extracts keywords such as "all screws", "corresponding nuts", "gently", etc. On the basis of this embodiment, the visual processor in the step S2 will obtain the recognition frames of all screws and nuts, obtain their features, then match the corresponding screws and nuts, and then call the `screw_control` function and tighten each screw and its corresponding nut in sequence according to the "gently" force.

[0105] Preferably, in order to enable the voice processor to better understand the instructions input by the user, the voice processor can be optimized for prompt words in the custom function.

[0106] The custom function mentioned in the present invention can be programmed by the user or installed from other existing custom function libraries. In another embodiment, the voice command input by the user in step S1 is: "Please put fruits and paper balls into the plate and the trash can respectively", then the voice processor extracts keywords: "fruits", "plate", "paper ball", "trash can", etc. The visual processor in step S2 will obtain the recognition frames of all fruits, plates, paper balls, and trash cans, and then match the corresponding fruits with plates or paper balls with trash cans, and call the custom function control_arm_grab in the custom function library multiple times to put the fruits into the plate or put the paper balls into the trash can in turn.

[0107] Preferably, in another embodiment, please refer to Figure 6 , the method and system for controlling a robotic arm based on voice control and multi-source sensing of the present invention further has a voice answering function. The robotic arm control system based on voice control and multi-source sensing further includes a speaker, and the above control device further includes a statement generation module 116 and a voice conversion module 118.

[0108] After step S12 executed by the semantic recognition module 114, the statement generation module 116 executes the steps: generate an answer statement for the start of the task through the multi-modal large language model Qwen2-VL 72B Instruct, and then the voice conversion module 118 executes the steps: convert the answer statement for the start of the task into a voice that can be played and transmit it to the speaker for playback.

[0109] When the target recognition module 122 executes step S2A, if the target object is not recognized, the statement generation module 116 executes the steps: generate an answer statement for recognition failure through the multi-modal large language model Qwen2-VL 72B Instruct, and then the voice conversion module 118 executes the steps: convert the answer statement for recognition failure into a voice that can be played and transmit it to the speaker for playback. For example: when the target recognition module 122 fails to recognize the target object, the recognition failure statement "Recognition failure, the target object was not found" is generated.

[0110] After the custom action executor 150 executes step S5, it also passes through the statement generation module 116 to execute the steps: generate an answer statement for task success through the multi-modal large language model Qwen2-VL 72B Instruct, and then the voice conversion module 118 executes the steps: convert the answer statement for task success into a voice that can be played and transmit it to the speaker for playback. For example, when the task is successfully executed, the task success statement generated is: "Task completed, the paper ball has been put into the trash can".

[0111] Compared with the existing control methods for robots or robotic arms, a robotic arm control method and system based on voice control and multi-source sensing according to the present invention can control the robotic arm to move and perform tasks through simple voice input. The sound sensor receives external sound signals, and then the voice processor analyzes the sound signals to generate keywords and response voices. When an object needs to be picked up, the depth camera takes pictures of the external environment, and the environmental image processor obtains the recognition frame of the target object and the depth information of the image from the environmental image captured by the depth camera. Then, the robotic arm motion calculator extracts the depth information of each pixel point in the recognition frame of the target object, and then calculates the position of the target object relative to the robotic arm and the picking posture of the robotic arm. Next, the robotic arm is moved to the position of the target object and the object is picked up with the calculated picking posture. When gripping the object, the visual tactile sensor contacts the object and generates a tactile image to control the gripping force of the robotic arm to ensure that the object does not slip. After successfully gripping the object, a function in the custom function library can be called to make the robotic arm complete a certain fixed action. The robotic arm control system based on voice control and multi-source sensing according to the present invention has the advantages of simple operation and high operating efficiency.

[0112] Based on the same inventive concept, the present application also provides an electronic device, which can be a server, a desktop computing device or a mobile computing device (such as a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.) and other terminal devices. The device includes one or more processors and a memory, where the processor is used to execute a program to implement the embodiments of the present invention; the memory is used to store a computer program executable by the processor.

[0113] Based on the same inventive concept, the present application also provides a computer-readable storage medium, corresponding to the embodiment of a robotic arm control method based on voice control and multi-source sensing. The computer-readable storage medium stores a computer program thereon, and when the program is executed by a processor, it implements the steps of a robotic arm control method based on voice control and multi-source sensing recorded in any of the above embodiments.

[0114] The present application may be in the form of a computer program product implemented on one or more storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain program code. Computer-usable storage media include both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0115] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and the present invention is also intended to include these modifications and improvements.

Claims

1. A robotic arm control method based on voice control and multi-source sensing, characterized in that: The following steps are involved: S1: Identify and parse voice operation instructions to generate target object keywords and action keywords; S2: Control the depth camera to capture the environment image; Analyze the environment image to obtain the depth information of each pixel in the image; and Performing target recognition on the environment image according to the target object keywords, and obtaining the recognition frame of the target object corresponding to the target object keywords; S3: extract the depth information of each pixel in the recognition frame of the target object, and then calculate the position of the target object relative to the robotic arm and the gripping posture of the robotic arm; S4: According to the position of the target object relative to the robotic arm and the gripping posture of the robotic arm, the robotic arm is controlled to move to the target position and grip the target object with the gripping posture, and then the gripping force is calculated and adjusted according to the tactile image obtained by the visual tactile sensor.

2. The method for controlling a robotic arm based on voice control and multi-source sensing according to claim 1, characterized in that: The method further includes step S5: determining the action to be performed on the target object according to the action keyword, and calling the custom function from the custom function library to perform the custom action on the target object.

3. The method for controlling a robotic arm based on voice control and multi-source sensing according to claim 1 or 2, characterized in that: In step S3: Calculating the position of the target object relative to the robotic arm, including the steps of: calculating the three-dimensional coordinate position information of the target object in the depth camera coordinate system according to the depth information of each pixel point in the identification frame of the target object; Then, the three-dimensional coordinate position information in the depth camera coordinate system is converted to the three-dimensional coordinate system of the robotic arm; Calculating the gripping posture of the robot arm, including steps; Input the recognition box of the target object into the SAM 2 model as the mask of the target object to obtain the mask area of ​​the target object; The coordinate point dataset of the mask area of ​​the target object in the depth camera is obtained through the depth camera, and converted into a three-dimensional coordinate point cloud dataset of the robot arm; The main axis direction and center point of the target object are calculated through the three-dimensional coordinate point cloud data set of the robot arm of the target object, and the center point of the target object perpendicular to the main axis direction is used as the gripping point of the robot arm.

4. The method for controlling a robotic arm based on voice control and multi-source sensing according to claim 3, characterized in that: In step S4, the magnitude of the clamping force is calculated in real time according to the tactile image obtained by the visual tactile sensor, and the specific steps are: Preprocess the tactile image and convert it into a format suitable for neural network processing; Convert the preprocessed tactile image into three feature maps of different sizes; The three feature maps of different sizes are fused to extract richer feature information and obtain the feature map after feature fusion; Analyze the feature map after feature fusion to obtain the normal force information and tangential force information when the robotic arm grasps the object.

5. The method for controlling a robotic arm based on voice control and multi-source sensing according to claim 4, characterized in that: When controlling the robotic arm to grip the target object in step S4, the target object is first gripped with a relatively small force, so that the target object moves a certain distance away from its support, so as to obtain the friction time series data after the target object completely leaves its support during the movement; the variance and the average value of the friction time series data are calculated, and if the ratio of the friction variance to the average value of the friction is greater than a certain value, it is judged that the gripping force is too small, and the robotic arm is controlled to increase the gripping force and re-grip; if the ratio of the friction variance to the average value of the friction is less than a certain value, it is judged that the gripping is successful; wherein, the friction is the tangential force obtained by the visual tactile sensor.

6. The method for controlling a robotic arm based on voice control and multi-source sensing according to claim 5, characterized in that: In step S2, a pre-trained YOLO-World model is used to perform target recognition on the environment image, and the target object keywords input into the YOLO-World model are tuned through the CLIP model.

7. The method for controlling a robotic arm based on voice control and multi-source sensing according to claim 6, characterized in that: The step S1 specifically includes the following steps: S11: recognizing the voice operation instruction received by the acoustic sensor and generating a text sentence; S12: Perform semantic recognition on text sentences to generate target object keywords and action keywords.

8. The method for controlling a robotic arm based on voice control and multi-source sensing according to claim 7, characterized in that: After step S12, the following steps are further performed: generating a response statement for the start of the task and converting it into voice for playback; When the target object is recognized in the environment image according to the target object keyword in step S2, if the target object is not recognized, the following steps are executed: generating a reply sentence of recognition failure and converting it into voice playback; After step S5, the following step is further performed: generating a response statement indicating successful task execution and converting it into voice for playback.

9. A control device for a robotic arm based on voice control and multi-source sensing, characterized in that: Includes voice processor, environmental image processor, robot arm motion calculator, gripping force calculator, custom function library and custom action actuator; A voice processor is used to recognize and analyze voice operation instructions and generate target object keywords and action keywords; The environmental image processor is used to control the depth camera to capture environmental images; analyze the environmental images to obtain the depth information of each pixel in the image; and performing target recognition on the environment image according to the target object keyword to obtain a recognition frame of the target object corresponding to the target object keyword; The robot arm motion calculator is used to extract the depth information of each pixel in the recognition frame of the target object, and then calculate the position of the target object relative to the robot arm and the gripping posture of the robot arm; A gripping force calculator, used to control the robot arm to move to the target position and grip the target object in the gripping posture according to the position of the target object relative to the robot arm and the gripping posture of the robot arm, and then calculate and adjust the gripping force in real time according to the tactile image obtained by the visual tactile sensor; A custom function library is used to store custom function functions that control the robot arm to perform actions; The custom action executor is used to determine the action to be performed on the target object according to the action keyword, and call the custom function from the custom function library to implement the custom action on the target object.

10. A robotic arm control system based on voice control and multi-source sensing, characterized in that: It includes a robotic arm, a visual-tactile sensor arranged at the gripping end of the robotic arm, a depth camera arranged on the robotic arm, an acoustic sensor, and a control device; the acoustic sensor receives voice control instructions; the depth camera captures environmental images; the visual-tactile sensor contacts the target object and generates a tactile image when the robotic arm grasps the target object; the control device receives the voice control instructions obtained by the acoustic sensor, the environmental image captured by the depth camera, and the tactile image of the visual-tactile sensor, and executes the robotic arm control method based on voice control and multi-source sensing as described in any one of claims 1-8, to control the robotic arm to move and grip objects.

Citation Information

Patent Citations

  • Grabbing device and assembling equipment

    CN110978033A

  • Self-transfer system of mechanical arm based on combination of single-lens camera and double-lens camera

    CN111267083A

  • Manipulator for synchronously detecting gripping friction

    CN111546365A

  • Grabbed object recognition method based on tactile vibration signal and visual image fusion

    CN112388655A

  • Article grabbing planning method and system

    CN114102585A

Cited By

  • Robot control method and device, electronic equipment, storage medium and program product

    CN120347761A

  • Robot control method and device, electronic equipment, storage medium and program product

    CN120347761B