Multi-mode intelligent interactive mechanical arm system
By introducing multi-mode intelligent interaction technology into the robotic arm system, combining AI and multiple interaction modes, the control problem of robotic arm in complex scenarios is solved, and efficient and accurate task execution is achieved.
Patent Information
- Application Number
- CN202510350012.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-24
AI Technical Summary
When faced with complex and changing work scenarios, existing robotic arm control methods are difficult to quickly adapt to environmental changes, and when the light is dim or the hardware foundation is weak, identification errors and operational errors are prone to occur.
The multi-mode intelligent interactive robot arm system is adopted, combining artificial intelligence and robot arm to provide four interaction modes: gesture recognition, voice control, web control and eye movement following based on machine vision, reducing the threshold for using robot arm technology and broadening the application field.
It realizes efficient control of the robotic arm in complex and changing scenarios, simplifies complex task processing, reduces operation difficulty, and improves the accuracy and efficiency of task execution.
Smart Images

Figure CN120080317A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robotic arms, specifically a multi-mode intelligent interactive robotic arm system. Background Technique
[0002] Due to its unique technical characteristics, robotic arms have been widely penetrated into many core fields such as industrial manufacturing, medical services, and logistics transportation, becoming a key factor in promoting the intelligent upgrading of industries and improving people's quality of life. In the field of industrial manufacturing, robotic arms are used for high-precision assembly of parts in automobile manufacturing, such as engine assembly. Compared with manual work, they are faster and can effectively avoid errors, significantly improving production efficiency and quality.
[0003] However, in the face of complex and changeable working scenarios where the production line needs to be temporarily adjusted to grasp different parts, the traditional robotic arm control method that relies on preset programs and fixed modes is difficult to quickly adapt. Reprogramming and debugging are time-consuming and affect the progress. In the field of medical services, robotic arms assist doctors with high precision and stability in minimally invasive brain surgery, reducing the surgical risk and increasing the success rate. However, in the case of remote medical surgery, traditional robotic arms are affected by poor signals and are prone to operating errors, affecting the surgical effect.
[0004] In the field of logistics transportation, robotic arms in large e-commerce warehouses quickly sort and transport goods according to system instructions, improving operation efficiency and reducing the human burden. However, in an environment with dim light and unclear object features, traditional robotic arms are prone to recognition errors, affecting the completion of tasks. Although current human-computer interaction robotic arms have achieved interaction and control functions to varying degrees, most of the existing control methods focus on single vision control, full teleoperation, or semi-automatic operation modes. These traditional control methods have obvious limitations when facing large environmental changes or relatively weak hardware bases. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a multi-mode intelligent interactive robotic arm system to solve the problems raised in the above background technique. The present invention combines artificial intelligence with robotic arms, reducing the technical use threshold of robotic arms and also diversifying the interactive functions of robotic arms.
[0006] To achieve the above purpose, the present invention is realized through the following technical solutions: A multi-mode intelligent interactive robotic arm system, the robotic arm system includes a first mode, a second mode, a third mode, and a fourth mode. The first mode is a grasping control method that combines a robotic arm with gestures based on machine vision, that is, an interactive robotic arm gesture following control mode; the second mode is a voice control mode of the interactive robotic arm; the third mode is a Web control mode of the interactive robotic arm; the fourth mode is an eye control following mode of the interactive robotic arm.
[0007] Further, the first mode includes the following steps:
[0008] S1.1. Obtain an image of the human hand gesture movement as sample data of the gesture movement, then build a gesture movement template according to the sample data of the gesture movement, and then use the gesture movement template as a test set;
[0009] S1.2. Process the sample data of the gesture movement into a gesture training set, and train according to the gesture movement training set to obtain a trained gesture recognition motion control model based on machine vision;
[0010] S1.3. Use a non-distorted camera to calibrate the position relationship between the center point of the object in the image and the robotic arm, and obtain the connection between the camera coordinate system and the physical coordinate system, and process the image to return the data parameters of the object center point;
[0011] S1.4. Convert the camera coordinate system and the physical coordinate system, calibrate the two coordinate systems according to a certain mathematical method using the obtained data parameters and the physical coordinate system data parameters, and then process the data into a G-code format that the robotic arm can recognize;
[0012] S1.5. Test multiple times to accurately align the camera coordinate system and the physical coordinate system. And to prevent the robotic arm from moving into the dead zone, add a region of interest (ROI) in the image, and stipulate that the robotic arm will only move within the green frame of the ROI;
[0013] S1.6. Transmit the G-code described in S1.4 to the robotic arm. The robotic arm will move according to the corresponding bone coordinate points, and judge the state of the gesture according to the image. If the gesture is zero (i.e., the hand makes a fist), the air pump will run and the robotic arm will perform an object suction operation; if the gesture is five (i.e., the palm is open), the air pump operation will stop, thus completing the grasping control of the combination of the robotic arm and the gesture.
[0014] Further, in step S1.1, obtain an image of the human hand gesture movement as sample data of the gesture movement, build a gesture movement template according to the sample data of the gesture movement, and then train the gesture recognition motion control model according to the training set of the model built in S1.2.
[0015] Further, the specific process of training the gesture recognition motion control model is:
[0016] Use a distortion-free camera to obtain a gesture sample image with a pixel size of 28*28; for the gesture sample image, construct gesture skeleton point data, which contains 21 points, and use the 9th point, that is, the palm center skeleton point, as the main coordinate point, and mark it in blue; after obtaining the skeleton point data, construct templates for different gesture actions, and use the gesture action templates as the test set; then repeat multiple times to obtain the required images to construct the training set of the model, and convert the image format to Tensor; use the training set as the input for training the model, and establish a three-layer neural network for training; test the trained model against the test set, and if the accuracy is high, it can be used, otherwise, increase the training set and continue training; when the gestures from zero to five can be accurately recognized, the gesture recognition motion control model is completed.
[0017] Further, the specific process of calibrating the two coordinate systems in step S1.4 according to a certain mathematical method and processing the data into the G-code format that the robotic arm can recognize is as follows:
[0018] Use a distortion-free camera to collect image sample data of different objects to be grasped; according to the image sample data, perform image processing to frame the outer contour of the object in the image; then use the obtained outer contour to calibrate the data parameters of the center point of the object in the image; process the data parameters of the center point to obtain the physical coordinate data parameters relative to the robotic arm; return the data parameters of the above center point and the physical coordinate data parameters; then, according to the corresponding data of the X-axis and Y-axis of the two coordinate systems, first compare the origin of the physical coordinate with the origin of the visual coordinate, and after adding or subtracting the initial value, make the origins of the two coordinate systems correspond, and then measure the coordinate change ratios of the X-axis and Y-axis of the two coordinate systems respectively, so as to obtain the corresponding linear transformation relationships of the X-axis and Y-axis of the two coordinate systems, so that the X-axis and Y-axis coordinates of the camera coordinate system and the physical coordinate system correspond one by one, and finally convert the processed data into the G-code format that the robotic arm can recognize.
[0019] Further, the G-code format is f"G1 X{100}Y{100}Z{100}F10000", where f is the data frame header, XYZ represent the X-axis, Y-axis, and Z-axis respectively, the content inside {} is the corresponding moving position, F is the moving speed, and " is the data frame tail.
[0020] Further, the second mode includes the following steps:
[0021] S2.1: Input the voice to the robotic arm system and call the speech large model;
[0022] S2.2: The speech large model recognizes the voice and converts the voice into data that it can recognize;
[0023] S2.3. Let the large model receive the above data and determine whether the received data matches the set control instructions. If it matches, after the data is parsed by the model, specific words will be sent out; if it does not match, voice input needs to be re - entered;
[0024] S2.4. The robotic arm will receive the G - code data corresponding to the control instruction, and the robotic arm will move along with the data;
[0025] S2.5. The robotic arm completes the operation corresponding to the instruction.
[0026] Furthermore, the third mode includes the following steps:
[0027] S3.1. First, write a Web page, edit the UI interface, which includes an AI dialogue window and buttons for corresponding instruction modes. Then, call the API on the web page side to deploy the large model. Next, in the dialogue window on the web page, enter the text of the operation you want to achieve;
[0028] S3.2. After the large model recognizes the above text, it converts the text into data and corresponding control instructions that the large model can recognize;
[0029] S3.3. Determine whether the received data matches the set control instructions. If it matches, after the data is parsed by the model, specific text will be sent back; if it does not match, re - input is required;
[0030] S3.4. The robotic arm will receive the G - code data corresponding to the control instruction, and the robotic arm will move along with the data;
[0031] S3.5. The robotic arm completes the operation corresponding to the instruction.
[0032] Furthermore, the fourth mode includes the following steps:
[0033] S4.1. Use the deep - learning algorithm based on YOLOV8 to train the labeled eye images of people, and use the optimal model obtained from the training to detect the left and right human eyes in the image, distinguish between the left eye and the right eye, and then use the region of interest to frame the left and right eyes respectively. The size of the frame is determined according to the size calibrated in the previous dataset;
[0034] S4.2. Use methods such as grayscale conversion, Gaussian filtering, Hough circle detection, and coordinate fitting in the OpenCV algorithm for the left and right eyes to fit the pupil coordinates of the two eyes to obtain a square box with the combined pupil coordinates. The size of the square box is formed by fitting the boxes of the left and right eyes, and the center coordinate of the square box is the origin of the coordinate after fitting;
[0035] S4.3. Convert the fitted coordinate system and the physical coordinate system. Align the two coordinate systems according to the mathematical method described in S1.4 using the obtained data parameters and physical coordinate data parameters, and then process the data into the G-code format that the robotic arm can recognize.
[0036] S4.4. Add a region of interest (ROI) to the image. It is specified that the robotic arm will only move within the green frame of the ROI to prevent the robotic arm from recognizing positions it cannot move to.
[0037] S4.5. If the human eye is not detected, the robotic arm will stop responding.
[0038] S4.6. Judge the state of the eyes according to the image. If the eyes are closed for more than one second, the air pump will run and the robotic arm will perform an object suction operation. If the eyes are closed again for more than one second, the air pump will stop running, thus completing the control method of the combination of the robotic arm and the eyes.
[0039] Advantages of the present invention:
[0040] 1. The multi-mode intelligent interactive robotic arm system controls the robotic arm for human-computer interaction through AI algorithms, speech large models, Web pages, gestures, and eye movements, and constructs a structural framework of four modes: gesture-speech-Web-eye movement, realizing the diversification of the four interaction modes of the robotic arm and expanding the application fields of the robotic arm.
[0041] 2. The multi-mode intelligent interactive robotic arm system adopts an interaction scheme of gesture-speech-Web-eye movement mode, which can simplify and improve the processing of complex tasks. The operation application difficulty threshold is very low, and the operation can be completed through simple human interaction. - Low difficulty
[0042] 3. Under the large model self-trained based on the present invention, the robotic arm can accurately complete the command actions, ensuring the accuracy of task execution. And in order to improve the efficiency and accuracy of use, it combines a neural network model and conducts a large number of trainings, improving the accuracy and precision of the robotic arm, enabling it to adapt to complex and changeable usage scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is the hardware structure framework diagram of the multi-mode interactive robotic arm in the embodiment of the present invention;
[0044] Figure 2 It is the schematic diagram of the robotic arm electric control module (left figure) and the voice control module (right figure) of the multi-mode interactive robotic arm in the embodiment of the present invention;
[0045] Figure 3 It is the hardware structure schematic diagram of the multi-mode interactive robotic arm in the embodiment of the present invention;
[0046] Figure 4 This is the flowchart of the gesture following control mode of the multi-mode interactive robotic arm in the embodiments of the present invention;
[0047] Figure 5 This is the implementation step diagram of gesture recognition in the gesture following control mode of the multi-mode interactive robotic arm in the embodiments of the present invention;
[0048] Figure 6 This is the schematic diagram of the flowchart of the gesture recognition motion control model (left figure) and gesture marking (right figure) of the multi-mode interactive robotic arm in the embodiments of the present invention;
[0049] Figure 7 This is the flowchart of the voice control mode of the multi-mode interactive robotic arm in the embodiments of the present invention;
[0050] Figure 8 This is the flowchart of the Web control mode of the multi-mode interactive robotic arm in the embodiments of the present invention;
[0051] Figure 9 This is the flowchart of the eye control following mode of the multi-mode interactive robotic arm in the embodiments of the present invention;
[0052] Figure 10 This is the left and right eye marking diagram of eye movement recognition YOLO in the eye control mode of the multi-mode interactive robotic arm in the embodiments of the present invention;
[0053] Figure 11 This is the image diagram of eye movement recognition YOLO training process in the eye control mode of the multi-mode interactive robotic arm in the embodiments of the present invention;
[0054] Figure 12 This is the principle block diagram of the multi-mode intelligent interactive robotic arm system of the present invention;
[0055] In the figure: 101, air pump suction cup; 102, robotic arm motor; 103, electronic control of the robotic arm module. Specific embodiments
[0056] To make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.
[0057] Please refer to Figures 1 to 12 , the present invention provides the following technical solutions: a multi-mode intelligent interactive robotic arm system, the robotic arm system includes a first mode, a second mode, a third mode and a fourth mode, the first mode is a grasping control method that combines a robotic arm and gestures based on machine vision, that is, an interactive robotic arm gesture following control mode; the second mode is the voice control mode of the interactive robotic arm; the third mode is the Web control mode of the interactive robotic arm; the fourth mode is the eye control following mode of the interactive robotic arm.
[0058] Example 1
[0059] In this embodiment, a first mode is provided, and this mode includes the following steps:
[0060] S1.1. Obtain an image of a human hand gesture as sample data of the gesture, then build a template of the gesture according to the sample data of the gesture, and then use the gesture template as a test set. Obtain an image of a human hand gesture as sample data of the gesture, build a template of the gesture according to the sample data of the gesture, and then train a gesture recognition motion control model according to the training set for building the model described in S1.2.
[0061] In this embodiment, the specific process of training the gesture recognition motion control model is as follows:
[0062] Use a non-distorted camera to obtain a gesture sample image with a pixel of 28*28; for the gesture sample image, build gesture skeleton point data, which contains 21 points, and use the 9th point, that is, the skeleton point of the palm center, as the main coordinate point and mark it with blue; after obtaining the skeleton point data, build templates of different gesture actions, and use the gesture action templates as a test set; repeat multiple times to obtain the required images to build the training set of the model, and convert the picture format to Tensor; use the training set as the input for training the model, establish a three-layer neural network for training; test the trained model against the test set, if the accuracy is high, it can be used, otherwise, increase the training set and continue training; when the gestures from zero to five can be accurately recognized, the gesture recognition motion control model is completed;
[0063] S1.2. Process the sample data of the gesture into a gesture training set, and train according to the gesture training set to obtain a trained gesture recognition motion control model based on machine vision;
[0064] S1.3. Use a non-distorted camera to calibrate the position relationship between the center point of the object in the image and the robotic arm, and obtain the connection between the camera coordinate system and the physical coordinate system, and process the image to return the data parameters of the object center point;
[0065] S1.4. Convert the camera coordinate system and the physical coordinate system, calibrate the two coordinate systems according to a certain mathematical method using the above-obtained data parameters and the data parameters of the physical coordinate system, and then process the data into the G-code format that the robotic arm can recognize.
[0066] The specific process of calibrating the two coordinate systems and processing the data into the G-code format that the robotic arm can recognize is as follows:
[0067] Collect image sample data of different objects to be grasped using a distortion-free camera; perform image processing based on the image sample data to frame the outer contour of the object in the image; then use the obtained outer contour to calibrate the data parameters of the center point of the object in the image; process the data parameters of the center point to obtain the physical coordinate data parameters relative to the robotic arm; return the data parameters of the center point and the physical coordinate data parameters; the given coordinate points are symmetrically transformed between two coordinate systems according to affine transformation, and then the data is processed into the G-code format recognizable by the robotic arm. The G-code format is f"G1 X{100}Y{100}Z{100}F10000", where f is the data frame header, XYZ represent the X-axis, Y-axis, and Z-axis respectively, the content inside {} is the corresponding moving position, F is the moving speed, and " is the data frame tail;
[0068] S1.5. Align the camera coordinate system and the physical coordinate system accurately through multiple tests. And to prevent the robotic arm from moving into the dead zone, add a region of interest (ROI) in the image, and stipulate that the robotic arm will only move within the green frame of the ROI;
[0069] S1.6. Transmit the G-code described in S1.4 to the robotic arm, and the robotic arm will move according to the corresponding bone coordinate points as Figure 6-2 shown, and judge the state of the gesture according to the image. If the gesture is zero (i.e., the hand makes a fist), the air pump will run and the robotic arm will perform the object suction operation; if the gesture is five (i.e., the palm is open), the air pump operation will stop, thus completing the grasping control combining the robotic arm and the gesture.
[0070] Embodiment 2
[0071] In this embodiment, a second mode is provided, which specifically includes the following steps:
[0072] S2.1. Input the voice to the robotic arm system and call the voice large model;
[0073] S2.2. The voice large model recognizes the voice and converts the voice into data that it can recognize;
[0074] S2.3. Let the large model receive the above data and judge whether the received data matches the set control instruction. If it matches, the data will emit a specific speech after being parsed by the model; if it does not match, a new voice input is required;
[0075] S2.4. The robotic arm will receive the G-code data of the corresponding control instruction and move along with the data;
[0076] S2.5. The robotic arm completes the operation of the corresponding instruction.
[0077] The voice interaction provided in this embodiment is another major core advantage of the multi-mode interaction robotic arm. With the built-in advanced artificial intelligence large model, the robotic arm can quickly and accurately recognize and translate voice commands issued by humans. Whether it is a simple action command or a complex task description, it can quickly understand and execute efficiently. This voice interaction method not only reduces the technical threshold for operators but also enables the robotic arm to achieve flexible and efficient control in working scenarios where the hands are busy or the environment is noisy.
[0078] Embodiment 3
[0079] In this embodiment, a third mode is provided, which specifically includes the following steps:
[0080] S3.1. First, write a Web page and edit the UI interface, which includes an AI dialogue window and buttons corresponding to instruction modes. Then, call the API on the web page side to deploy the large model. Next, in the dialogue window on the web page, enter the text for the operation you want to achieve.
[0081] S3.2. After the large model recognizes the above text, it converts the text into data and corresponding control instructions that the large model can recognize.
[0082] S3.3. Determine whether the received data matches the set control instructions. If it matches, the data is sent back with specific text after being parsed by the model. If it does not match, re-entry is required.
[0083] S3.4. The robotic arm will receive the G-code data corresponding to the control instructions and move along with the data.
[0084] S3.5. The robotic arm completes the operation corresponding to the instruction.
[0085] Embodiment 4
[0086] In this embodiment, a fourth mode is provided, which specifically includes the following steps:
[0087] S4.1. Use the deep learning algorithm based on YOLOV8 to train the labeled eye images of people, and use the optimal model obtained from the training to detect the left and right human eyes in the image and distinguish between the left and right eyes. Then, use the region of interest to frame the left and right eyes respectively. The size of the frame is determined according to the size calibrated in the previous dataset.
[0088] S4.2. Use methods such as grayscale conversion, Gaussian filtering, Hough circle detection, and coordinate fitting in the OpenCV algorithm for the left and right eyes to fit the pupil coordinates of the two eyes to obtain a square box with the combined pupil coordinates. The size of the square box is formed by fitting the boxes of the left and right eyes, and the center coordinates of the square box are the origin of the coordinates after fitting.
[0089] S4.3. Convert the fitted coordinate system and the physical coordinate system. Align the two coordinate systems according to the mathematical method described in S1.4 using the obtained data parameters and physical coordinate data parameters, and then process the data into the G-code format that the robotic arm can recognize.
[0090] S4.4. Add a region of interest (ROI) to the image and specify that the robotic arm will only move within the green frame of the ROI to prevent the robotic arm from recognizing positions where it cannot move.
[0091] S4.5. If the human eye is not detected, the robotic arm will stop responding.
[0092] S4.6. Judge the state of the eyes according to the image. If the eyes are closed for more than one second, the air pump will run and the robotic arm will perform an object suction operation. If the eyes are closed again for more than one second, the air pump will stop running, thus completing the control method of the combination of the robotic arm and the eyes.
[0093] In this embodiment, the eye control following interaction mode integrates a high-precision eye movement tracking device, which can capture the eye movements and line-of-sight directions of the operator in real time, and accurately analyze the intention. When the operator's gaze focuses on the target object, the robotic arm can quickly locate and act according to the preset program, realizing precise grasping or moving. It has significant advantages in scenarios with high concentration of attention and limited operation space, greatly improving the operation convenience and precision, and further expanding the application scenarios and interaction capabilities.
[0094] The above shows and describes the basic principles, main features and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic features of the present invention, the present invention can be implemented in other specific forms.
[0095] In addition, it should be understood that although this specification is described according to the embodiments, not each embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. Multi-mode intelligent interactive robotic arm system, characterized by: The robotic arm system includes a first mode, a second mode, a third mode and a fourth mode. The first mode is a grasping control method based on a combination of a robotic arm and gestures based on machine vision, namely, an interactive robotic arm gesture following control mode; the second mode is a voice control mode of an interactive robotic arm; the third mode is an interactive robotic arm Web control mode; and the fourth mode is an interactive robotic arm eye control following mode.
2. The multi-mode intelligent interactive robotic arm system according to claim 1, characterized in that: The first mode comprises the following steps: S1.
1. Obtain images of human hand gestures as sample data of gestures, build a gesture template based on the sample data of gestures, and then use the gesture template as a test set; S1.2, processing the sample data of the gesture action into a gesture training set, performing training according to the gesture action training set, and obtaining a trained machine vision-based gesture recognition motion control model; S1.
3. Use a distortion-free camera to calibrate the positional relationship between the center point of the object in the image and the robotic arm, and obtain the connection between the camera coordinate system and the physical coordinate system, and process the image to return the data parameters of the center point of the object; S1.4, convert the camera coordinate system to the physical coordinate system, calibrate the two coordinate systems according to a certain mathematical method using the above-obtained data parameters and the physical coordinate system data parameters, and then process the data into a G-code format that can be recognized by the robot arm; S1.
5. Perform multiple tests to accurately align the camera coordinate system with the physical coordinate system. In order to prevent the robot arm from moving into a dead zone, add a region of interest (ROI) to the image, and stipulate that the robot arm will only move in the green frame of the ROI. S1.
6. Transfer the G-code described in S1.4 to the robotic arm. The robotic arm will move according to the corresponding bone coordinate points and judge the state of the gesture based on the image. If the gesture is zero (i.e., the hand is clenched into a fist), the air pump will run and the robotic arm will perform the object suction operation; if the gesture is five (i.e., the palm is open), the air pump will stop running, thereby completing the grasping control of the robotic arm and the gesture.
3. The multi-mode intelligent interactive robotic arm system according to claim 2, characterized in that: In the step S1.1, an image of a human hand gesture is obtained as sample data of the gesture, a template of the gesture is constructed according to the sample data of the gesture, and then the gesture recognition motion control model is trained according to the training set of the model constructed in S1.
2.
4. The multi-mode intelligent interactive robotic arm system according to claim 3, characterized in that: The specific process of training the gesture recognition motion control model is as follows: Use a distortion-free camera to obtain gesture sample images with a pixel size of 28*28; for gesture sample images, build gesture skeleton point data; build templates of different gesture actions, and use the gesture action templates as a test set; repeatedly obtain the required images to build the training set of the model, and convert the image format to Tensor; use the training set as the input of the training model, and establish a three-layer neural network for training; test the trained model against the test set; when the gestures from zero to five can be accurately recognized, the gesture recognition motion control model is completed.
5. The multi-mode intelligent interactive robotic arm system according to claim 2, characterized in that: The specific process of step S1.4 calibrating the two coordinate systems according to a certain mathematical method and processing the data into a G-code format that can be recognized by the robot arm is as follows: Use a distortion-free camera to collect image sample data of different objects to be grasped; perform image processing based on the image sample data to frame the outer contour of the object in the image; then use the obtained outer contour to calibrate the data parameters of the center point of the object in the image; return the data parameters of the above center point and the physical coordinate data parameters; then based on the corresponding data of the X-axis and Y-axis of the two coordinate systems, first compare the origin of the physical coordinate with the origin of the visual coordinate, add and subtract the initial values to make the origins of the two coordinate systems correspond, and then measure the coordinate change ratio of the two coordinate systems of the X-axis and Y-axis respectively, and you can get the linear transformation relationship corresponding to the X-axis and Y-axis of the two coordinate systems, so that the X-axis and Y-axis coordinates of the camera coordinate system and the physical coordinate system correspond one by one, and finally change the processed data into a G-code format that the robotic arm can recognize.
6. The multi-mode intelligent interactive robotic arm system according to claim 5, characterized in that: The G-code format is f"G1 X{100}Y{100}Z{100}F10000", f" is the data frame header, XYZ represents the X-axis, Y-axis and Z-axis respectively, {} is the corresponding moving position, F is the moving speed, and " is the data frame tail.
7. The multi-mode intelligent interactive robotic arm system according to claim 1, characterized in that: The second mode comprises the following steps: S2.
1. Input the speech into the robot arm system and call the speech model; S2.2, the speech is recognized by the speech model and converted into data that it can recognize; S2.3, the above data is received by the big model, and it is determined whether the received data matches the set control instructions. If it matches, the data is analyzed by the model and a specific speech is issued; if it does not match, the voice input is required again; S2.4, the robot arm will receive the G-code data corresponding to the control command, and the robot arm will move according to the data; S2.
5. The robot arm completes the operation corresponding to the instruction.
8. The multi-mode intelligent interactive robotic arm system according to claim 1, characterized in that: The third mode comprises the following steps: S3.
1. First, write a web page, then call the API on the web page, deploy the large model, and then enter the text you want to operate in the dialog window on the web page; S3.2, after the large model recognizes the above text, it converts the text into data and corresponding control instructions that the large model can recognize; S3.3, determine whether the received data matches the set control instruction; S3.4, the robot arm will receive the G-code data corresponding to the control command, and the robot arm will move according to the data; S3.
5. The robotic arm completes the operation corresponding to the instruction.
9. The multi-mode intelligent interactive robotic arm system according to claim 1, characterized in that: The fourth mode comprises the following steps: S4.
1. Use the YOLOV8-based deep learning algorithm to train the labeled human eye images, and use the trained optimal model to detect the left and right eyes in the image and distinguish the left and right eyes. Then use the region of interest to frame the left and right eyes respectively. The size of the frame is determined according to the size calibrated in the previous dataset. S4.
2. Use grayscale, Gaussian filtering, Hough circle detection, coordinate fitting and other methods in the OpenCV algorithm to fit the pupil coordinates of the left and right eyes to obtain a square box with the pupil coordinates combined into one. The size of the square box is fitted by the square boxes of the left and right eyes, and the center coordinates of the square box are the coordinate origin after fitting; S4.3, converting the fitted coordinate system to the physical coordinate system, aligning the two coordinate systems using the above-obtained data parameters and the physical coordinate data parameters according to the mathematical method described in S1.4, and then processing the data into a G-code format that can be recognized by the robot arm; S4.4, and add a region of interest ROI in the image, stipulating that the robot arm will move only in the green frame of the ROI, to prevent the robot arm from identifying a position that it cannot move to; S4.5, if the human eye is not detected, the robot arm will stop responding; S4.
6. Determine the state of the eyes based on the image. If the eyes are closed for more than one second, the air pump will start and the robotic arm will perform the object suction operation. If the eyes are closed for more than one second again, the air pump will stop running, thereby completing the control method of combining the robotic arm and the eyes.
Citation Information
Patent Citations
Service robot control platform system and multimode intelligent interaction and intelligent behavior realizing method thereof
CN102323817A
Intelligent cognitive robot system architecture
CN116100548A
Mechanical arm gesture control method and system based on visual control guidance
CN118769252A
Gesture recognition algorithm and system based on color attention in AR operation environment
CN119356524A
Method for Selecting Interactivity Mode
US20150227275A1