Multi-modal intelligent interactive robotic arm system

By combining gesture, voice, web, and eye-tracking control, the multi-mode intelligent interactive robotic arm system solves the adaptability problem of traditional robotic arms in complex scenarios, achieves efficient and accurate task execution, and expands application scenarios.

CN120080317BActive Publication Date: 2025-12-16HUBEI NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510350012.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-12-16
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

Existing robotic arms struggle to adapt quickly to production line adjustments when faced with complex and ever-changing work scenarios. They are also prone to operational errors due to signal interference during remote medical surgeries and identification errors in logistics environments. Traditional control methods have limitations.

Method used

The system employs a multi-mode intelligent interactive robotic arm system, combining artificial intelligence with the robotic arm to provide four interaction modes: gesture, voice, web, and eye tracking. It is controlled through AI algorithms, large voice models, and neural network models.

Benefits of technology

It has enabled the robotic arm to have diversified interactive functions, reduced the difficulty of operation, improved the accuracy and efficiency of task execution in complex scenarios, and broadened the application field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120080317B_ABST
    Figure CN120080317B_ABST
Patent Text Reader

Abstract

The application provides a multi-mode intelligent interactive mechanical arm system, which comprises a first mode, a second mode, a third mode and a fourth mode, the first mode is a machine vision-based mechanical arm and gesture combined grasping control method, namely an interactive mechanical arm gesture following control mode, the second mode is a voice control mode of the interactive mechanical arm, the third mode is a Web control mode of the interactive mechanical arm, and the fourth mode is an eye control following mode of the interactive mechanical arm. The application controls the mechanical arm to realize human-computer interaction through AI algorithms, a voice large model, a Web page, gestures and eye movements, constructs a structure framework of the four modes of gestures, voice, Web and eye movements, realizes the diversification of the four interactive modes of the mechanical arm, widens the application field of the mechanical arm, improves the precision and accuracy of the mechanical arm, and enables the mechanical arm to adapt to complex and changeable use scenarios and complete tasks with low errors in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of mechanical arms, in particular to a multi-mode intelligent interactive mechanical arm system. BACKGROUND

[0002] Mechanical arms have been widely used in industrial manufacturing, medical services, logistics transportation and other core fields due to their unique technical characteristics, and have become a key factor in promoting the intelligent upgrading of industries and improving the quality of people's lives. In the field of industrial manufacturing, mechanical arms can assemble parts with high precision in automobile manufacturing, such as engine assembly, which is faster than manual work and can effectively avoid errors, significantly improving production efficiency and quality.

[0003] However, in the face of complex and variable work scenarios where the production line needs to grab different parts, the traditional mechanical arm control method relying on preset programs and fixed modes is difficult to adapt quickly, and reprogramming and debugging takes time and affects progress. In the field of medical services, mechanical arms can assist doctors with high precision and stability in minimally invasive brain surgery, reducing the risk of surgery and improving the success rate. However, in remote medical surgery, traditional mechanical arms are affected by poor signals and are prone to operation errors, affecting the effectiveness of the surgery.

[0004] In the field of logistics transportation, mechanical arms in large e-commerce warehouses can quickly sort and transport goods according to system instructions, improving operational efficiency and reducing labor burden. However, in dimly lit environments where object features are not obvious, traditional mechanical arms are prone to recognition errors, affecting task completion. Although current human-robot interaction mechanical arms have achieved interaction and control functions to varying degrees, existing control methods mostly focus on single visual control, full teleoperation or semi-automatic operation modes. These traditional control methods have obvious limitations when facing large environmental changes or relatively weak hardware foundations. SUMMARY

[0005] To address the deficiencies of the prior art, the present application aims to provide a multi-mode intelligent interactive mechanical arm system to solve the problems raised in the background art. The present application combines artificial intelligence with mechanical arms to reduce the technical threshold for using mechanical arms and also diversify the interactive functions of mechanical arms.

[0006] To achieve the above-mentioned purpose, the present application is implemented by the following technical solution: a multi-mode intelligent interactive mechanical arm system, the mechanical arm system includes a first mode, a second mode, a third mode and a fourth mode, the first mode is a mechanical arm and gesture combination based on machine vision grabbing control method, namely interactive mechanical arm gesture following control mode; the second mode is a voice control mode of the interactive mechanical arm; the third mode is a Web control mode of the interactive mechanical arm; the fourth mode is an eye control following mode of the interactive mechanical arm.

[0007] Further, the first mode comprises the following steps:

[0008] S1.1, acquire the image of the human hand gesture action as the sample data of the gesture action, and then build a gesture action template according to the sample data of the gesture action, and then take the gesture action template as a test set;

[0009] S1.2, process the sample data of the gesture action into a gesture training set, train according to the gesture action training set, and obtain a trained gesture recognition motion control model based on machine vision;

[0010] S1.3, use a non-distorted camera to calibrate the position relationship between the center point of the object in the image and the mechanical arm, and obtain the relationship between the camera coordinate system and the physical coordinate system, and process the image to return the data parameters of the object center point;

[0011] S1.4, convert the camera coordinate system and the physical coordinate system, calibrate the two coordinate systems according to certain mathematical methods by using the above obtained data parameters and physical coordinate system data parameters, and process the data into G-code format that can be recognized by the mechanical arm;

[0012] S1.5, multiple tests are used to accurately align the camera coordinate system and the physical coordinate system, and in order to prevent the mechanical arm from moving to the dead zone, an area of interest (ROI, Region of Interest) is added in the image, and it is stipulated that only in the ROI green frame the mechanical arm will move;

[0013] S1.6, deliver the G-code in S1.4 to the mechanical arm, and the mechanical arm will move according to the corresponding skeletal coordinate points, and judge the state of the gesture according to the image, if the gesture is zero (i.e. the palm is clenched), the air pump runs, and the mechanical arm performs object suction operation; if the gesture is five (i.e. the palm is open), stop the air pump running, thereby completing the grasping control of the combination of the mechanical arm and the gesture.

[0014] Further, in the step S1.1, the image of the human hand gesture action is acquired as the sample data of the gesture action, and a gesture action template is built according to the sample data of the gesture action, and then the gesture recognition motion control model is trained according to the training set of the model in S1.2.

[0015] Further, the specific process of training the gesture recognition motion control model is as follows:

[0016] The gesture sample image with pixels of 28*28 is obtained by using an anamorphic camera; for the gesture sample image, a gesture skeleton point data is built, containing 21 points, and the 9th point, i.e. the palm center skeleton point, is taken as the main coordinate point and marked with blue; after obtaining the skeleton point data, a sample of different gesture actions is built, and the gesture action sample is taken as the test set; then the required images are repeatedly obtained to build the training set of the model, and the picture format is converted into Tensor; the training set is taken as the input of the training model, a three-layer neural network is established for training; the trained model is tested against the test set, if the accuracy is high, it can be used, otherwise the training set will be increased for further training; when the gesture can be accurately recognized from zero to five, the gesture recognition motion control model is completed.

[0017] Further, the two coordinate systems in step S1.4 are calibrated according to a certain mathematical method, and the specific process of processing the data into the G-code format that can be recognized by the mechanical arm is as follows:

[0018] The image sample data of different objects to be grasped is collected by using an anamorphic camera; according to the image sample data, image processing is performed to frame the peripheral contour of the object in the image; then the data parameters of the center point of the object in the image are calibrated using the obtained peripheral contour; the data parameters of the center point are processed to obtain the physical coordinate data parameters relative to the mechanical arm; the data parameters of the center point and the physical coordinate data parameters are returned; then according to the corresponding data of the X-axis and Y-axis of the two coordinate systems, the origins of the two coordinate systems are compared after adding or subtracting the initial value, and then the origins of the two coordinate systems are corresponded, and the coordinate change proportions of the X-axis and Y-axis of the two coordinate systems are measured respectively, so that the linear transformation relationship corresponding to the X-axis and Y-axis of the two coordinate systems is obtained, so that the X-axis and Y-axis coordinates of the camera coordinate system and the physical coordinate system are one-to-one corresponding, and finally the processed data is changed into the G-code format that can be recognized by the mechanical arm.

[0019] Further, the G-code format is f"G1 X{100}Y{100}Z{100}F10000", f" is the data frame header, XYZ represents X-axis, Y-axis and Z-axis respectively, the inside of {} is the corresponding moving position, F is the moving speed, and " is the data frame trailer.

[0020] Further, the second mode includes the following steps:

[0021] S2.1, input the voice to the mechanical arm system, and call the voice large model;

[0022] S2.2, the voice is recognized by the voice large model, and the voice is converted into data that can be recognized by the voice large model;

[0023] S2.3, the above data, let the large model receive, and judge whether the received data matches the set control instruction, if it matches, the data is sent after model analysis Special speech; if it does not match, it needs to be re-inputted;

[0024] S2.4, the mechanical arm will receive the G-code data corresponding to the control instruction, and the mechanical arm will move with the data;

[0025] S2.5, the mechanical arm completes the operation corresponding to the instruction.

[0026] Further, the third mode includes the following steps:

[0027] S3.1, first write a Web page, edit the UI interface, which contains the AI dialogue window, the key corresponding to the instruction mode, then call the API on the Web side, deploy the large model, and input the text of the operation you want to realize in the dialogue window on the Web page;

[0028] S3.2, the large model identifies the above text and converts the text into data and corresponding control instructions that the large model can recognize;

[0029] S3.3, judge whether the received data matches the set control instruction, if it matches, the data is sent after model analysis Special text; if it does not match, it needs to be re-inputted;

[0030] S3.4, the mechanical arm will receive the G-code data corresponding to the control instruction, and the mechanical arm will move with the data;

[0031] S3.5, the mechanical arm completes the operation corresponding to the instruction.

[0032] Further, the fourth mode includes the following steps:

[0033] S4.1, use the deep learning algorithm based on YOLOV8 to train the labeled eye image of the person, and use the optimal model obtained by training to detect the left and right eyes in the image and distinguish the left and right eyes, and then use the region of interest to frame the left and right eyes respectively, the size of the frame is determined according to the size of the pre-data set calibration;

[0034] S4.2, the left and right eyes are fitted with the pupil coordinates of the two eyes using the methods of gray scale, Gaussian filter, Hough circle detection, coordinate fitting, etc. in OpenCV algorithm, a frame with combined pupil coordinates is obtained, the size of the frame is fitted from the left and right eye frames, and the center coordinates of the frame are the coordinate origin after fitting;

[0035] S4.3, convert the fitted coordinate system with the physical coordinate system, use the data parameters obtained above and the physical coordinate data parameters to align the two coordinate systems according to the mathematical method described in S1.4, and then process the data into a G-code format that can be recognized by the mechanical arm;

[0036] S4.4, and add a region of interest (ROI) in the image, specify that the mechanical arm will only move in the ROI green color frame, and prevent the mechanical arm from recognizing positions that cannot be moved to;

[0037] S4.5, if the human eye is not detected, the mechanical arm will stop responding;

[0038] S4.6, determine the state of the eye according to the image, if the eye is closed for more than one second, the air pump operates, and the mechanical arm performs object suction operation; if the eye is closed again for more than one second, the air pump stops operating, thereby completing the control method of the mechanical arm combined with the eye.

[0039] Advantages of the present application:

[0040] 1. The multi-mode intelligent interactive mechanical arm system controls the mechanical arm for human-computer interaction through AI algorithm, voice large model, Web page, gesture and eye movement, and constructs a structure framework of gesture-voice-Web-eye movement four modes, realizes the diversification of the four interactive modes of the mechanical arm, and expands the application field of the mechanical arm.

[0041] 2. The multi-mode intelligent interactive mechanical arm system adopts the interactive scheme of gesture-voice-Web-eye movement mode, can simplify and perfect the processing of complex tasks, has very low threshold in operation and application difficulty, and can complete the operation through simple human interaction.

[0042] 3. The multi-mode intelligent interactive mechanical arm system under the large model trained by the present application can accurately complete the instruction action, ensure the accuracy of task execution, and combine the neural network model to improve the efficiency and accuracy of use, and has been trained a lot, improves the precision and accuracy of the mechanical arm, and makes it can adapt to complex and changeable use scene. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 It is a hardware structure framework diagram of the multi-mode interactive mechanical arm in the embodiment of the present application;

[0044] Figure 2 It is a schematic diagram of the mechanical arm electric control module (left) and the voice control module (right) of the multi-mode interactive mechanical arm in the embodiment of the present application;

[0045] Figure 3 It is a hardware structure schematic diagram of the multi-mode interactive mechanical arm in the embodiment of the present application;

[0046] Figure 4 The step flow chart of the gesture follow-up control mode of the multi-mode interactive mechanical arm in the embodiment of the application;

[0047] Figure 5 The gesture recognition implementation step chart of the gesture follow-up control mode of the multi-mode interactive mechanical arm in the embodiment of the application;

[0048] Figure 6 The gesture recognition motion control model flow chart (left chart) and gesture mark (right chart) of the multi-mode interactive mechanical arm in the embodiment of the application;

[0049] Figure 7 The step flow chart of the voice control mode of the multi-mode interactive mechanical arm in the embodiment of the application;

[0050] Figure 8 The step flow chart of the Web control mode of the multi-mode interactive mechanical arm in the embodiment of the application;

[0051] Figure 9 The step flow chart of the eye control follow-up mode of the multi-mode interactive mechanical arm in the embodiment of the application;

[0052] Figure 10 The eye movement recognition YOLO left and right eye mark chart of the eye control mode of the multi-mode interactive mechanical arm in the embodiment of the application;

[0053] Figure 11 The eye movement recognition YOLO training processing image chart of the eye control mode of the multi-mode interactive mechanical arm in the embodiment of the application;

[0054] Figure 12 The principle block diagram of the multi-mode intelligent interactive mechanical arm system of the application;

[0055] In the figure: 101, air pump suction disc; 102, mechanical arm motor; 103, mechanical arm module electric control. DETAILED DESCRIPTION

[0056] In order to make the technical means, creative features, purposes and effects realized by the application easy to understand, the application will be further described below in combination with specific embodiments.

[0057] Please refer to Figures 1 to 12 The application provides the following technical solutions: a multi-mode intelligent interactive mechanical arm system, the mechanical arm system comprising a first mode, a second mode, a third mode and a fourth mode, the first mode being a mechanical arm and gesture mutual combination based on machine vision grabbing control method, namely, an interactive mechanical arm gesture follow-up control mode; the second mode being an interactive mechanical arm voice control mode; the third mode being an interactive mechanical arm Web control mode; and the fourth mode being an interactive mechanical arm eye control follow-up mode.

[0058] Embodiment 1

[0059] The first mode provided in this embodiment includes the following steps:

[0060] S1.1, obtaining an image of a human hand gesture action as sample data of the gesture action, and then building a gesture action template according to the sample data of the gesture action, and then taking the gesture action template as a test set. Obtain an image of a human hand gesture action as sample data of the gesture action, build a gesture action template according to the sample data of the gesture action, and then train the gesture recognition motion control model according to the training set of the model built in S1.2.

[0061] In this embodiment, the specific process of training the gesture recognition motion control model is as follows:

[0062] Use an undistorted camera to obtain a gesture sample image with 28*28 pixels; for the gesture sample image, build gesture skeleton point data, which contains 21 points, and take the 9th point, i.e. the palm center skeleton point, as the main coordinate point and mark it with blue color; after obtaining the skeleton point data, build a template of different gesture actions, and take the gesture action template as a test set; then repeat the process of obtaining the required images to build the training set of the model, and convert the image format to Tensor; take the training set as the input of the training model, build a three-layer neural network for training; test the trained model against the test set, if the accuracy is high, the model can be used, otherwise increase the training set for further training; when the gesture recognition motion control model can accurately recognize gestures from zero to five, the training is completed.

[0063] S1.2, process the sample data of the gesture action into a gesture training set, train according to the gesture action training set, and obtain a trained gesture recognition motion control model based on machine vision;

[0064] S1.3, use an undistorted camera to calibrate the position relationship between the center point of the object in the image and the mechanical arm, and obtain the relationship between the camera coordinate system and the physical coordinate system, and process the image to return the data parameters of the object center point;

[0065] S1.4, convert the camera coordinate system and the physical coordinate system, calibrate the two coordinate systems according to certain mathematical methods using the data parameters obtained above and the physical coordinate system data parameters, and then process the data into G-code format that can be recognized by the mechanical arm.

[0066] The specific process of calibrating the two coordinate systems and processing the data into G-code format that can be recognized by the mechanical arm is as follows:

[0067] Collecting image sample data of different objects to be grabbed by using an undistorted camera; performing image processing according to the image sample data to frame the peripheral contour of the object in the image; using the obtained peripheral contour to calibrate the data parameters of the center point of the object in the image; processing the data parameters of the center point to obtain physical coordinate data parameters relative to the mechanical arm; returning the data parameters of the center point and the physical coordinate data parameters; and performing affine transformation on the given coordinate points to make the two coordinate systems symmetrical, and processing the data into a G-code format recognizable by the mechanical arm. The G-code format is f"G1 X{100}Y{100}Z{100}F10000", f" is the data frame header, XYZ respectively represent X axis, Y axis and Z axis, the internal {} is the corresponding movement position, F is the movement speed, and " is the data frame trailer.

[0068] S1.5, multiple tests are performed to accurately align the camera coordinate system with the physical coordinate system, and in order to prevent the mechanical arm from moving to a dead zone, a region of interest (ROI) is added in the image, and it is stipulated that the mechanical arm will only move in the ROI green frame;

[0069] S1.6, the G-code described in S1.4 is transmitted to the mechanical arm, and the mechanical arm will move according to the corresponding bone coordinate point as shown in Figure 6-2 , and the state of the gesture is judged according to the image, if the gesture is zero (i.e. the palm is clenched), the air pump runs, and the mechanical arm performs object suction operation; if the gesture is five (i.e. the palm is open), the air pump stops running, thereby completing the grabbing control of the combination of the mechanical arm and the gesture.

[0070] Embodiment 2

[0071] The second mode is provided in this embodiment, which specifically includes the following steps:

[0072] S2.1, inputting voice to the mechanical arm system, and calling a large voice model;

[0073] S2.2, identifying the voice by the large voice model, and converting the voice into data recognizable by the large voice model;

[0074] S2.3, inputting the above-mentioned data into the large model, and judging whether the received data matches the set control instruction, if yes, the data is analyzed by the model to output a specific speech, if not, the voice needs to be input again;

[0075] S2.4, the mechanical arm receives G-code data corresponding to the control instruction, and the mechanical arm moves according to the data;

[0076] S2.5, the mechanical arm completes the operation corresponding to the instruction.

[0077] The voice interaction provided in this embodiment is another major advantage of the multi-mode interactive robot arm. With the built-in advanced artificial intelligence large model, the robot arm can quickly and accurately recognize and translate voice commands from humans. Whether it is a simple action command or a complex task description, it can quickly understand and efficiently execute. This voice interaction method not only reduces the technical threshold of the operator, but also enables the robot arm to be flexible and efficient in control even in a noisy environment.

[0078] Embodiment 3

[0079] The third mode is provided in this embodiment, which specifically includes the following steps:

[0080] S3.1, first write a Web page, edit a UI interface, which contains an AI conversation window and a button corresponding to the instruction mode, then call the API on the Web page, deploy the large model, and input the text of the operation to be implemented in the conversation window on the Web page;

[0081] S3.2, after the large model recognizes the above text, the text is converted into data and corresponding control instructions that the large model can recognize;

[0082] S3.3, determine whether the received data matches the set control instructions, if they match, the data is sent back to the specific text after being parsed by the model; if they do not match, they need to be re-inputted;

[0083] S3.4, the robot arm will receive G-code data corresponding to the control instructions, and the robot arm will move according to the data;

[0084] S3.5, the robot arm completes the operation corresponding to the instructions.

[0085] Embodiment 4

[0086] The fourth mode is provided in this embodiment, which specifically includes the following steps:

[0087] S4.1, use a deep learning algorithm based on YOLOV8 to train the labeled eye images of people, and use the optimal model obtained by training to detect the left and right eyes in the image and distinguish the left and right eyes, then use the region of interest to frame the left and right eyes, and the size of the frame is determined according to the size of the pre-data set calibration;

[0088] S4.2, use OpenCV algorithm methods such as graying, Gaussian filtering, Hough circle detection, and coordinate fitting to fit the pupil coordinates of the two eyes to obtain a frame with combined pupil coordinates, and the size of the frame is fitted from the frames of the left and right eyes, and the center coordinates of the frame are the original coordinates after fitting;

[0089] S4.3, converting the fitted coordinate system with the physical coordinate system, using the data parameters obtained above and the physical coordinate data parameters to align the two coordinate systems according to the mathematical method described in S1.4, and then processing the data into a G-code format that can be recognized by the mechanical arm;

[0090] S4.4, and adding a region of interest (ROI) in the image, and stipulating that the mechanical arm will only move in the ROI green color frame to prevent the mechanical arm from recognizing a position that cannot be moved to;

[0091] S4.5, if the human eye is not detected, the mechanical arm will stop responding;

[0092] S4.6, judging the state of the eye according to the image, if the eye is closed for more than one second, the air pump is operated, and the mechanical arm performs the object suction operation; if the eye is closed again for more than one second, the air pump is stopped, thereby completing the control method of the mechanical arm and the eye.

[0093] In this embodiment, the eye control following interaction mode integrates a high-precision eye movement tracking device, captures the eye movement and gaze direction of the operator in real time, accurately analyzes the intention, and when the operator's gaze focuses on a target object, the mechanical arm can quickly position and act according to the preset program, realizing precise grabbing or moving. In the scene where attention is highly concentrated and the operation space is limited, the advantages are significant, and the operation convenience and accuracy are greatly improved, further expanding the application scenarios and interaction capabilities.

[0094] The above shows and describes the basic principles and main features of the present application and the advantages of the present application. It is obvious for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and the present application can be realized in other specific forms without departing from the spirit or essential characteristics of the present application.

[0095] In addition, it should be understood that although the present specification is described in terms of embodiments, not every embodiment contains only one independent technical solution, and the description manner of the specification is only for the sake of clarity, those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be appropriately combined to form other embodiments that those skilled in the art can understand.

Claims

1. A multi-mode intelligent interactive robotic arm system, characterized in that: The robotic arm system includes a first mode, a second mode, a third mode, and a fourth mode. The first mode is a grasping control method that combines machine vision with gestures, i.e., an interactive robotic arm gesture-following control mode. The second mode is a voice control mode for the interactive robotic arm. The third mode is a web control mode for the interactive robotic arm. The fourth mode is an eye-tracking mode for the interactive robotic arm. The first mode includes the following steps: S1.

1. Obtain images of human hand gestures as sample data of gestures, then build a template of gestures based on the sample data of gestures, and then use the template of gestures as a test set. S1.2 Process the sample data of gesture actions into a gesture training set, and train according to the gesture action training set to obtain a trained gesture recognition motion control model based on machine vision. S1.3 Use a distortion-free camera to calibrate the positional relationship between the center point of the object in the image and the robotic arm, and obtain the connection between the camera coordinate system and the physical coordinate system, as well as process the image to return the data parameters of the object's center point. S1.

4. Convert the camera coordinate system to the physical coordinate system. Using the data parameters obtained above and the physical coordinate system data parameters, calibrate the two coordinate systems according to mathematical methods, and then process the data into G-code format that the robotic arm can recognize. S1.

5. After multiple tests, the camera coordinate system and the physical coordinate system were precisely aligned. In order to prevent the robotic arm from moving into the dead zone, a region of interest was added to the image, and it was stipulated that the robotic arm would only move within the green ROI. S1.

6. The G-code described in S1.4 is transmitted to the robotic arm. The robotic arm will move according to the corresponding skeletal coordinate points and judge the state of the gesture based on the image. If the gesture is zero, the air pump will run and the robotic arm will perform the object suction operation; if the gesture is five, the air pump will stop running, thereby completing the grasping control of the combination of robotic arm and gesture.

2. The multi-mode intelligent interactive robotic arm system according to claim 1, characterized in that: In step S1.1, images of human hand gestures are obtained as sample data of gestures. Based on the sample data of gestures, a template of gestures is built. Then, the gesture recognition motion control model is trained based on the gesture training set described in S1.

2.

3. The multi-mode intelligent interactive robotic arm system according to claim 2, characterized in that: The specific process of training the gesture recognition motion control model is as follows: Use a distortion-free camera to acquire 28*28 pixel gesture sample images; for the gesture sample images, construct gesture skeleton point data; construct templates for different gesture actions, and use the gesture action templates as the test set; repeatedly acquire the required images to construct the training set of the model, and convert the image format to Tensor; use the training set as the input to train the model, and build a three-layer neural network for training; test the trained model against the test set; when the gesture is accurately recognized from zero to five, the gesture recognition motion control model is complete.

4. The multi-mode intelligent interactive robotic arm system according to claim 1, characterized in that: The specific process of step S1.4, which involves calibrating the two coordinate systems using mathematical methods and processing the data into a G-code format that the robotic arm can recognize, is as follows: Image sample data of different objects to be grasped is collected using a distortion-free camera; image processing is performed based on the image sample data to define the outer contour of the object in the image; the obtained outer contour is then used to determine the data parameters of the center point of the object in the image; the data parameters of the center point and the physical coordinate data parameters are returned; based on the corresponding data of the X and Y axes of the two coordinate systems, the origin of the physical coordinate system is first compared with the origin of the visual coordinate system, and after adding or subtracting initial values, the origins of the two coordinate systems are aligned; then the coordinate change ratios of the X and Y axes of the two coordinate systems are measured respectively to obtain the linear transformation relationship of the X and Y axes of the two coordinate systems, so that the X and Y axes of the camera coordinate system correspond one-to-one with the physical coordinate system; finally, the processed data is converted into G-code format that the robotic arm can recognize.

5. The multi-mode intelligent interactive robotic arm system according to claim 4, characterized in that: The G-code format is f"G1 X{100} Y{100} Z{100} F10000", where f" is the data frame header, XYZ represent the X-axis, Y-axis, and Z-axis respectively, the area inside {} is the corresponding movement position, F is the movement speed, and " is the data frame tail.

6. The multi-mode intelligent interactive robotic arm system according to claim 1, characterized in that: The second mode includes the following steps: S2.1 Input the voice into the robotic arm system and call the large voice model; S2.2, Recognize speech using a large speech model and convert the speech into data that it can recognize; S2.

3. The above data is received by the large model, and it is determined whether the received data matches the set control command. If they match, the data is parsed by the model and a specific voice is emitted; if they do not match, the voice input needs to be repeated. S2.4 The robotic arm will receive the G-code data corresponding to the control command, and the robotic arm will move according to the data; S2.5 The robotic arm completes the operation corresponding to the instruction.

7. The multi-mode intelligent interactive robotic arm system according to claim 1, characterized in that: The third mode includes the following steps: S3.1 First, write a web page, then call the API on the web page to deploy the large model, and then enter the text you want to perform the operation in the dialog window on the web page; S3.2 After the large model recognizes the above text, it converts the text into data and corresponding control commands that the large model can recognize. S3.3 Determine whether the received data matches the set control commands; S3.4 The robotic arm will receive the G-code data corresponding to the control command, and the robotic arm will move according to the data; S3.5 The robotic arm completes the operation corresponding to the instruction.

8. The multi-mode intelligent interactive robotic arm system according to claim 1, characterized in that: The fourth mode includes the following steps: S4.

1. Use a YOLOv8-based deep learning algorithm to train labeled human eye images, and use the trained optimal model to detect the left and right eyes in the image. After distinguishing the left and right eyes, use the region of interest to outline the left and right eyes respectively. The size of the outline depends on the size of the previously labeled dataset. S4.

2. Using one of the OpenCV algorithms, namely grayscale conversion, Gaussian filtering, Hough circle detection, and coordinate fitting, the pupil coordinates of the two eyes are fitted to obtain a box with the combined pupil coordinates. The size of the box is fitted from the boxes of the left and right eyes, and the center coordinates of the box are the origin of the fitted coordinates. S4.

3. Transform the fitted coordinate system with the physical coordinate system. Using the data parameters and physical coordinate data parameters obtained above, align the two coordinate systems according to the mathematical method described in S1.

4. Then process the data into a G-code format that the robotic arm can recognize. S4.4, and add a region of interest (ROI) to the image, stipulating that the robotic arm will only move within the green box of the ROI to prevent the robotic arm from recognizing a position that it cannot move to; S4.5 If the human eye is not detected, the robotic arm will stop responding; S4.

6. Determine the state of the eyes based on the image. If the eyes are closed for more than one second, the air pump will start and the robotic arm will perform the object picking operation. If the eyes are closed again for more than one second, the air pump will stop running, thus completing the control method of combining the robotic arm and the eyes.

Citation Information

Patent Citations

  • Service robot control platform system and multimode intelligent interaction and intelligent behavior realizing method thereof

    CN102323817A

  • Gesture recognition algorithm and system based on color attention in AR operation environment

    CN119356524A