A collaborative robot gesture recognition system based on computer vision
Through the collaborative robot gesture recognition system based on computer vision, using image acquisition, model training and gesture recognition modules, the problem that the collaborative robot gesture recognition system is difficult to quickly locate and recognize operator gestures, and the improvement of high accuracy and human-computer interaction is achieved.
Patent Information
- Application Number
- CN202210638712.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-07
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-06-07
AI Technical Summary
In the prior art, it is difficult for collaborative robot gesture recognition systems to quickly locate and recognize operator gestures, resulting in timeliness and variability problems in gesture operation errors.
A collaborative robot gesture recognition system based on computer vision is adopted, including an image acquisition module, a model training module and a gesture recognition module. Gesture images are collected through the image acquisition module, training sets are constructed, and neural network models are constructed through the model training module to train to obtain gesture recognition model. The gesture recognition module recognizes the current gesture according to the recognition model and sends operation instructions.
The collaborative robot has realized the rapid gesture positioning and autonomous learning update recognition model, which improves the accuracy and human-computer interaction of gesture recognition.
Smart Images

Figure CN114863571B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot control, and particularly to a collaborative robot gesture recognition system based on computer vision. Background Art
[0002] With the development of technology, robots are gradually replacing monotonous, highly repetitive, and dangerous work in industrial production. Collaborative robots are also slowly infiltrating into various industrial fields to work with humans. With the development of computer vision technology, non-contact gesture control of robots is gradually replacing traditional contact control methods. However, due to the timeliness and variability of gestures, gesture operation errors occur from time to time.
[0003] Therefore, the present invention proposes a collaborative robot gesture recognition system based on computer vision for rapid gesture positioning and autonomous learning to update the recognition model. Summary of the Invention
[0004] The present invention provides a collaborative robot gesture recognition system based on computer vision, which can perform rapid gesture positioning, autonomously learn and update the recognition model, and continuously improve the gesture recognition accuracy of the robot.
[0005] The present invention provides a collaborative robot gesture recognition system based on computer vision, comprising:
[0006] An image acquisition module, configured to acquire gesture images of an operator, classify the gesture images, and construct a gesture recognition training set according to the classification results;
[0007] A model training module, configured to construct a neural network for gesture recognition, and train the neural network using the gesture recognition training set to obtain a gesture recognition model;
[0008] A gesture recognition module, configured to recognize the current gesture of the operator according to the gesture recognition model, and send a corresponding operation instruction to the robot according to the gesture recognition result.
[0009] Preferably, a collaborative robot gesture recognition system based on computer vision further comprises:
[0010] A human-computer interaction module connected to the robot control panel, configured to allow the operator to manually change the operation instruction through the robot control panel when the operation instruction sent by the gesture recognition module to the robot does not achieve the expected operation effect of the operator, wherein the robot preferentially executes the manual operation instruction;
[0011] A log storage module, configured to generate a manual operation log and record the misrecognized gestures of the robot when a manual operation instruction appears on the robot.
[0012] Preferably, the log storage module is further configured to record the entire gesture operation process of the operator, freeze the final state of each operation gesture, obtain a final state gesture diagram, and store it after adding a time tag and an operation instruction tag to the final state gesture diagram.
[0013] Preferably, the image acquisition module includes:
[0014] A gesture capture unit, configured to capture the hand movements of the operator through a shooting device on the robot to obtain a first hand image of the operator;
[0015] An image classification unit, configured to obtain gesture features on the hand image, classify according to the gesture features, and obtain multiple gesture image sets;
[0016] A training set construction unit, configured to respectively extract multiple second hand images from the multiple gesture image sets to generate a gesture recognition training set.
[0017] Preferably, the gesture capture unit further includes:
[0018] A capture subunit, configured to determine the hand movement trajectory of the operator based on gesture positioning frame positioning, and track and capture the hand movements of the operator according to the movement trajectory to obtain a first hand image;
[0019] A quality detection subunit, configured to detect the image quality of the first hand image, eliminate blurred images according to the detection results, and send the remaining first hand images to the image classification unit.
[0020] Preferably, the gesture capture unit further includes:
[0021] A positioning subunit, configured to establish a basic hand model according to the human hand bone model, obtain a large number of human images and a third hand image of the person in the human image;
[0022] Compare the differences between the third hand images, obtain the hand contour differences, and determine the amplitude of the human hand data differences according to the hand contour differences;
[0023] Determine the contour parameters of the hand model according to the amplitude of the human hand data differences, and adjust the basic hand model based on the contour parameters;
[0024] Locate the hand of the person in the human image, perform matte extraction to obtain a fourth hand image, and obtain the first hand data of the fourth hand image;
[0025] Horizontally compare the first hand data with the second hand data corresponding to the third hand image of the corresponding person to determine the first influence of the gesture change of the same person on the hand contour;
[0026] Vertically compare the first hand data with the second hand data corresponding to the third hand image of the corresponding person to determine the second influence on the hand contour;
[0027] Based on the first influence and the second influence, obtain a hand change loss function, and according to the hand change loss function and the contour parameters, obtain the hand change amplitude range;
[0028] According to the hand change amplitude range, optimize the adjusted basic hand model to obtain an optimized hand model;
[0029] Based on the optimized hand model, locate the hand positions of the large number of person images respectively to obtain a large number of marked person images;
[0030] According to the marked person images, obtain the first positional relationship between the person's body and the person's hand;
[0031] Obtain the first positional relationships of all the person images, establish a relationship comparison set, and based on the relationship comparison set, determine the hand positioning range of the person;
[0032] Based on the hand positioning range and according to the optimized hand model, use a gesture positioning frame to locate the operator's hand.
[0033] Preferably, the model training module includes:
[0034] A first construction unit for constructing a neural network for gesture recognition;
[0035] A training set processing unit for preprocessing the second hand image, eliminating the image background, and updating the gesture recognition training set;
[0036] A second construction unit for training the neural network through the gesture recognition training set to obtain a gesture recognition model.
[0037] Preferably, a collaborative robot gesture recognition system based on computer vision further includes: a judgment module for judging the reasons for gesture recognition errors when a manual operation instruction appears in the gesture recognition module, including:
[0038] An information acquisition unit for scaling the hand image corresponding to the gesture error operation instruction to the standard image size to obtain a gesture comparison map;
[0039] Determine multiple human hand key points according to the human hand bone model, search for the multiple human hand key points on the gesture comparison graph, and mark them;
[0040] According to the marking result, determine the finger positions on the gesture comparison graph to obtain the first position feature;
[0041] At the same time, connect the multiple human hand key points and obtain the connection characteristics of the connection lines;
[0042] Compare the connection characteristics with the bone connection characteristics on the human hand bone model to obtain the first direction feature;
[0043] Based on the first position feature and the first direction feature, determine the first gesture feature of the gesture on the gesture comparison graph;
[0044] According to the bone features and the manual operation log on the gesture comparison graph, obtain the first standard gesture image corresponding to the wrong operation instruction and the second standard gesture image corresponding to the manual operation instruction;
[0045] Obtain the second gesture feature and the third gesture feature corresponding to the first standard gesture image and the second standard gesture image respectively;
[0046] A comparison unit is used to compare the second gesture feature and the third gesture feature to obtain a feature difference, and according to the feature difference, determine the gesture difference degree between the gesture corresponding to the wrong operation instruction and the gesture corresponding to the manual operation instruction;
[0047] A first determination unit is used to, when the gesture difference degree is greater than a preset value, determine that the current gesture recognition error type is the first error type, compare the first gesture feature with the second gesture feature, if the first gesture feature is consistent with the second gesture feature, determine that the current gesture recognition error is an operator's gesture application error, and send a first error notification to the control module;
[0048] If the first gesture feature is inconsistent with the second gesture feature, compare the first gesture feature with the third gesture feature, if the first gesture feature is inconsistent with the third gesture feature, determine that the current gesture recognition error is an operator's gesture application error, and send a first error notification to the control module;
[0049] If the first gesture feature is consistent with the third gesture feature, after adding a mandatory label to the gesture comparison graph, send it to the data update unit and send a second error notification to the control module;
[0050] A second determination unit, configured to, when the gesture difference degree is less than or equal to a preset value, determine that the current gesture recognition error type is a second error type, and determine the gesture difference position according to the feature difference;
[0051] Mark on the gesture comparison graph according to the gesture difference position to obtain a marked image, send the marked image to the data update unit, and send a second error notification to the control module.
[0052] Preferably, the determination module further includes:
[0053] A data update unit, configured to, when receiving an image with a mandatory label, send the image with the mandatory label to the corresponding gesture image set, and reselect a gesture recognition training set;
[0054] When the image with the mandatory label is a marked image, mark all the images in the reselected gesture recognition training set according to the marks on the marked image.
[0055] Preferably, a collaborative robot gesture recognition system based on computer vision further includes a control module, including:
[0056] A receiving unit, configured to receive the error notification sent by the determination module;
[0057] A control unit, configured to, after receiving the first error notification, control the alarm unit to send an operation error warning to the operator;
[0058] The control unit is further configured to, after receiving the second error notification, train the gesture recognition model according to the gesture recognition training set reselected by the data update module to obtain a new gesture recognition model.
[0059] Other features and advantages of the present invention will be described in the following specification, and part of them will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings.
[0060] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0061] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0062] Figure 1 It is a schematic structural diagram of the collaborative robot gesture recognition system based on computer vision of the present invention;
[0063] Figure 2 This is a schematic structural diagram of the image acquisition module of the collaborative robot gesture recognition system based on computer vision of the present invention;
[0064] Figure 3 This is a schematic structural diagram of the model training module of the collaborative robot gesture recognition system based on computer vision of the present invention;
[0065] Figure 4 This is a schematic structural diagram of the judgment module of the collaborative robot gesture recognition system based on computer vision of the present invention;
[0066] Figure 5 This is a schematic structural diagram of the control module of the collaborative robot gesture recognition system based on computer vision of the present invention. Detailed implementation manners
[0067] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0068] Embodiment 1:
[0069] The present invention provides a collaborative robot gesture recognition system based on computer vision, as Figure 1 shown, including:
[0070] An image acquisition module, configured to acquire the gesture images of the operator, classify the gesture images, and construct a gesture recognition training set according to the classification results;
[0071] A model training module, configured to construct a neural network for gesture recognition, and train the neural network by using the gesture recognition training set to obtain a gesture recognition model;
[0072] A gesture recognition module, configured to recognize the current gesture of the operator according to the gesture recognition model, and send corresponding operation instructions to the robot according to the gesture recognition results.
[0073] Advantages of the above technical solution: The gesture recognition training set of the present invention trains the neural network to obtain a gesture recognition model, and uses this gesture recognition model to quickly locate and recognize the current gesture of the operator, control the robot to achieve the target operation, and enhance the human-computer interaction of the collaborative robot.
[0074] Embodiment 2:
[0075] On the basis of Embodiment 1, a collaborative robot gesture recognition system based on computer vision further includes:
[0076] A human-machine interaction module connected to the robot control panel, which is used for the operator to manually change the operation instruction through the robot control panel when the operation instruction sent by the gesture recognition module to the robot does not achieve the expected operation effect of the operator. Among them, the robot preferentially executes the manual operation instruction;
[0077] A log storage module, which is used to generate a manual operation log and record the misrecognized gestures of the robot when the robot receives a manual operation instruction.
[0078] Beneficial effects of the above technical solution: The present invention is provided with a human-machine interaction module, which is convenient for the operator to change and stop incorrect operations, manually control the robot to complete the target operation. At the same time, when the robot receives a manual operation instruction, a manual operation log is generated to record the misrecognized gestures of the robot, providing a basis for the autonomous learning and update of the gesture recognition model.
[0079] Embodiment 3:
[0080] Based on Embodiment 2, the log storage module is further used to record the entire gesture operation process of the operator, freeze the final state of each operation gesture, obtain the final state gesture diagram, and store it after adding a time tag and an operation instruction tag to the final state gesture diagram.
[0081] In this embodiment, the final state refers to the gesture when the operator's gesture operation is completed, and the final state gesture diagram refers to the gesture image when the operator's gesture operation is completed.
[0082] In this embodiment, the operation instruction tag refers to a tag added according to the gesture when the operator's gesture operation is completed and the corresponding operation instruction recognized based on the gesture recognition model.
[0083] Beneficial effects of the above technical solution: The present invention records the entire gesture operation process of the operator, which is beneficial for investigating responsibilities when the collaborative robot makes mistakes.
[0084] Embodiment 4:
[0085] Based on Embodiment 1, the image acquisition module, as Figure 2 shown, includes:
[0086] A gesture capture unit, which is used to capture the hand movements of the operator through the shooting device on the robot to obtain the first hand image of the operator;
[0087] An image classification unit, which is used to obtain the gesture features on the hand image, classify them according to the gesture features, and obtain multiple gesture image sets;
[0088] A training set construction unit, configured to respectively extract a plurality of second hand images from the plurality of gesture image sets to generate a gesture recognition training set.
[0089] In this embodiment, the first hand image refers to the hand image of the operator captured by the capturing device.
[0090] In this embodiment, the gesture features include finger positions, finger pointing directions, wrist pointing, etc.
[0091] In this embodiment, the gesture image set refers to a set of hand images of the same class placed according to the gesture features in the first hand images.
[0092] In this embodiment, the second hand image refers to a plurality of first hand images respectively extracted from different gesture image sets, and the extracted first hand images are placed in the same set, namely the gesture recognition training set.
[0093] Advantages of the above technical solution: The present invention classifies the collected first hand images according to gesture features to obtain a plurality of gesture image sets, and extracts a plurality of second gesture images from different gesture image sets respectively as the training set of the gesture recognition model, ensuring the diversity of training samples, which is beneficial to improving the accuracy of gesture recognition of the gesture recognition model.
[0094] Embodiment 5:
[0095] Based on Embodiment 4, the gesture capturing unit, as Figure 2 shown, further includes:
[0096] A capturing subunit, configured to determine the hand movement trajectory of the operator based on the gesture positioning frame and track and capture the hand movement of the operator according to the movement trajectory to obtain a first hand image;
[0097] A quality detection subunit, configured to detect the image quality of the first hand image, eliminate blurred images according to the detection results, and send the remaining first hand images to the image classification unit.
[0098] In this embodiment, the gesture positioning frame positioning refers to a rectangular frame for positioning the hand of the operator.
[0099] In this embodiment, the image quality is calculated as follows:
[0100]
[0101] Among them, φ represents the image quality of the first hand image; M represents the length of the first hand image, N represents the width of the first hand image, and M*N represents the size of the first hand image; S represents the length of the gesture positioning frame, T represents the width of the gesture positioning frame, and S*T represents the size of the image selected by the gesture positioning frame; f(i,j) represents the gray value of the pixel at the coordinate (i,j) on the image selected by the gesture positioning frame.
[0102] In this embodiment, a blurred image refers to a first hand image with relatively low quality.
[0103] Beneficial effects of the above technical solution: The present invention uses a gesture positioning frame for positioning the operator's hand, ensuring that the operator's hand movements are accurately captured during the gesture capture process. The image clarity of the collected first hand image is detected. According to the detection results, blurred images are removed, and the remaining first hand images are sent to the image classification unit to filter out invalid images, eliminating interference for the training of the hand neural network and avoiding the problem of low accuracy of the gesture recognition model caused by the low quality of the images used in the training.
[0104] Embodiment 6:
[0105] Based on Embodiment 4, the gesture capture unit, as Figure 2 shown, further includes:
[0106] A positioning sub-unit, configured to establish a basic hand model according to the human hand bone model, obtain a large number of human images and the third hand image of the person in the human image;
[0107] Compare the differences between the third hand images, obtain the hand contour differences, and determine the amplitude of the human hand data differences according to the hand contour differences;
[0108] Determine the contour parameters of the hand model according to the amplitude of the human hand data differences, and adjust the basic hand model based on the contour parameters;
[0109] Locate the hand of the person in the human image and perform matte extraction to obtain a fourth hand image, and obtain the first hand data of the fourth hand image;
[0110] Horizontally compare the first hand data with the second hand data corresponding to the third hand image of the corresponding person to determine the first influence of the gesture change of the same person on the hand contour;
[0111] Vertically compare the first hand data with the second hand data corresponding to the third hand image of the corresponding person to determine the second influence on the hand contour;
[0112] Based on the first influence and the second influence, obtain a hand variation loss function, and according to the hand variation loss function and the contour parameters, obtain the hand variation amplitude range;
[0113] According to the hand variation amplitude range, optimize the adjusted basic hand model to obtain an optimized hand model;
[0114] Based on the optimized hand model, respectively locate the hand positions of the large number of person images to obtain a large number of labeled person images;
[0115] According to the labeled person images, obtain the first position relationship between the person's body and the person's hand;
[0116] Obtain the first position relationship of all person images, establish a relationship comparison set, and based on the relationship comparison set, determine the hand positioning range;
[0117] Based on the hand positioning range and according to the optimized hand model, use a gesture positioning frame to locate the operator's hand.
[0118] In this embodiment, the third hand image refers to the hand image of a person in a person image, and this hand image is a directly captured hand image.
[0119] In this embodiment, the hand contour difference refers to the differences in the size, length of the palm, and length and thickness of the fingers of different persons; the contour parameters refer to the size, length of the palm of the hand model, and length and thickness of the fingers.
[0120] In this embodiment, the human hand data difference amplitude refers to the range of hand data of a normal human body, and this hand data refers to the size, length of the palm, and length and thickness of the fingers.
[0121] In this embodiment, the fourth hand image refers to the hand image cropped from the person's body.
[0122] In this embodiment, the first hand data refers to the hand data corresponding to the fourth hand image; the second hand data refers to the hand data corresponding to the third hand image.
[0123] In this embodiment, the basic hand model refers to a hand model established based on the human hand bones without any processing.
[0124] In this embodiment, the first influence refers to the influence of the gesture change of the same person on the hand data.
[0125] In this embodiment, the second influence refers to the influence of the gesture change of different persons on the hand data, and this influence may be caused by different shooting angles due to differences in the heights of different tasks.
[0126] In this embodiment, the hand change loss function refers to the change function of hand data directly in hand image acquisition.
[0127] In this embodiment, the hand change amplitude range refers to the range of possible changes in hand data during gesture acquisition.
[0128] In this embodiment, the optimized hand model refers to the optimized basic hand model.
[0129] In this embodiment, the marked person image refers to the person image in which the position of the person's hand is calibrated using the optimized hand model.
[0130] In this embodiment, the first position relationship refers to the relationship between the person's hand and the person's body in the marked person image; the relationship comparison set refers to the set where the first position relationship data is placed.
[0131] In this embodiment, the person's hand positioning range refers to the range where the person's hand can appear.
[0132] Advantages of the above technical solution: According to the human hand bone model, the present invention establishes a basic hand model, and then optimizes the basic model according to a series of data differences between the person's photo and the directly captured image of the person's hand; at the same time, according to the comparison of the first position relationship between the person's body and the person's hand, the person's hand positioning range is determined, and within the person's hand positioning range, according to the optimized hand model, a gesture positioning frame is used to quickly locate the operator's hand, realizing the fast and accurate acquisition of the person's gesture image.
[0133] Embodiment 7:
[0134] Based on Embodiment 1, the model training module, as Figure 3 shown, includes:
[0135] The first construction unit is used to construct a neural network for gesture recognition;
[0136] The training set processing unit is used to preprocess the second hand image, eliminate the image background, and update the gesture recognition training set;
[0137] The second construction unit is used to train the neural network through the gesture recognition training set to obtain a gesture recognition model.
[0138] Advantages of the above technical solution: Before training the neural network, the present invention eliminates the image background of the images in the training set, which is beneficial to reducing interference and improving the fault tolerance rate of the system's gesture recognition.
[0139] Embodiment 8:
[0140] Based on Embodiment 1, a collaborative robot gesture recognition system based on computer vision further includes: a judgment module for judging the reasons for gesture recognition errors when a manual operation instruction appears in the gesture recognition module, such as Figure 4 shown, including:
[0141] An information acquisition unit for scaling the hand image corresponding to the gesture error operation instruction to the standard image size to obtain a gesture comparison map;
[0142] According to the human hand bone model, determine multiple human hand key points, and find and mark the multiple human hand key points on the gesture comparison map;
[0143] According to the marking result, determine the finger positions on the gesture comparison map to obtain a first position feature;
[0144] At the same time, connect the multiple human hand key points and obtain the connection characteristics of the connection lines;
[0145] Compare the connection characteristics with the bone connection characteristics on the human hand bone model to obtain a first direction feature;
[0146] Based on the first position feature and the first direction feature, determine the first gesture feature of the gesture on the gesture comparison map;
[0147] According to the bone features and the manual operation log on the gesture comparison map, obtain a first standard gesture image corresponding to the error operation instruction and a second standard gesture image corresponding to the manual operation instruction;
[0148] Respectively obtain a second gesture feature and a third gesture feature corresponding to the first standard gesture image and the second standard gesture image;
[0149] A comparison unit for comparing the second gesture feature and the third gesture feature to obtain a feature difference, and determining the gesture difference degree between the gesture corresponding to the error operation instruction and the gesture corresponding to the manual operation instruction according to the feature difference;
[0150] A first determination unit for, when the gesture difference degree is greater than a preset value, judging that the current gesture recognition error type is a first error type, comparing the first gesture feature with the second gesture feature, and if the first gesture feature is consistent with the second gesture feature, determining that the current gesture recognition error is an operator's gesture application error and sending a first error notification to the control module;
[0151] If the first gesture feature is inconsistent with the second gesture feature, compare the first gesture feature with the third gesture feature. If the first gesture feature is inconsistent with the third gesture feature, determine that the current gesture recognition error is an incorrect gesture application by the operator, and send a first error notification to the control module;
[0152] If the first gesture feature is consistent with the third gesture feature, after adding a mandatory label to the gesture comparison graph, send it to the data update unit, and send a second error notification to the control module;
[0153] The second determination unit is configured to, when the gesture difference degree is less than or equal to a preset value, determine that the current gesture recognition error type is the second error type, and determine the gesture difference position according to the feature difference;
[0154] According to the gesture difference position, perform marking on the gesture comparison graph to obtain a marked image, send the marked image to the data update unit, and send a second error notification to the control module.
[0155] In this embodiment, the standard image refers to the image of each gesture stored in the system for gesture judgment.
[0156] In this embodiment, the gesture comparison graph refers to the corresponding hand image when a gesture recognition error occurs.
[0157] In this embodiment, the human hand key points refer to the positioning points that can calibrate the finger positions. For example, the positioning points located at the finger joints, fingertips, palm center, and palm edge.
[0158] In this embodiment, the first position feature refers to the position feature of the fingers. For example, on the gesture comparison graph, the finger positioning points of the index finger and the middle finger do not change significantly, but the positioning points of other fingers coincide with some palm point positioning points, indicating that other fingers are in a bent state.
[0159] In this embodiment, the connection characteristic refers to the direction of the connection line after connecting the human hand key points on the gesture comparison graph, where the fingertip pointing direction is the positive direction.
[0160] In this embodiment, the first direction feature refers to the fingers and the direction of the fingers on the gesture comparison graph.
[0161] In this embodiment, the first gesture feature includes the first position feature and the first direction feature.
[0162] In this embodiment, the first standard gesture image refers to the image of the gesture corresponding to the incorrect operation instruction stored in the system for gesture judgment; the second standard gesture image refers to the image of the gesture corresponding to the manual operation instruction stored in the system for gesture judgment.
[0163] In this embodiment, the second gesture feature refers to features such as the position of the hand on the first standard gesture image and the directions of the fingers and wrist; the third gesture feature refers to features such as the position of the hand on the second standard gesture image and the directions of the fingers and wrist.
[0164] In this embodiment, the feature difference refers to the positions of the hands on the first standard gesture image and the second standard gesture image and the directions of the fingers and wrist after comparison, and the positions that are the same.
[0165] In this embodiment, the gesture difference degree refers to the difference degree after comparing the gesture corresponding to the wrong operation instruction and the gesture corresponding to the manual operation instruction, and the specific calculation is as follows:
[0166]
[0167] Where ω represents the difference degree after comparing the gesture corresponding to the wrong operation instruction and the gesture corresponding to the manual operation instruction, that is, the gesture difference degree; α1 refers to the difference weight ratio of the fingers, and the value range is (0.6, 0.8); α2 refers to the difference weight ratio of the wrist, and the value range is (0.2, 0.4), and α1 + α2 = 1; A i refers to the difference angle of the i-th finger on the gestures corresponding to the wrong operation instruction and the manual operation instruction; M i refers to the angle of the i-th finger on the gesture corresponding to the wrong operation instruction on the first standard gesture image; B refers to the difference angle at the wrist of the gestures corresponding to the wrong operation instruction and the manual operation instruction; B0 refers to the angle of the wrist on the gesture corresponding to the wrong operation instruction on the first standard gesture image.
[0168] Where the wrist directions of the gestures corresponding to the wrong operation instruction and the manual operation instruction are completely different means that the directions of the gestures corresponding to the wrong operation instruction and the manual operation instruction are completely opposite.
[0169] In this embodiment, the first error type refers to a gesture recognition error where the gesture difference degree is greater than the preset value (that is, the difference degree after comparing the gesture corresponding to the wrong operation instruction and the gesture corresponding to the manual operation instruction is relatively large); the second error type refers to a gesture recognition error where the gesture difference degree is less than or equal to the preset value (that is, the difference degree after comparing the gesture corresponding to the wrong operation instruction and the gesture corresponding to the manual operation instruction is relatively small).
[0170] In this embodiment, the first error notification refers to a notification sent to the control module when the gesture recognition error is an error in the operator's gesture application.
[0171] In this embodiment, the second error notification is a notification sent to the control module when the gesture recognition error is caused by a system recognition error.
[0172] In this embodiment, a mandatory label refers to a label that, when a gesture recognition error is caused by a system recognition error, is added to the hand image that causes the error, i.e., the gesture comparison image, and this image must be selected as an image in the gesture recognition training set.
[0173] In this embodiment, an annotated image refers to a gesture comparison image in which, when the gesture recognition error is of the second error type, the positions where the gesture comparison image is different from the first standard gesture image and the second standard gesture image are marked.
[0174] Advantages of the above technical solution: By carefully positioning the gesture on the gesture comparison image, the present invention extracts the gesture features on the gesture comparison image as detailed as possible. At the same time, based on the bone features on the gesture comparison image and the manual operation log, it determines the hand (e.g., the right hand) used by the operator during operation, and then obtains the standard gesture images corresponding to the incorrect operation instruction and the hand corresponding to the manual instruction, providing accurate comparison data for subsequent gesture comparison.
[0175] After comparing the gesture features of the gesture corresponding to the incorrect operation instruction and the gesture corresponding to the manual operation instruction, according to the degree of difference between the two, the error type is initially determined, and then compared with the gesture features on the gesture comparison image to determine whether it is a human error or a system error, which is beneficial to timely discovering the drawbacks of system recognition.
[0176] Embodiment 9:
[0177] Based on Embodiment 1, the judgment module, as Figure 4 shown, further includes:
[0178] A data update unit, configured to, when receiving an image with a mandatory label, send the image with the mandatory label to the corresponding gesture image set and reselect the gesture recognition training set;
[0179] When the image with the mandatory label is an annotated image, annotate all the images in the reselected gesture recognition training set according to the annotation on the annotated image.
[0180] Advantages of the above technical solution: The present invention updates the gesture image set and reselects the gesture recognition training set according to the mandatory image sent to the data update unit after analyzing the reason for the gesture recognition error, providing a basis for the training of the gesture recognition model. At the same time, when the image with the mandatory label is an annotated image, all the images in the reselected gesture recognition training set are annotated according to the annotation on the annotated image. The annotated training set emphasizes image details, which is beneficial to improving the accuracy of the new gesture recognition model.
[0181] Embodiment 10:
[0182] Based on Embodiment 1, a collaborative robot gesture recognition system based on computer vision further includes a control module, as Figure 5 shown, including:
[0183] A receiving unit, configured to receive error notifications sent by the judgment module;
[0184] A control unit, configured to control the alarm unit to send an operation error warning to the operator after receiving the first error notification;
[0185] The control unit is further configured to, after receiving the second error notification, train the gesture recognition model according to the gesture recognition training set newly selected by the data update module to obtain a new gesture recognition model.
[0186] In this embodiment, the error notification includes a first error notification and a second error notification.
[0187] In this embodiment, the operation error warning is a gesture application error reminder sent by the alarm unit to the user.
[0188] Beneficial effects of the above technical solution: According to the error notification sent by the judgment module, the present invention determines the reason for the gesture recognition error, and when it is determined that the reason for the gesture error is a defect of the system itself, trains the gesture recognition model according to the gesture recognition training set newly selected by the data update module to obtain a new gesture recognition model, realizing the autonomous learning of the collaborative robot gesture recognition system.
[0189] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A collaborative robot gesture recognition system based on computer vision, characterized in that, include: An image acquisition module is used to acquire gesture images of operators, classify the gesture images, and construct a gesture recognition training set based on the classification results; A model training module, used to construct a neural network for gesture recognition, and train the neural network using the gesture recognition training set to obtain a gesture recognition model; A gesture recognition module, used to recognize the operator's current gesture according to the gesture recognition model, and send corresponding operation instructions to the robot according to the gesture recognition result; Wherein, the image acquisition module comprises: A gesture capturing unit, used to capture the hand movements of the operator through a capturing device on the robot to obtain a first hand image of the operator; An image classification unit, used to obtain gesture features on the hand image, classify the gestures according to the gesture features, and obtain multiple gesture image sets; A training set construction unit, used to extract a plurality of second hand images from the plurality of gesture image sets respectively to generate a gesture recognition training set; Wherein, the gesture capturing unit comprises: A capture subunit, used to determine the hand motion trajectory of the operator based on the gesture positioning frame positioning, and track and shoot the hand motion of the operator according to the motion trajectory to obtain a first hand image; a quality detection subunit, configured to detect the image quality of the first hand image, remove blurred images according to the detection result, and send the remaining first hand images to the image classification unit; A positioning subunit, used to establish a basic hand model according to a human hand skeleton model, and obtain a large number of person images and a third hand image of the person in the person images; Comparing the differences between the third hand images, obtaining hand contour differences, and determining the difference amplitude of human hand data according to the hand contour differences; Determining contour parameters of a hand model according to the difference amplitude of the human hand data, and adjusting the basic hand model based on the contour parameters; Positioning the hand of the person on the person image, performing cutout to obtain a fourth hand image, and acquiring first hand data of the fourth hand image; Comparing the first hand data with the second hand data corresponding to the third hand image of the corresponding person horizontally to determine a first impact of the hand gesture change of the same person on the hand contour; longitudinally comparing the first hand data with second hand data corresponding to a third hand image of the corresponding person to determine a second influence of the hand contour; Based on the first influence and the second influence, a hand variation loss function is obtained, and according to the hand variation loss function and the contour parameter, a hand variation amplitude range is obtained; According to the hand change amplitude range, optimizing the adjusted basic hand model to obtain an optimized hand model; Based on the optimized hand model, respectively locate the hand positions of the large number of person images to obtain a large number of marked person images; According to the marked person image, a first position relationship between the person's body and the person's hand is obtained; Obtaining first positional relationships of all person images, establishing a relationship comparison set, and determining a person hand positioning range based on the relationship comparison set; Based on the hand positioning range and according to the optimized hand model, a gesture positioning frame is used to position the operator's hand.
2. The collaborative robot gesture recognition system based on computer vision according to claim 1, characterized in that, It further includes: A human-machine interaction module connected to the robot control panel, which is used for the operator to manually change the operation instruction through the robot control panel when the operation instruction sent by the gesture recognition module to the robot does not achieve the expected operation effect of the operator. Among them, the robot preferentially executes the manual operation instruction; A log storage module, which is used to generate a manual operation log and record the misrecognized gestures of the robot when the robot receives a manual operation instruction.
3. The collaborative robot gesture recognition system based on computer vision according to claim 2, characterized in that: The log storage module is further used to record the entire gesture operation process of the operator, freeze the final state of each operation gesture, obtain the final state gesture diagram, and store it after adding a time tag and an operation instruction tag to the final state gesture diagram.
4. The collaborative robot gesture recognition system based on computer vision according to claim 1, characterized in that, The model training module includes: A first construction unit, which is used to construct a neural network for gesture recognition; A training set processing unit, which is used to preprocess the second hand image, eliminate the image background, and update the gesture recognition training set; A second construction unit, which is used to train the neural network through the gesture recognition training set to obtain a gesture recognition model.
5. A collaborative robot gesture recognition system based on computer vision according to claim 1, characterized in that, It further includes: A judgment module, which is used to judge the reason for gesture recognition error when the gesture recognition module receives a manual operation instruction, including: An information acquisition unit, which is used to scale the hand image corresponding to the gesture error operation instruction to the standard image size to obtain a gesture comparison diagram; According to the human hand bone model, multiple human hand key points are determined, and the multiple human hand key points are searched for and marked on the gesture comparison diagram; According to the marking result, the finger positions on the gesture comparison diagram are determined to obtain a first position feature; At the same time, the multiple human hand key points are connected, and the connection characteristics of the connection lines are obtained; The connection characteristics are compared with the bone connection characteristics on the human hand bone model to obtain a first direction feature; Based on the first position feature and the first direction feature, the first gesture feature of the gesture on the gesture comparison diagram is determined; According to the bone features on the gesture comparison diagram and the manual operation log, a first standard gesture image corresponding to the wrong operation instruction and a second standard gesture image corresponding to the manual operation instruction are obtained; The second gesture feature and the third gesture feature corresponding to the first standard gesture image and the second standard gesture image are respectively obtained; A comparison unit, which is used to compare the second gesture feature and the third gesture feature to obtain a feature difference, and determine the gesture difference degree between the gesture corresponding to the wrong operation instruction and the gesture corresponding to the manual operation instruction according to the feature difference; A first determination unit, which is used to judge that the current gesture recognition error type is the first error type when the gesture difference degree is greater than a preset value, compare the first gesture feature with the second gesture feature, and if the first gesture feature is consistent with the second gesture feature, judge that the current gesture recognition error is an operator's gesture application error, and send a first error notification to the control module; If the first gesture feature is inconsistent with the second gesture feature, compare the first gesture feature with the third gesture feature. If the first gesture feature is inconsistent with the third gesture feature, determine that the current gesture recognition error is an operator gesture application error, and send a first error notification to the control module; If the first gesture feature is consistent with the third gesture feature, after adding a mandatory label to the gesture comparison diagram, send it to the data update unit, and send a second error notification to the control module; A second determination unit, configured to, when the gesture difference degree is less than or equal to a preset value, determine that the current gesture recognition error type is a second error type, and determine the gesture difference position according to the feature difference; According to the gesture difference position, perform annotation on the gesture comparison diagram to obtain an annotated image, and send the annotated image to the data update unit, and send a second error notification to the control module.
6. A collaborative robot gesture recognition system based on computer vision according to claim 5, characterized in that, The judgment module further includes: A data update unit, configured to, when receiving an image with a mandatory label, send the image with the mandatory label to the corresponding gesture image set, and re-select a gesture recognition training set; When the image with the mandatory label is an annotated image, perform annotation on all images in the re-selected gesture recognition training set according to the annotation on the annotated image.
7. A collaborative robot gesture recognition system based on computer vision according to claim 1, characterized in that, It further includes: A control module, including: A receiving unit, configured to receive error notifications sent by the judgment module; A control unit, configured to, after receiving the first error notification, control the alarm unit to send an operation error warning to the operator; The control unit is further configured to, after receiving the second error notification, train the gesture recognition model according to the gesture recognition training set re-selected by the data update module to obtain a new gesture recognition model.
Citation Information
Patent Citations
Gesture recognition method based on deep learning
CN111104820A