Gesture interaction method and device, electronic equipment, chip and medium
By detecting hand areas and key points in vehicle-computer interaction, the problem of low accuracy and recall of finger gesture recognition is solved, and higher recognition accuracy and user experience is achieved.
Patent Information
- Application Number
- CN202311453778.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-06
AI Technical Summary
In car-computer interaction, the accuracy and recall of finger gesture recognition in the image is low, mainly due to the difference in the skin tone of the hand and the length of the finger.
By detecting the hand area, determining the hand key points, and performing gesture recognition based on these key points, the search space for gesture recognition is narrowed and the accuracy of recognition is improved.
It improves the accuracy and recall rate of finger pointing recognition, achieves more accurate and reliable gesture interaction, and improves the user's interactive experience in the vehicle cockpit.
Smart Images

Figure CN119937767A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of human-computer interaction, and in particular to a gesture interaction method, device, electronic device, chip and medium. Background Art
[0002] The vehicle-computer interaction in the vehicle cabin is carried out through keyboard, touch, voice and other means. In order to obtain a more convenient and intelligent interactive experience, relevant technologies have proposed natural human-computer interaction through finger pointing. Specifically, based on finger pointing image processing, finger pointing classification is performed through Convolutional Neural Network (CNN). However, in most scenarios, the background area regarded as noise in the finger pointing image occupies a large part. At the same time, it is affected by the difference in hand skin color and finger length, resulting in low accuracy and recall rate of finger pointing recognition. Summary of the invention
[0003] The present disclosure provides a gesture interaction method, device, electronic device, chip and medium to solve the problems of low accuracy and low recall rate of finger gesture recognition in finger pointing images in related technologies. By detecting the hand area, determining the hand key points, and performing gesture recognition based on the hand key points, the accuracy and recall rate of finger pointing recognition are improved.
[0004] A first aspect of the present disclosure provides a gesture interaction method, which is applied to a vehicle computer system. The method includes:
[0005] Acquire a hand area image;
[0006] Based on the hand region image, the hand key points are detected by a hand key point detector;
[0007] Based on the hand key points, the finger pointing category is determined by the finger pointing classifier;
[0008] The vehicle is controlled to execute a first function corresponding to the finger pointing category.
[0009] In one embodiment of the present disclosure, acquiring a hand region image includes:
[0010] Acquire target images;
[0011] According to the target image, a hand region is detected by a hand region detector;
[0012] Determine a first confidence level of the hand region based on a distance between a first coordinate point in the hand region and a second coordinate point in the target image, wherein the distance is inversely proportional to the first confidence level;
[0013] Determining whether the first confidence level is greater than a preset confidence level threshold;
[0014] If the first confidence is greater than a preset confidence threshold, the hand region is used as the hand region image;
[0015] If the first confidence level is less than or equal to a preset confidence threshold, a hand region image is acquired based on a next frame image of the target image.
[0016] In one embodiment of the present disclosure, the process of determining the second coordinate point includes:
[0017] According to the target application currently running in the vehicle system and the correspondence between the application and the reference point position saved in advance, the target reference point corresponding to the target application is determined, and the target reference point is used as the second coordinate point.
[0018] In one embodiment of the present disclosure, based on the hand region image, detecting the hand key points by using a hand key point detector includes:
[0019] Extracting a first feature map of the hand region image through a first branch of a residual neural network of a first feature extraction network;
[0020] Extracting a second feature map of the hand region graphic through a second branch of the residual neural network of the first feature extraction network, wherein the number of channels of the convolutional layers of the first branch and the second branch are both half of the convolutional layer of the residual neural network, the activation layers of the first branch and the second branch are different, and / or the pooling layers of the first branch and the second branch are different;
[0021] fusing the first feature map and the second feature map into a third feature map;
[0022] Based on the third feature map of the hand region image, a second number of hand key point coordinates are predicted through a fully connected layer of the first feature extraction network.
[0023] In one embodiment of the present disclosure, determining the finger pointing category by a finger pointing classifier based on the hand key points includes:
[0024] Using the second number of hand key point coordinates as input of the finger pointing classifier, performing feature analysis on the second number of hand key point coordinates based on the multi-layer fully connected layer of the finger pointing classifier, and determining a prediction probability value corresponding to each category in the finger pointing category;
[0025] The category with the largest predicted probability value is taken as the category that the finger is pointing to.
[0026] A second aspect of the present disclosure provides a gesture interaction device, which is applied to a vehicle system. The device includes:
[0027] An acquisition module, used for acquiring a hand area image;
[0028] A key point detection module, used for detecting hand key points through a hand key point detector based on the hand region image;
[0029] A pointing classification module, used to determine the finger pointing category through a finger pointing classifier based on the hand key points;
[0030] The control module is used to control the vehicle to execute a first function corresponding to the finger pointing category.
[0031] The third aspect embodiment of the present disclosure proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the methods in the first aspect embodiment of the present disclosure.
[0032] The fourth aspect embodiment of the present disclosure proposes a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to enable a computer to execute the method in the first aspect embodiment of the present disclosure.
[0033] The fifth aspect embodiment of the present disclosure proposes a computer program product, characterized in that it includes a computer program, and when the computer program is executed by a processor, it implements any method in the first aspect embodiment of the present disclosure.
[0034] The sixth aspect embodiment of the present disclosure proposes a chip, characterized in that it includes one or more interface circuits and one or more processors; the interface circuit is used to receive signals from the memory of the electronic device and send signals to the processor, the signals include computer instructions stored in the memory, and when the processor executes the computer instructions, the electronic device executes any one of the methods in the first aspect embodiment of the present disclosure.
[0035] In summary, according to the gesture interaction method proposed in the present disclosure, the hand area image is obtained, which reduces the search space of gesture recognition; based on the hand area image, the hand key points are detected by the hand key point detector, and the coordinate information of the hand key points is obtained, providing accurate and reliable data for gesture recognition; based on the hand key points, the finger pointing category is determined by the finger pointing classifier, and an accurate recognition result of the finger pointing category is obtained; the vehicle is controlled to execute the first function corresponding to the finger pointing category, realizing the quick control of the vehicle through natural human-computer interaction of finger pointing. Let the user get a better human-computer interaction experience in the vehicle cabin.
[0036] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0038] Figure 1 is a flow chart of a gesture interaction method according to an embodiment of the present disclosure;
[0039] Figure 2 A flow chart of acquiring a hand area image according to an embodiment of the present disclosure;
[0040] Figure 3 A flowchart of an embodiment of the present disclosure for determining a first confidence level of a hand region based on a distance between a first coordinate point in the hand region and a second coordinate point in a target image;
[0041] Figure 4 A flowchart of detecting hand key points by a hand key point detector based on a hand region image according to an embodiment of the present disclosure;
[0042] Figure 5 A flowchart of determining a finger pointing category through a finger pointing classifier based on hand key points according to an embodiment of the present disclosure;
[0043] Figure 6 is a structural schematic diagram of a gesture interaction device according to an embodiment of the present disclosure;
[0044] Figure 7 It is a block diagram of an electronic device for implementing the gesture interaction method disclosed in the present invention according to an exemplary embodiment. DETAILED DESCRIPTION
[0045] Embodiments of the present disclosure are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout identify the same or similar originals or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0046] The present disclosure aims to solve the problems of low recognition accuracy and low recall rate of finger pointing in gesture interaction of vehicle computers, extract hand area images through target detection algorithm; detect hand key point coordinates using hand key point detection algorithm; and recognize finger pointing gestures based on hand key point coordinates.
[0047] The method proposed in the present disclosure is applied to a vehicle computer system, and its application has a wide range of scenarios. The vehicle computer is controlled through human gesture interaction and can be used in the following fields:
[0048] 1. Car driving assistance: Through gesture interaction, the driver can more conveniently control the vehicle's wipers, air conditioning, audio, navigation and other functions to improve driving safety and comfort.
[0049] 2. Car sales: Through gesture interaction, sales staff can more intuitively show customers the various functions and features of the vehicle, thereby improving sales efficiency.
[0050] 3. Car service: Through gesture interaction, car owners can more conveniently query vehicle status information, make appointments for repairs and maintenance, and other services.
[0051] There is no limitation on the application scenario in the embodiments of the present disclosure.
[0052] The gesture interaction method provided by the present application is described in detail below with reference to the accompanying drawings.
[0053] Figure 1 is a flow chart of a gesture interaction method according to an embodiment of the present disclosure. The method is applied to a vehicle system, based on Figure 1 In the illustrated embodiment, the gesture interaction method includes:
[0054] Step 101, obtaining a hand area image.
[0055] In this embodiment, the hand area image is an image that includes the area where the hand is located. The hand area image can be determined by detecting the hand area using a target detection algorithm, or by segmenting the hand area from the original image or processed image captured by the image sensor using a target segmentation algorithm. The image determined by the rectangular area of the minimum bounding box where the hand area is located is the hand area image.
[0056] Step 102: Based on the hand region image, detect the hand key points by using a hand key point detector.
[0057] In this embodiment, the hand key point detector is a hand key point detection algorithm, which uses a feature extraction neural network to extract the hand key points in the hand area image based on the determined hand area image to obtain the position information of the hand key points.
[0058] Step 103, based on the hand key points, determine the finger pointing category through a finger pointing classifier.
[0059] In this embodiment, the finger pointing classifier is a classification algorithm, which can be implemented by a machine learning algorithm or a neural network algorithm. According to the key points of the hand, the finger pointing category can be determined by classifying the finger pointing category using a classification algorithm.
[0060] In one implementation of this embodiment, the finger pointing classifier includes but is not limited to a multi-layer perceptron (MLP), a support vector machine (SVM), and a logistic regression (LR).
[0061] Step 104 , controlling the vehicle to execute a first function corresponding to the finger pointing category.
[0062] In this embodiment, the vehicle is controlled to execute a function corresponding to the classification result according to the classification result of the identified finger pointing category. For example, by pointing the thumb upward, the volume of the vehicle computer is adjusted to be louder; by pointing the thumb downward, the volume of the vehicle computer is adjusted to be lower; or by pointing the thumb to the left, the vehicle air conditioner is turned on, and by pointing the thumb to the right, the vehicle air conditioner is turned off.
[0063] In summary, according to the gesture interaction method proposed in the present disclosure, the hand area image is obtained, which reduces the search space of gesture recognition; based on the hand area image, the hand key points are detected by the hand key point detector, and the coordinate information of the hand key points is obtained, providing accurate and reliable data for gesture recognition; based on the hand key points, the finger pointing category is determined by the finger pointing classifier, and the recognition result of the finger pointing category with high accuracy and high recall rate is obtained; the vehicle is controlled to execute the first function corresponding to the finger pointing category, realizing the quick control of the vehicle through the natural human-computer interaction of finger pointing. Let the user get a better human-computer interaction experience in the vehicle cabin.
[0064] Figure 2 FIG. 1 is a flow chart of obtaining a hand area image according to an embodiment of the present disclosure. Figure 2 right Figure 1 Step 101 is further described based on Figure 2 The embodiment shown comprises the following steps:
[0065] Step 201, capturing a target image.
[0066] In this embodiment, the target image can be a RAW original image, an RGB color image or a grayscale image captured by any one of the two types of image sensors, such as a charge coupled device (CCD) and a metal oxide semiconductor (CMOS), or a depth image captured by a depth camera such as a structured light, TOF or binocular.
[0067] Step 202: Detect the hand region using a hand region detector according to the target image.
[0068] In this embodiment, the hand region detector is a target detection algorithm, and the hand region is a rectangular region including the left hand or the right hand. After the target image is acquired, the hand region is detected based on the target image data by the hand region detector, that is, the target detection algorithm.
[0069] In one implementation, the hand region detector includes two-stage target detection, single-stage target detection, and transformer-based target detection. Two-stage target detection includes Faster R-CNN and Mask R-CNN; single-stage target detection includes YOLO and SSD; transformer-based target detection includes DETR and TRDET.
[0070] Step 203, determining a first confidence level of the hand region based on a distance between a first coordinate point in the hand region and a second coordinate point in the target image, wherein the distance is inversely proportional to the first confidence level.
[0071] In this embodiment, the first coordinate point is the position of the hand area in the target image coordinate system, which can be the center position of the hand area, or the origin position of the upper left corner of the hand area, preferably the center position. The coordinates of the first coordinate point are not limited here. The second coordinate point is the position of a pixel point in the target image, which can be the center position of the target image, or the origin position of the upper left corner of the target image, preferably the center position, and the coordinates of the second coordinate point are not limited here. Through the hand area detected in the target image and the target image, the first coordinate point in the hand area and the second coordinate point of the target image are determined, and the first confidence confidence of the hand area is calculated according to the distance between the first coordinate point and the second coordinate point according to the preset relationship formula, wherein the first confidence reflects the credibility of the hand area detection result, and the distance between the first coordinate point and the second coordinate point is inversely proportional to the first confidence. In one embodiment, the first coordinate point is the center point of the hand area, and the second coordinate point is the center point of the target image. The farther the distance between the first coordinate point and the second coordinate point is, that is, the more the hand area deviates from the central area of the target image, the lower the first confidence, indicating that the reliability of the hand area data is lower. On the contrary, the closer the distance between the first coordinate point and the second coordinate point is, that is, the closer the hand region is to the center of the target image, the higher the reliability of the hand region data is.
[0072] Step 204: determine whether the first confidence level is greater than a preset confidence level threshold.
[0073] In this embodiment, the first confidence corresponding to the calculated hand region is compared with a preset confidence threshold to determine whether the detected hand region is reliable.
[0074] Step 205: If the first confidence level is greater than a preset confidence threshold, the hand region is used as a hand region image.
[0075] In this embodiment, if the first confidence is greater than the preset confidence threshold, it means that the hand state has been recorded relatively clearly and accurately in the detected hand area, so the hand area can be directly cropped from the target image as a hand area image.
[0076] Step 206: If the first confidence level is less than or equal to a preset confidence threshold, a hand region image is acquired based on the next frame image of the target image.
[0077] In this embodiment, if the first confidence is less than or equal to the preset confidence threshold, it means that the hand state cannot be clearly or accurately reflected in the detected hand area. For example, the detected hand area is in the boundary area of the target image, which may cause the subsequent key point detection algorithm to be unable to accurately detect the gesture key points and make it difficult to complete the recognition of finger pointing. Therefore, according to the small current value of the first confidence, the processing of the subsequent algorithm can be skipped, the hand area can be detected from the next frame of the target image, and the hand area can be cropped as the hand area image.
[0078] In this embodiment, the hand area image is determined by detecting the hand area in the captured image, providing target data for further gesture key point recognition. By comparing the first confidence of the hand area with the preset confidence threshold, it is determined whether the detected hand area accurately and clearly reflects the hand state. If the confidence is higher than the preset confidence threshold, the hand area image is obtained with the current clear and accurate hand area, otherwise, the hand area is detected from the next frame image. Thus, the input data for hand key point recognition is determined, which reduces the search space for key point detection and improves the detection speed. Among them, the first confidence of the hand area directly affects the obtained hand area image.
[0079] Figure 3 The present invention is a flowchart of determining a first confidence level of a hand region based on a distance between a first coordinate point in the hand region and a second coordinate point in a target image according to an embodiment of the present invention. Figure 3 right Figure 2 Step 203 is described in detail. Figure 3 The steps include:
[0080] Step 301 , according to the target application currently running in the vehicle system and the correspondence between the application and the reference point position saved in advance, determine the target reference point corresponding to the target application, and use the target reference point as the second coordinate point.
[0081] In this embodiment, the target application is the application currently running in the vehicle system, which can be a system application or a third-party application. The reference point is the location of a pixel point selected in the target image, which is used to provide a position reference for the detected hand area, and the second coordinate point is the reference point. Among them, the target application is related to the second coordinate point. When the algorithm runs on the vehicle system, according to the mapping table of gestures and vehicle functions that has been set, when a gesture that already exists in the mapping table is detected, the corresponding function can be triggered. When using gestures in the cockpit, if the gesture is detected, the currently running application is checked to ensure that the triggering of the function of the application is more suitable for the user's operating posture. For example, the preset gesture function is to turn on or off the seat heating by a right-hand gesture. Since the right-hand hand area in the target image is on the right side, the second coordinate point (reference point) is currently at the center right position of the target image, for example, (0.75*W, 0.5*H) as the coordinate of the second coordinate point. Through this method, the gestures made by the user more conveniently and comfortably can be recognized when the operating space in the car is limited.
[0082] In one implementation of this embodiment, after the hand region is detected by a hand region detection algorithm, such as SSD or YOLO, a quaternion value (l, c, x, y) is obtained, including the horizontal coordinate l of the origin of the hand region, the vertical coordinate c of the origin, the horizontal coordinate x of the center position of the hand region, and the vertical coordinate y of the center position. The first position coordinate is the position coordinate of the hand region in the target image, that is, the center position coordinate (x, y).
[0083] According to the coordinates of the center position, width and height of the target image, and the first position coordinates of the hand area, the first confidence of the hand area is calculated to reflect the degree to which the position of the hand area in the target image can clearly and accurately express the state of the hand. The width of the target image is W, the height is H, and the center position is The hand region detector detects that the center position of the hand region is (a, b), and the first confidence can be determined according to the following formula:
[0084]
[0085] In this embodiment, the first confidence of the hand area is determined by the position of the hand area and the size information of the target image, which provides a reliability judgment basis for determining the hand area image. After obtaining the hand area image, it is necessary to detect the key points of the hand to provide an accurate data source for gesture recognition.
[0086] Figure 4 The flowchart of the embodiment of the present disclosure is to detect the hand key points through the hand key point detector based on the hand area image. Figure 4right Figure 1 Step 102 is specifically described, based on Figure 4 The embodiment shown comprises the following steps:
[0087] Step 401: extract a first feature map of a hand image through a first branch of a residual neural network of a first feature extraction network.
[0088] In this embodiment, the residual neural network is a deep learning model for extracting image features. The first feature extraction network is a neural network model for extracting features of hand key points in the hand area image, which is composed of a residual neural network and a fully connected layer. The first branch is obtained by splitting the convolution layer, activation layer and pooling layer of the residual neural network. The number of channels of the convolution layer of the first branch is half of the number of channels of the convolution layer of the residual neural network. Preferably, the activation layer uses a relu activation function, and the pooling layer uses average pooling avgpooling. The first feature map is the hand feature extracted by the first branch of the residual neural network of the first feature extraction network through the hand area image.
[0089] Step 402, extracting a second feature map of the hand image through the second branch of the residual neural network of the first feature extraction network, wherein the number of channels of the convolutional layers of the first branch and the second branch are both half of the convolutional layer of the residual neural network, the activation layers of the first branch and the second branch are different, and / or the pooling layers of the first branch and the second branch are different.
[0090] In this embodiment, the second branch is obtained by splitting the convolution layer, activation layer and pooling layer of the residual neural network. The number of channels of the convolution layer of the second branch is the same as the number of channels of the convolution layer of the first branch, which is half of the number of channels of the convolution layer of the residual neural network. Preferably, the activation layer uses a sigmoid activation function, and the pooling layer uses maximum pooling maxpooling. The second feature map is the hand feature extracted by the second branch of the residual neural network of the first feature extraction network through the hand area image.
[0091] By splitting the convolution layer, activation layer and pooling layer of the residual neural network into the first branch and the second branch, the richness of the residual neural network for hand feature extraction is improved.
[0092] Step 403: fuse the first feature map and the second feature map into a third feature map.
[0093] In this embodiment, the first feature map and the second feature map are directly spliced together to form a new feature map as the third feature map. The third feature map may contain all the information of the two feature maps. Alternatively, the first feature map and the second feature map are fused together in a certain manner to form a new feature map. The fusion method may be weighted fusion, product fusion, maximum fusion, minimum fusion, etc. The fused third feature map may contain important information in the two feature maps while eliminating redundant information therein.
[0094] Step 404: Based on the third feature map, predict the coordinates of a second number of hand key points through the fully connected layer of the first feature extraction network.
[0095] In this embodiment, based on the third feature map, the coordinates of the key points of the hand are estimated using the fully connected layer of the first feature extraction network. The output layer of the fully connected layer has N neural network nodes for obtaining the coordinates of the N key points of the hand. Preferably, N=21. That is, the coordinate values of x and y of 21 key points are detected from the hand area image through the first feature extraction network.
[0096] In this embodiment, the first feature extraction network is used as a hand key point detector. After the first feature extraction network is trained by providing hand key point sample data, it can predict the hand key points in the hand area image. The hand key point detector is obtained by training the first feature extraction network by using the hand area image and the key points in the hand area image as labels to form training sample data.
[0097] In this embodiment, by constructing a first feature extraction network as a hand key point detector, the coordinates of the hand key points are extracted from the hand area image, providing high-quality input data for gesture recognition.
[0098] Figure 5 This is a flow chart of determining the finger pointing category through a finger pointing classifier based on hand key points according to an embodiment of the present disclosure. Figure 5 right Figure 1 Step 103 and Figure 4 Step 404 is specifically described based on the following Figure 5 The embodiment shown comprises the following steps:
[0099] Step 501, using the second number of hand key point coordinates as the input of the finger pointing classifier, performing feature analysis on the second number of hand key point coordinates based on the multi-layer fully connected layer of the finger pointing classifier, and determining the predicted probability value corresponding to each category in the finger pointing category.
[0100] The finger pointing classifier can use a multi-layer perceptron, logistic regression, support vector machine, etc. as a classification model. Preferably, a multi-layer perceptron is used as the finger pointing classifier model.
[0101] In one implementation of this embodiment, the input layer of the finger pointing classifier model has 21 neural network nodes for receiving N=21 point coordinates as input, and the input data is analyzed using multi-layer full connection in the middle hidden layer. The output layer is provided with neural network nodes with a preset number of gestures to be recognized, and the output value of each neural network node is obtained by a softmax function to obtain a predicted probability value. That is, on each neural network node of the multi-finger pointing category of the output layer, the obtained output value is subjected to softmax processing, and the softmax is as follows:
[0102] where z i is the i-th neural network node of the output layer, and M is the number of neural network nodes in the output layer, that is, the number of multi-finger pointing categories.
[0103] The finger pointing classification model uses a large number of 21 finger key point coordinates and their corresponding gesture labels as training samples, and trains them through the back propagation algorithm to obtain the weight set of the neural network nodes in the middle hidden layer when the loss is lowest. This weight set is used as the finger pointing classifier model.
[0104] The multi-layer perceptron is used as a pointing classifier to classify the finger pointing category. The classification result of the pointing classifier may include at least two finger pointing categories. For example, the thumb points to the right and the thumb points to the left. The pointing classifier is trained by the multi-layer perceptron using the training sample data consisting of the key point coordinate sequence data and the labels.
[0105] Step 502: The category with the largest predicted probability value is taken as the finger pointing category.
[0106] The second number of hand key point coordinates output by the first feature extraction network is used as input data, and the predicted probability value corresponding to each gesture is obtained through the pointing classifier model. The gesture category corresponding to the maximum predicted probability value is used as the classification result of the finger pointing category.
[0107] Through the embodiment of the gesture interaction method disclosed in the present invention, the hand area image is obtained, and the search space of gesture recognition is reduced; based on the hand area image, the hand key points are detected by the hand key point detector, and the coordinate information of the hand key points is obtained, providing accurate and reliable data for gesture recognition; based on the hand key points, the finger pointing category is determined by the finger pointing classifier, and the recognition result of the finger pointing category with high accuracy and high recall rate is obtained; the vehicle is controlled to execute the first function corresponding to the finger pointing category, and the natural human-computer interaction through finger pointing is realized to quickly control the vehicle. Let the user get a better human-computer interaction experience in the vehicle cabin.
[0108] Corresponding to the methods provided in the above-mentioned embodiments, the present disclosure also provides a gesture interaction device. Since the device provided in the embodiment of the present disclosure corresponds to the methods provided in the above-mentioned embodiments, the implementation method of the method is also applicable to the device provided in this embodiment and will not be described in detail in this embodiment.
[0109] Figure 6 FIG. 6 is a schematic diagram of a gesture interaction device 600 according to an embodiment of the present disclosure. Figure 6 As shown, the gesture interaction device includes:
[0110] The acquisition module 610 is used to acquire a hand area image.
[0111] The key point detection module 620 is used to detect the hand key points through a hand key point detector based on the hand area image.
[0112] The pointing classification module 630 is used to determine the finger pointing category through a finger pointing classifier based on the hand key points.
[0113] The control module 640 is used to control the vehicle to execute a first function corresponding to the finger pointing category.
[0114] In some embodiments, the acquisition module 610 is used to:
[0115] Acquire target images;
[0116] According to the target image, a hand region is detected by a hand region detector;
[0117] Determine a first confidence level of the hand region based on a distance between a first coordinate point in the hand region and a second coordinate point in the target image, wherein the distance is inversely proportional to the first confidence level;
[0118] Determining whether the first confidence level is greater than a preset confidence level threshold;
[0119] If the first confidence is greater than a preset confidence threshold, the hand region is used as the hand region image;
[0120] If the first confidence level is less than or equal to a preset confidence threshold, a hand region image is acquired based on a next frame image of the target image.
[0121] In some embodiments, the acquisition module 610 determines the first confidence level of the hand region based on the distance between the first coordinate point in the hand region and the second coordinate point in the target image in the following manner:
[0122] According to the target application currently running in the vehicle system and the correspondence between the application and the reference point position saved in advance, the target reference point corresponding to the target application is determined, and the target reference point is used as the second coordinate point. In some embodiments, the key point detection module 620 is used to:
[0123] Extracting a first feature map of the hand region image through a first branch of a residual neural network of a first feature extraction network;
[0124] Extracting a second feature map of the hand region graphic through a second branch of the residual neural network of the first feature extraction network, wherein the number of channels of the convolutional layers of the first branch and the second branch are both half of the convolutional layer of the residual neural network, the activation layers of the first branch and the second branch are different, and / or the pooling layers of the first branch and the second branch are different;
[0125] fusing the first feature map and the second feature map into a third feature map;
[0126] Based on the third feature map, a second number of hand key point coordinates are predicted by a fully connected layer of the first feature extraction network. In some embodiments, the pointing classification module 630 is used to:
[0127] Using the second number of hand key point coordinates as input of the finger pointing classifier, performing feature analysis on the second number of hand key point coordinates based on the multi-layer fully connected layer of the finger pointing classifier, and determining a prediction probability value corresponding to each category in the finger pointing category;
[0128] The category with the largest predicted probability value is taken as the category that the finger is pointing to.
[0129] In summary, a hand area image is obtained through a gesture interaction device; based on the hand area image, the hand key points are detected through a hand key point detector; based on the hand key points, the finger pointing category is determined through a finger pointing classifier; and the vehicle is controlled to execute the first function corresponding to the finger pointing category. The device solves the problems of low accuracy and low recall rate of finger pointing recognition in the related art, and improves the accuracy of finger pointing recognition. In addition, the device has rich application scenarios, and uses finger pointing to realize the designated seat heating area in the cabin and the opening and retracting of the rear screen. For example: when the hand of the person in the car appears within the shooting range of the on-board camera, it is recognized that the thumb points to the left, indicating that the left side of the seat is desired to be heated, and the thumb points to the right, indicating that the right side of the seat is desired to be heated. For another example: when the hand of the rear passenger appears within the shooting range of the on-board camera, it is recognized that the thumb points down, indicating that the rear screen is turned on, and the thumb points up, indicating that the rear screen is turned off. The disclosure aims to innovate the user's experience of using touch or voice interaction in the cabin, replacing it with a choice of functions such as seat heating area that can be realized with only one gesture, which greatly improves the user experience.
[0130] In the embodiments provided in the present application, the methods and devices provided in the embodiments of the present application are introduced. In order to implement the functions in the methods provided in the embodiments of the present application, the electronic device may include a hardware structure and a software module, and implement the functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. A function of the functions may be executed in the form of a hardware structure, a software module, or a hardware structure plus a software module.
[0131] Figure 7 is a block diagram of an electronic device 700 for implementing the above-mentioned gesture interaction method according to an exemplary embodiment.
[0132] For example, the electronic device 700 may be a mobile phone, a computer, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0133] Reference Figure 7 , the electronic device 700 may include one or more of the following components: a processing component 702 , a memory 704 , a power component 706 , a multimedia component 708 , an audio component 710 , an input / output (I / O) interface 712 , a sensor component 714 , and a communication component 716 .
[0134] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.
[0135] The memory 704 is configured to store various types of data to support operations on the electronic device 700. Examples of such data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, videos, etc. The memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0136] The power supply component 706 provides power to the various components of the electronic device 700. The power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 700.
[0137] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0138] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC), and when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 704 or sent via the communication component 716. In some embodiments, the audio component 710 also includes a speaker for outputting audio signals.
[0139] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0140] The sensor assembly 714 includes one or more sensors for providing various aspects of status assessment for the electronic device 700. For example, the sensor assembly 714 can detect the open / closed state of the electronic device 700, the relative positioning of components, such as the display and keypad of the electronic device 700, and the sensor assembly 714 can also detect the position change of the electronic device 700 or a component of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and the temperature change of the electronic device 700. The sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 714 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 714 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0141] The communication component 716 is configured to facilitate wired or wireless communication between the electronic device 700 and other devices. The electronic device 700 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewR7dio) or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0142] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0143] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, and the instructions can be executed by a processor 720 of an electronic device 700 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0144] The embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the gesture interaction method described in the above embodiments of the present disclosure.
[0145] The embodiments of the present disclosure further provide a computer program product, including a computer program, which executes the gesture interaction method described in the above embodiments of the present disclosure when a processor is used to execute the computer program.
[0146] An embodiment of the present disclosure also proposes a chip, which includes one or more interface circuits and one or more processors; the interface circuit is used to receive signals from a memory of an electronic device and send signals to the processor, the signals include computer instructions stored in the memory, and when the processor executes the computer instructions, the electronic device executes the gesture interaction method described in the above embodiments of the present disclosure.
[0147] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0148] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0149] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.
[0150] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processing module, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (control method), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing in a suitable manner if necessary, and then stored in a computer memory.
[0151] It should be understood that the various parts of the embodiments of the present disclosure can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0152] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0153] In addition, each functional unit in each embodiment of the present disclosure may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0154] Although the embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present disclosure. A person skilled in the art may make changes, modifications, substitutions and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A gesture interaction method, characterized in that: Applied to a vehicle computer system, the method includes: Acquire a hand area image; Based on the hand region image, detecting hand key points by a hand key point detector; Based on the hand key points, determining the finger pointing category by a finger pointing classifier; The vehicle is controlled to execute a first function corresponding to the finger pointing category.
2. The method according to claim 1, characterized in that The acquiring of the hand region image comprises: Acquire target images; According to the target image, detecting a hand region by a hand region detector; Determining a first confidence level of the hand region based on a distance between a first coordinate point in the hand region and a second coordinate point in the target image, wherein the distance is inversely proportional to the first confidence level; Determining whether the first confidence level is greater than a preset confidence level threshold; If the first confidence is greater than the preset confidence threshold, taking the hand region as the hand region image; If the first confidence is less than or equal to the preset confidence threshold, the hand area image is acquired based on the next frame image of the target image.
3. The method according to claim 2, characterized in that The process of determining the second coordinate point includes: According to the target application currently running in the vehicle system and the correspondence between the application and the reference point position saved in advance, the target reference point corresponding to the target application is determined, and the target reference point is used as the second coordinate point.
4. The method according to claim 1, characterized in that: The detecting the hand key points by a hand key point detector based on the hand region image comprises: Extracting a first feature map of the hand region image through a first branch of a residual neural network of a first feature extraction network; Extracting a second feature map of the hand region graphic through a second branch of the residual neural network of the first feature extraction network, wherein the number of channels of the convolutional layers of the first branch and the second branch is half of the convolutional layer of the residual neural network, the first branch and the second branch have different activation layers, and / or the first branch and the second branch have different pooling layers; fusing the first feature map and the second feature map into a third feature map; Based on the third feature map, a second number of hand key point coordinates are predicted through a fully connected layer of the first feature extraction network.
5. The method according to claim 4, characterized in that The determining the finger pointing category by a finger pointing classifier based on the hand key points includes: Using the second number of the hand key point coordinates as input of a finger pointing classifier, performing feature analysis on the second number of the hand key point coordinates based on the multi-layer fully connected layers of the finger pointing classifier, and determining a prediction probability value corresponding to each category in the finger pointing category; The category with the largest predicted probability value is taken as the finger pointing category.
6. A gesture interaction device, characterized in that: Applied to a vehicle system, the device comprises: An acquisition module, used for acquiring a hand area image; A key point detection module, used for detecting hand key points through a hand key point detector based on the hand area image; A pointing classification module, configured to determine a finger pointing category through a finger pointing classifier based on the hand key points; The control module is used to control the vehicle to execute a first function corresponding to the finger pointing category.
7. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.
9. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 5.
10. A chip, characterized in that: It comprises one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from a memory of an electronic device and send the signal to the processor, the signal includes a computer instruction stored in the memory, and when the processor executes the computer instruction, the electronic device executes the method described in any one of claims 1 to 5.
Citation Information
Cited By
Image recognition method and device and electronic equipment
CN118212647A