Finger pointing identification method and device, electronic equipment and storage medium
By combining a hand region detector and a key point detection model, finger key point features are extracted and fused, solving the problem of inaccurate finger pointing in gesture recognition systems under complex backgrounds, and achieving more accurate target finger recognition and control.
Patent Information
- Application Number
- CN202410534062.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-29
- Publication Date
- 2025-10-31
AI Technical Summary
In existing human-computer interaction systems based on gesture recognition, the inability to accurately identify the user's finger pointing in complex backgrounds leads to inaccurate control.
By combining a hand region detector, a key point detection model, and a classification model, the key point features of fingers in hand images are extracted and fused, reducing the dependence on background information and achieving accurate identification of the target finger's direction.
It improves the accuracy of finger pointing recognition, reduces the reliance on background information of hand images, supports more finger pointing recognition and a wider angle range, and reduces the impact of image distortion and camera angle.
Smart Images

Figure CN120877324A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of gesture recognition technology, and more specifically, to a method, device, electronic device, and storage medium for recognizing finger pointing. Background Technology
[0002] Gesture recognition technology can be applied to human-computer interaction within car cabins, such as controlling devices in the cabin through gestures. Currently, in gesture-based human-computer interaction, cameras installed in the car cabin typically capture images of the user's hands. The captured hand images are then segmented to obtain the hand portion, which is then input into a trained detection model. The model classifies the hand portion to determine the user's finger orientation, thereby determining the corresponding control command and responding accordingly.
[0003] However, when using current detection models to detect the direction of a target finger in a partial image of the hand, the model is highly dependent on the background information of the acquired hand image. This makes it difficult to accurately identify the user's finger direction when the background of the acquired hand image is complex. Summary of the Invention
[0004] To address the problem of inaccurate recognition of user finger pointing during human-computer interaction, embodiments of this application provide a finger pointing recognition method, device, electronic device, and storage medium.
[0005] In a first aspect, embodiments of this application provide a method for recognizing finger pointing, including:
[0006] The original hand image of the identified user is input into the hand region detector, which detects and outputs the hand position information in the original hand image.
[0007] Based on the hand position information, the hand region image is extracted from the original hand image. The hand region image is then input into a key point detection model for detection to obtain the position information of each finger key point in the hand region image. Finally, hand feature information containing the position information of each finger key point is output.
[0008] The hand region image and the hand feature information are input into a classification model. The classification model is used to extract and fuse the key points of the target finger contained in the hand region image and the hand feature information to determine and output the direction of the target finger.
[0009] As an optional implementation of this application, the hand region detector includes: a first feature extraction network, a first feature fusion network, and a first detection network; the step of inputting the identified user's original hand image into the hand region detector, and detecting and outputting the hand position information in the original hand image through the hand region detector includes:
[0010] The original hand image is input into the first feature extraction network for image feature extraction, and the extracted original hand image features are input into the first feature fusion network for fusion to obtain the fused features of the original hand image;
[0011] The fused features of the hand image are input into the first detection network to detect the hand position in the original hand image, and the hand position information in the original hand image is obtained and output.
[0012] As an optional implementation of this application, the keypoint detection model includes: a second feature extraction network and a second detection network; the step of inputting the hand region image into the keypoint detection model for detection, obtaining the position information of each finger keypoint in the hand region image, and outputting hand feature information containing the position information of each finger keypoint includes:
[0013] The hand region image is input into the second feature extraction network for feature extraction to obtain the hand region image features;
[0014] The hand region image features are input into the second detection network to detect the positions of key points of each finger in the hand region image, obtain the position information of each key point of the finger, and output hand feature information containing the position information of each key point of the finger.
[0015] As an optional implementation of this application, the classification model includes: a third feature extraction network, a third feature fusion network, and a third detection network; the step of inputting the hand region image and the hand feature information into the classification model, and using the classification model to extract and fuse features of the target finger key points contained in the hand region image and the hand feature information, to determine and output the direction of the target finger, includes:
[0016] The hand region image and the hand feature information are input into the third feature extraction network for feature extraction to obtain the hand region image features corresponding to the hand region image and the target finger key point features contained in the hand feature information;
[0017] The hand region image features and the target finger key point features are input into the third feature fusion network for fusion to obtain the fused features of the target finger;
[0018] The fused features of the target finger are input into the third detection network for classification, and the direction of the target finger is determined and output.
[0019] As an optional implementation of this application, the keypoint detection model further outputs gesture features corresponding to the hand feature information; before inputting the hand region image and the hand feature information into the classification model, the method includes:
[0020] Determine whether the gesture feature corresponding to the hand feature information output by the key point detection model is a referential gesture feature;
[0021] If so, the hand region image and the hand feature information are input into the classification model.
[0022] After determining and outputting the direction of the target finger, the method includes:
[0023] The direction from the root key point of the target finger to the fingertip key point in the original hand image is taken as the first direction, and the direction of the X-axis in the image coordinate system is taken as the second direction. It is determined whether the angle between the first direction and the second direction is within a preset angle range.
[0024] If the included angle is within a preset included angle range, then determine whether the positional relationship between the root key point and the fingertip key point of the target finger in the original hand image is consistent with the direction of the target finger determined by the classification model;
[0025] If they match, then it is determined whether the target finger is within the valid area of the original hand image;
[0026] If the target finger is within the valid area of the original hand image, then the object pointed to by the target finger is determined as the target object.
[0027] As an optional implementation of this application, after determining and outputting the direction of the target finger, the method includes:
[0028] Using the moment when the original hand image is identified as a reference moment, it is determined whether a user's voice control command is received within a preset time period; wherein, the voice control command is used to instruct the execution of a target control operation on an object inside the car cabin.
[0029] If a user's voice control command is received within a preset time period, the target control operation is performed on the target object indicated by the target finger.
[0030] Secondly, embodiments of this application provide a finger pointing recognition device, comprising:
[0031] The image detection module is used to input the original hand image of the user into the hand region detector, and to detect and output the hand position information in the original hand image through the hand region detector.
[0032] The key point detection module is used to extract the hand region image from the original hand image based on the hand position information, input the hand region image into the key point detection model for detection, obtain the position information of each finger key point in the hand region image, and output hand feature information containing the position information of each finger key point.
[0033] The classification module is used to input the hand region image and the hand feature information into the classification model, and use the classification model to extract and fuse the key points of the target finger contained in the hand region image and the hand feature information to determine and output the direction of the target finger.
[0034] As an optional implementation of this application, the hand region detector includes: a first feature extraction network, a first feature fusion network, and a first detection network; the image detection module is specifically used to input the original hand image into the first feature extraction network for image feature extraction, and input the extracted original hand image features into the first feature fusion network for fusion to obtain the fused features of the original hand image;
[0035] The fused features of the hand image are input into the first detection network to detect the hand position in the original hand image, and the hand position information in the original hand image is obtained and output.
[0036] As an optional implementation of this application, the key point detection model includes: a second feature extraction network and a second detection network; the key point detection module is specifically used to input the hand region image into the second feature extraction network for feature extraction to obtain hand region image features;
[0037] The hand region image features are input into the second detection network to detect the positions of key points of each finger in the hand region image, obtain the position information of each key point of the finger, and output hand feature information containing the position information of each key point.
[0038] As an optional implementation of this application, the classification model includes: a third feature extraction network, a third feature fusion network, and a third detection network; the classification module is specifically used to input the hand region image and the hand feature information into the third feature extraction network for feature extraction, to obtain the hand region image features corresponding to the hand region image and the target finger key point features contained in the hand feature information;
[0039] The hand region image features and the target finger key point features are input into the third feature fusion network for fusion to obtain the fused features of the target finger;
[0040] The fused features of the target finger are input into the third detection network for classification, and the direction of the target finger is determined and output.
[0041] As an optional implementation of this application, the key point detection model also outputs the gesture features corresponding to the hand feature information; the device further includes: a judgment module, used to determine whether the gesture features corresponding to the hand feature information output by the key point detection model are referential gesture features before inputting the hand region image and the hand feature information into the classification model;
[0042] If so, the hand region image and the hand feature information are input into the classification model.
[0043] As an optional implementation of this application, the judgment module is further configured to, after determining and outputting the direction of the target finger, take the direction from the root key point of the target finger in the original hand image to the fingertip key point of the target finger as the first direction and the direction of the X-axis in the image coordinate system as the second direction, and determine whether the angle between the first direction and the second direction is within a preset angle range.
[0044] If the included angle is within a preset included angle range, then determine whether the positional relationship between the root key point and the fingertip key point of the target finger in the original hand image is consistent with the direction of the target finger determined by the classification model;
[0045] If they match, then it is determined whether the target finger is within the valid area of the original hand image;
[0046] If the target finger is within the valid area of the original hand image, then the object pointed to by the target finger is determined as the target object.
[0047] As an optional implementation of this application, the device further includes: a response module, configured to, after determining and outputting the direction of the target finger, determine whether a user's voice control command is received within a preset time period, taking the moment when the original hand image is recognized as a reference time; wherein the voice control command is used to instruct the execution of a target control operation on an object inside the car cabin;
[0048] If a user's voice control command is received within a preset time period, the target control operation is performed on the target object indicated by the target finger.
[0049] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the finger pointing recognition method described in the first aspect or any optional embodiment of the first aspect when the computer program is invoked.
[0050] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the finger pointing recognition method described in the first aspect or any optional implementation of the first aspect.
[0051] Fifthly, embodiments of this application provide a vehicle equipped with a finger-pointing recognition device as described in the second aspect, or an electronic device as described in the third aspect, or a storage medium as described in the fourth aspect.
[0052] The technical solution provided in this application has the following advantages compared with the prior art:
[0053] This application provides a method, apparatus, electronic device, and storage medium for recognizing finger pointing. The method includes: inputting an original hand image of a user into a hand region detector; detecting and outputting hand position information in the original hand image using the hand region detector; extracting a hand region image from the original hand image based on the hand position information; inputting the hand region image into a keypoint detection model for detection to obtain position information of keypoints of each finger in the hand region image; and outputting hand feature information containing the position information of each finger keypoint; inputting the hand region image and the hand feature information into a classification model; using the classification model to extract and fuse features of the target finger keypoints contained in the hand region image and the hand feature information; and determining and outputting the pointing of the target finger. This application embodiment detects the position information of key points of each finger in a hand region image, uses a classification model to extract and fuse the key point features of the target finger and the hand region image features to obtain the direction of the target finger. Since the key point features of the target finger are obtained, rather than simply classifying the direction of the target finger based on the hand image, the dependence on background information in the hand image is reduced, and the direction of the target finger can be accurately identified. Attached Figure Description
[0054] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0055] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0056] Figure 1 This is a flowchart of a finger pointing recognition method provided according to one or more embodiments of this application;
[0057] Figure 2 This is a flowchart of a finger pointing recognition method provided according to one or more embodiments of this application;
[0058] Figure 3 This is a flowchart of a finger pointing recognition method provided according to one or more embodiments of this application;
[0059] Figure 4 A structural block diagram of a finger pointing recognition device provided in one or more embodiments of this application;
[0060] Figure 5 A structural block diagram of a finger pointing recognition device provided in one or more embodiments of this application;
[0061] Figure 6 This is an internal structural diagram of an electronic device provided according to one or more embodiments of this application. Detailed Implementation
[0062] To make the objectives, implementation methods and advantages of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only some embodiments of this application, and not all embodiments.
[0063] Based on the exemplary embodiments described in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the appended claims. Furthermore, although the disclosures in this application are presented by way of one or more exemplary examples, it should be understood that each aspect of these disclosures can constitute a complete implementation on its own. It should be noted that the brief descriptions of terminology in this application are merely for the convenience of understanding the embodiments described below, and are not intended to limit the implementation of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0064] First, the application scenarios of the embodiments of this application are described. The method provided in the embodiments of this application can be applied to the cabin of a car. Users can control the devices on the vehicle based on pointing gestures and voice, such as controlling the air conditioner, windows, and other devices by pointing with their index finger. For example, in one scenario, after the user issues the voice command "turn it on," they point to the device to be turned on with their index finger. A camera installed in the vehicle cabin captures the user's gesture image, determines the direction of the index finger in the gesture image, and identifies the pointed device as the device to be turned on. For example, if the gesture image shows the index finger pointing to the air conditioner, the air conditioner is turned on.
[0065] Based on this, embodiments of this application provide a method, device, electronic device, and storage medium for recognizing finger pointing. The method determines the pointing of the target finger by detecting the position information of key points corresponding to each finger joint in a hand image. It has low requirements for the background information of the hand image and can accurately identify the pointing direction of the user's target finger.
[0066] The finger pointing recognition method provided in this application embodiment can be executed by the electronic device provided in this application embodiment, or it can be implemented by the finger pointing recognition device provided in this application embodiment, or it can be implemented by one or more functional entities on a vehicle. This application embodiment does not make any specific limitations.
[0067] The following detailed embodiments illustrate the finger pointing recognition method provided in this application. All finger pointing recognition methods described in the following embodiments can be executed on the vehicle side.
[0068] Figure 1 The flowchart of the finger pointing recognition method provided in the embodiments of this application is shown below. Figure 1 As shown, the finger pointing recognition method provided in this embodiment includes the following steps:
[0069] S11. Input the original hand image of the identified user into the hand region detector, and detect and output the hand position information in the original hand image through the hand region detector.
[0070] For example, a camera installed in the car cabin can be used to identify the user's original hand image. The original hand image can be an RGB image, an IR image, a TOF image, etc., and the size of the original hand image can be any size. This application embodiment does not impose any specific limitations.
[0071] In some embodiments, when the vehicle receives a user's activation command for the automatic gesture recognition function, in response to the activation command, the camera is activated to acquire the user's original hand image, and the acquired original hand image is input into a hand region detector. The hand region detector described in this embodiment is a trained hand region detector.
[0072] The hand region detector is used to detect hand position information in the original hand image. The network structure of the hand region detector can be divided into three parts: a first feature extraction network, a first feature fusion network, and a first detection network. For example, the hand position information in the original hand image can be detected through the following process: the original hand image is input into the first feature extraction network for image feature extraction; the extracted original hand image features are input into the first feature fusion network for fusion to obtain the fused features of the original hand image; the fused features of the hand image are input into the first detection network to detect the hand position in the original hand image, thereby obtaining and outputting the hand position information in the original hand image.
[0073] The system comprises three main components: a first feature extraction network (a backbone network used to extract image features from the original hand image, such as Mobilenetv3 for a good balance between speed and accuracy); a first feature fusion network (a network structure for feature fusion, such as a neck network formed by stacked FPNs to obtain feature information of hands of different sizes and distances in the original hand image); and a first detection network (a network for over-detection, such as a prediction head network with an anchor-free structure). In short, the hand region detector includes a backbone network, a neck network, and a prediction head network.
[0074] S12. Based on the hand position information, extract the hand region image from the original hand image, input the hand region image into the key point detection model for detection, obtain the position information of each finger key point in the hand region image, and output hand feature information containing the position information of each finger key point.
[0075] The key points of each finger include the joint nodes and fingertips of each finger. For example, each finger in the hand image can be considered to contain four key points, and the palm joint node corresponds to one key point. That is, each hand in the hand image includes 21 finger key points. The hand region image extracted from the original hand image can be a rectangular box containing the hand region. The obtained rectangular box is input into the key point detection model for detection.
[0076] In some embodiments, the keypoint detection model may include a second feature extraction network and a second detection network. The second feature extraction network is a backbone network for extracting features from a hand region image, and the second detection network is a detection head network for detecting keypoints in the hand region image. For example, ResNet18 can be selected as the backbone network, where each of the four stages (stage1, stage2, stage3, and stage4) of ResNet18 contains two base blocks, and each base block contains two convolutional neural networks for extracting image features. The second detection network can use regression to obtain the location information of each keypoint.
[0077] For example, the hand region image is input into the second feature extraction network for feature extraction to obtain hand region image features; the hand region image features are input into the second detection network to detect the position of each finger key point in the hand region image to obtain the position information of each finger key point, and hand feature information containing the position information of each finger key point is output.
[0078] In some embodiments, while outputting hand feature information containing the positional information of each finger keypoint, gesture features of the hand feature information are also output. These gesture features indicate whether the gesture in the hand feature information is a referential or non-referential gesture. For example, it can be determined whether the gesture feature corresponding to the hand feature information output by the keypoint detection model is a referential gesture feature. If yes, the hand region image and the hand feature information are input into a classification model; if no, the current frame image is filtered out, and subsequent gesture category detection is no longer performed on the current frame image. Specifically, if it is determined that the gesture feature corresponding to the hand feature information output by the keypoint detection model is a referential gesture feature, then the hand region image is determined to be a referential gesture image; if it is determined that the gesture feature corresponding to the hand feature information output by the keypoint detection model is a non-referential gesture, then the hand region image is determined to be a non-referential gesture image.
[0079] Filtering out non-referential gesture images can reduce the amount of subsequent computation.
[0080] S13. Input the hand region image and the hand feature information into the classification model, and use the classification model to extract and fuse the key points of the target finger contained in the hand region image and the hand feature information to determine and output the direction of the target finger.
[0081] In some embodiments, the classification model includes: a third feature extraction network, a third feature fusion network, and a third detection network; the third feature extraction network is used to extract image features from hand region images, for example, by extracting features from hand region images through multi-layer convolutional neural networks, or by using a multi-layer perceptron (MLP) to extract features from key points of the target finger in hand feature information; the third feature fusion network is a neural network for performing feature fusion, and the third detection network is a neural network for classifying the gestures of the target finger. The third detection network can be a network composed of multi-layer fully connected neural networks, such as a 5-layer fully connected neural network.
[0082] For example, the process of determining the direction of the target finger may include: inputting the hand region image and the hand feature information into the third feature extraction network for feature extraction to obtain the hand region image features corresponding to the hand region image and the target finger key point features contained in the hand feature information; inputting the hand region image features and the target finger key point features into the third feature fusion network for fusion to obtain the fused features of the target finger; inputting the fused features of the target finger into the third detection network for classification to determine the direction of the target finger and output the direction of the target finger.
[0083] For example, a two-layer convolutional neural network extracts hand region image features from a hand region image. A multimodal neural network (MLP) is used to determine the target finger and extract its keypoint features. A third feature fusion network fuses the hand region image features and the target finger's keypoint features to obtain the fused features of the target finger. These fused features are then input into a five-layer fully connected neural network for classification to obtain the target finger's direction, which is then output. Here, the target finger can be the index finger. In this embodiment, the direction of the target finger is obtained through a multimodal recognition method, including directions such as up, left, right, left front, right front, left back, and right back. Correspondingly, when the target finger is the index finger, the classification model can output multiple directions for the index finger, such as: index finger pointing up, index finger pointing left, index finger pointing right, index finger pointing left front, index finger pointing right front, index finger pointing left back, and index finger pointing right back.
[0084] Compared to existing technologies, the finger pointing recognition method provided in this application embodiment can support the recognition of more finger pointing, no longer limited to simple directions such as up, left, and right. It supports a wider range and a larger angle area, which can reduce the impact of image distortion and camera angle on finger pointing recognition.
[0085] This application provides a method for recognizing finger pointing. The method includes: inputting an original hand image of a user into a hand region detector; detecting and outputting hand position information in the original hand image using the hand region detector; extracting a hand region image from the original hand image based on the hand position information; inputting the hand region image into a keypoint detection model for detection to obtain position information of keypoints of each finger in the hand region image; and outputting hand feature information containing the position information of each finger keypoint; inputting the hand region image and the hand feature information into a classification model; using the classification model to extract and fuse features of target finger keypoints contained in the hand region image and the hand feature information; and determining and outputting the pointing of the target finger. This application obtains the pointing of the target finger by detecting the position information of keypoints of each finger in the hand region image, extracting and fusing target finger keypoint features and hand region image features using a classification model. Because it obtains the keypoint features of the target finger, rather than simply classifying the pointing of the target finger based on the hand image, it reduces the dependence on background information in the hand image and can accurately identify the pointing of the target finger.
[0086] Figure 2 A flowchart of a finger pointing recognition method provided in another embodiment of this application is shown. Figure 1 Based on the illustrated embodiment, after step S13, the following steps S21 to S22 are also included, as shown below. Figure 2 As shown.
[0087] S21. Taking the direction from the root key point of the target finger in the original hand image to the fingertip key point of the target finger as the first direction, and the direction of the X-axis in the image coordinate system as the second direction, determine whether the angle between the first direction and the second direction is within a preset angle range.
[0088] If the included angle is within the preset included angle range, then the following step S22 is executed; if the included angle is outside the preset included angle range, then the direction of the target finger is determined to be a non-pointing gesture, and the determination of the target object and the execution of the control operation corresponding to the voice control command are no longer performed. This application embodiment does not limit the specific value of the preset included angle range.
[0089] S22. Determine whether the positional relationship between the root key point and the fingertip key point of the target finger in the original hand image is consistent with the direction of the target finger determined by the classification model.
[0090] If they match, then execute S23; if they do not match, then determine that the target finger is pointing to a non-pointing gesture, and no longer determine the target object or execute the control operation corresponding to the voice control command.
[0091] For example, if the target finger points to the left front and the root keypoint of the target finger is located to the right of the fingertip keypoint, the positional relationship between the root keypoint and the fingertip keypoint of the target finger in the original hand image is consistent with the direction of the target finger determined by the classification model; if the target finger points to the right rear, the root keypoint of the target finger must be located to the left of the fingertip keypoint. That is, if the target finger points to the right rear and the root keypoint of the target finger is located to the left of the fingertip keypoint, the positional relationship between the root keypoint and the fingertip keypoint of the target finger in the original hand image is consistent with the direction of the target finger determined by the classification model.
[0092] S23. Determine whether the target finger is within the valid area of the original hand image.
[0093] If the target finger is determined to be within the valid area of the original hand image, then step S24 is executed. If the target finger is determined to be outside the valid area of the original hand image, then the direction of the target finger is determined to be a non-pointing gesture, and the determination of the target object and the execution of the control operation corresponding to the voice control command are no longer performed.
[0094] S24. The object pointed to by the target finger is identified as the target object.
[0095] In some embodiments, after identifying the target object, it can be determined whether a user's voice control command has been received. If a user's voice control command has been received, the target operation corresponding to the voice control command is executed on the target object. For example, using the moment when the original hand image is recognized as a reference time, it is determined whether a user's voice control command has been received within a preset time period; wherein, the voice control command is used to instruct the execution of a target control operation on an object inside the car cabin; if a user's voice control command has been received within the preset time period, the target control operation is executed on the target object indicated by the target finger.
[0096] In this embodiment, the effective area of the original hand image refers to the non-edge area of the original hand image. Through logical judgments based on the angle between the direction the target finger is pointing in the original hand image and the X-axis direction in the image coordinate system, the positional relationship between the root key point and the fingertip key point of the target finger in the original hand image, and whether the target finger is within the effective area of the original hand image, a secondary judgment of the target finger's direction can be achieved, making the determined target finger's direction more accurate. While reducing dependence on the image background, it can support all interaction needs of one or two rows of users in the entire car cabin, further enhancing the user experience.
[0097] For example, Figure 3 This is a flowchart of a finger pointing recognition method provided in another embodiment of this application. Figure 1Based on the illustrated embodiment, after step S13, the following steps S31 to S32 are also included, as shown below. Figure 3 As shown.
[0098] S31. Using the moment when the original hand image is recognized as a reference moment, determine whether the user's voice control command is received within a preset time period.
[0099] The voice control commands are used to instruct the execution of target control operations on objects within the vehicle cabin, including at least action instructions, i.e., the control operations to be performed, such as "turn it on", "turn it off", "heat this", etc.
[0100] In this embodiment, the preset duration for which the moment when the original hand image is identified is used as the reference moment can include any time within a first preset duration before the reference moment and a second preset duration after the reference moment. If a user's voice control command is received within the preset duration, step S32 is executed; if no user's voice control command is received within the preset duration, the target control operation is not performed on the equipment in the car cabin.
[0101] S32. Perform target control operation on the target object indicated by the target finger.
[0102] The target object can be any device inside the car cabin. The target object inside the car cabin can be determined based on the direction of the target finger. When the target finger points to the left rear, the left rear window can be identified as the target object. When the target finger points upward, the sunroof can be identified as the target object. When the target finger points to the right front, the right front window of the vehicle can be identified as the target object.
[0103] For example, when the target finger points to the sunroof, if the user's voice control command "open it" is received within a preset time period (with the time when the original hand image is recognized as a reference time), the control operation to open the sunroof is executed.
[0104] In this embodiment, by using the moment the original hand image is recognized as a reference moment, it is determined whether a user's voice control command has been received within a preset time period, thus avoiding accidental operation. After determining the direction of the target finger using the above method, voice command control operations are performed on the target object pointed to by the target finger, helping the user control the equipment in the cockpit and improving the efficiency of human-computer interaction.
[0105] Based on the same inventive concept, as an implementation of the above method, this application embodiment also provides a finger pointing recognition device that performs the above embodiment. This device embodiment corresponds to the aforementioned method embodiment. For ease of reading, this device embodiment will not repeat the details of the aforementioned method embodiment one by one, but it should be clear that the finger pointing recognition device in this embodiment can correspondingly implement all the contents of the aforementioned method embodiment.
[0106] Figure 4 This is a schematic diagram of the structure of a finger pointing recognition device provided in one embodiment of this application, as shown below. Figure 4 As shown, the finger pointing recognition device 400 provided in this embodiment includes:
[0107] The image detection module 410 is used to input the original hand image of the user into the hand region detector, and to detect and output the hand position information in the original hand image through the hand region detector.
[0108] The key point detection module 420 is used to extract a hand region image from the original hand image based on the hand position information, input the hand region image into the key point detection model for detection, obtain the position information of each finger key point in the hand region image, and output hand feature information containing the position information of each finger key point.
[0109] The classification module 430 is used to input the hand region image and the hand feature information into the classification model, and use the classification model to extract and fuse the key points of the target finger contained in the hand region image and the hand feature information to determine and output the direction of the target finger.
[0110] As an optional implementation of this application, the hand region detector includes: a first feature extraction network, a first feature fusion network, and a first detection network; the image detection module 410 is specifically used to input the original hand image into the first feature extraction network for image feature extraction, input the extracted original hand image features into the first feature fusion network for fusion to obtain the fused features of the original hand image; input the fused features of the hand image into the first detection network to detect the hand position in the original hand image, and obtain and output the hand position information in the original hand image.
[0111] As an optional implementation of this application, the key point detection model includes: a second feature extraction network and a second detection network; the key point detection module 420 is specifically used to input the hand region image into the second feature extraction network for feature extraction to obtain hand region image features; input the hand region image features into the second detection network to detect the position of each finger key point in the hand region image to obtain the position information of each finger key point, and output hand feature information containing the position information of each key point.
[0112] As an optional implementation of this application, the classification model includes: a third feature extraction network, a third feature fusion network, and a third detection network; the classification module 430 is specifically used to input the hand region image and the hand feature information into the third feature extraction network for feature extraction, to obtain the hand region image features corresponding to the hand region image and the target finger key point features contained in the hand feature information; to input the hand region image features and the target finger key point features into the third feature fusion network for fusion, to obtain the fused features of the target finger; to input the fused features of the target finger into the third detection network for classification, to determine and output the direction of the target finger.
[0113] Figure 5 This is a schematic diagram of the structure of a finger pointing recognition device provided in one embodiment of this application. Figure 4 Based on the device shown, it also includes:
[0114] The judgment module 510 further outputs the gesture features corresponding to the hand feature information in the key point detection model; the judgment module 510 is used to determine whether the gesture features corresponding to the hand feature information output by the key point detection model are referential gesture features before inputting the hand region image and the hand feature information into the classification model; if so, the hand region image and the hand feature information are input into the classification model.
[0115] As an optional implementation of this application, the judgment module 510 is further configured to, after determining and outputting the direction of the target finger, determine whether the angle between the first direction and the second direction is within a preset angle range, taking the direction from the root key point of the target finger to the fingertip key point in the original hand image as the first direction and the direction of the X-axis in the image coordinate system as the second direction; if the angle is within the preset angle range, determine whether the positional relationship between the root key point and the fingertip key point of the target finger in the original hand image is consistent with the direction of the target finger determined by the classification model; if consistent, determine whether the target finger is within the effective area of the original hand image; if the target finger is within the effective area of the original hand image, determine the object pointed to by the target finger as the target object.
[0116] As an optional implementation of this application, the device further includes: a response module 520, configured to, after determining and outputting the direction of the target finger, determine whether a user's voice control command is received within a preset time period, taking the time when the original hand image is recognized as a reference time; wherein, the voice control command is used to instruct the execution of a target control operation on an object inside the car cabin; if the user's voice control command is received within the preset time period, then the target control operation is executed on the target object indicated by the target finger.
[0117] In one embodiment, an electronic device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of any of the finger pointing recognition methods described in the above method embodiments.
[0118] For example, Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device provided in this embodiment includes a memory 61 and a processor 62. The memory 61 is used to store computer programs; the processor 62 is used to execute the steps of the finger pointing recognition method provided in the above-described method embodiment when the computer program is invoked. Its implementation principle and technical effects are similar, and will not be repeated here. Those skilled in the art will understand that... Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0119] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the finger pointing recognition methods described in the above method embodiments.
[0120] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM), etc.
[0121] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0122] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the discussion in some embodiments above is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the embodiments and various different variations of the embodiments suitable for specific application considerations.
Claims
1. A method for recognizing finger pointing, characterized in that, include: The original hand image of the identified user is input into the hand region detector, which detects and outputs the hand position information in the original hand image. Based on the hand position information, the hand region image is extracted from the original hand image. The hand region image is then input into a key point detection model for detection to obtain the position information of each finger key point in the hand region image. Finally, hand feature information containing the position information of each finger key point is output. The hand region image and the hand feature information are input into a classification model. The classification model is used to extract and fuse the key points of the target finger contained in the hand region image and the hand feature information to determine and output the direction of the target finger.
2. The method according to claim 1, characterized in that, The hand region detector includes: a first feature extraction network, a first feature fusion network, and a first detection network; the step of inputting the identified user's original hand image into the hand region detector, and detecting and outputting the hand position information in the original hand image through the hand region detector includes: The original hand image is input into the first feature extraction network for image feature extraction, and the extracted original hand image features are input into the first feature fusion network for fusion to obtain the fused features of the original hand image; The fused features of the hand image are input into the first detection network to detect the hand position in the original hand image, and the hand position information in the original hand image is obtained and output.
3. The method according to claim 1, characterized in that, The keypoint detection model includes: a second feature extraction network and a second detection network; the step of inputting the hand region image into the keypoint detection model for detection, obtaining the position information of each finger keypoint in the hand region image, and outputting hand feature information containing the position information of each finger keypoint includes: The hand region image is input into the second feature extraction network for feature extraction to obtain the hand region image features; The hand region image features are input into the second detection network to detect the positions of key points of each finger in the hand region image, obtain the position information of each key point of the finger, and output hand feature information containing the position information of each key point of the finger.
4. The method according to claim 1, characterized in that, The classification model includes: a third feature extraction network, a third feature fusion network, and a third detection network; the step of inputting the hand region image and the hand feature information into the classification model, and using the classification model to extract and fuse features of the target finger key points contained in the hand region image and the hand feature information, to determine and output the direction of the target finger, includes: The hand region image and the hand feature information are input into the third feature extraction network for feature extraction to obtain the hand region image features corresponding to the hand region image and the target finger key point features contained in the hand feature information; The hand region image features and the target finger key point features are input into the third feature fusion network for fusion to obtain the fused features of the target finger; The fused features of the target finger are input into the third detection network for classification, and the direction of the target finger is determined and output.
5. The method according to claim 1, characterized in that, The keypoint detection model also outputs gesture features corresponding to the hand feature information; before inputting the hand region image and the hand feature information into the classification model, the method includes: Determine whether the gesture feature corresponding to the hand feature information output by the key point detection model is a referential gesture feature; If so, the hand region image and the hand feature information are input into the classification model.
6. The method according to any one of claims 1-5, characterized in that, After determining and outputting the direction of the target finger, the method includes: The direction from the root key point of the target finger to the fingertip key point in the original hand image is taken as the first direction, and the direction of the X-axis in the image coordinate system is taken as the second direction. It is determined whether the angle between the first direction and the second direction is within a preset angle range. If the included angle is within a preset included angle range, then determine whether the positional relationship between the root key point and the fingertip key point of the target finger in the original hand image is consistent with the direction of the target finger determined by the classification model; If they match, then it is determined whether the target finger is within the valid area of the original hand image; If the target finger is within the valid area of the original hand image, then the object pointed to by the target finger is determined as the target object.
7. The method according to any one of claims 1-5, characterized in that, After determining and outputting the direction of the target finger, the method includes: Using the moment when the original hand image is identified as a reference moment, it is determined whether a user's voice control command is received within a preset time period; wherein, the voice control command is used to instruct the execution of a target control operation on an object inside the car cabin; If a user's voice control command is received within a preset time period, the target control operation is performed on the target object indicated by the target finger.
8. A finger pointing recognition device, characterized in that, include: The image detection module is used to input the original hand image of the user into the hand region detector, and to detect and output the hand position information in the original hand image through the hand region detector. The key point detection module is used to extract the hand region image from the original hand image based on the hand position information, input the hand region image into the key point detection model for detection, obtain the position information of each finger key point in the hand region image, and output hand feature information containing the position information of each finger key point. The classification module is used to input the hand region image and the hand feature information into the classification model, and use the classification model to extract and fuse the key points of the target finger contained in the hand region image and the hand feature information to determine and output the direction of the target finger.
9. An electronic device, comprising: A memory and a processor, the memory storing a computer program, characterized in that the processor, when executing the computer program, implements the finger pointing recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the finger pointing recognition method according to any one of claims 1 to 7.
11. A vehicle, characterized in that, The vehicle is equipped with the finger pointing recognition device as described in claim 8, or the electronic device as described in claim 9, or the storage medium as described in claim 10.