Dynamic gesture recognition method and intelligent vehicle-mounted device
By adopting dynamic gesture recognition methods in intelligent vehicle-mounted equipment and utilizing feature transfer fusion and classification technology, efficient dynamic gesture recognition is achieved under limited computing resources, solving the problem of huge computing resource consumption and ensuring recognition accuracy and security.
Patent Information
- Application Number
- CN202210618569.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2042-06-01
AI Technical Summary
Existing dynamic gesture recognition solutions consume huge computing resources and are not suitable for scenarios with limited computing resources for in-vehicle smart devices.
A dynamic gesture recognition method is adopted. By acquiring several frames of images containing dynamic gestures, the action recognition model is called to perform feature transfer fusion and classification, and the spatiotemporal dimension feature extraction and classification are used to reduce dependence on computing resources.
Efficient dynamic gesture recognition is achieved in intelligent vehicle-mounted equipment, saving computing resources while ensuring the accuracy of action recognition and avoiding the potential safety hazards of attention being diverted from the road due to operating the central control.
Smart Images

Figure CN114973414B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the field of intelligent driving, and in particular, to a dynamic gesture recognition method and an intelligent vehicle-mounted device. BACKGROUND
[0002] At present, dynamic gesture recognition is mainly based on a depth camera to obtain a depth map or a three-dimensional point cloud, and then a three-dimensional convolution (3DCNN) is used to calculate the action classification result with the help of a recurrent neural network (RNN / LSTM) or a variant network of the recurrent neural network.
[0003] However, the above recognition scheme requires a large amount of computing resources and computing power, and is not suitable for the limited computing resources of a vehicle-mounted intelligent device. SUMMARY
[0004] The present specification provides a dynamic gesture recognition method and an intelligent vehicle-mounted device to solve or partially solve the technical problem that the current dynamic gesture recognition scheme is not suitable for the limited resource scenario of a vehicle-mounted intelligent device due to huge consumption of computing resources.
[0005] To solve the above technical problem, the present specification provides a dynamic gesture recognition method, which comprises:
[0006] In the dynamic gesture recognition mode, a plurality of frames of images containing a dynamic gesture are acquired;
[0007] An action recognition model is called to perform feature transmission fusion and classification on the plurality of frames of images to obtain an action category of the plurality of frames of images;
[0008] An action category corresponding to the dynamic gesture is determined from the action category of the plurality of frames of images.
[0009] Preferably, before the plurality of frames of images containing a dynamic gesture are acquired, the method further comprises:
[0010] A hand action is photographed by using a camera module;
[0011] Whether the hand action has a trigger condition is detected by using a hand detection model;
[0012] If yes, the dynamic gesture recognition mode is started.
[0013] Preferably, whether the hand action has a trigger condition is detected by using a hand detection model, specifically comprising:
[0014] Whether a hovering time of the hand action exceeds a preset time is detected;
[0015] If yes, it indicates that the hand action has the trigger condition.
[0016] Preferably, the calling action recognition model performs feature transfer fusion and classification on the plurality of frames of images to obtain action categories of the plurality of frames of images, and specifically includes:
[0017] In the plurality of frames of images, the spatial dimension features of the associated previous adjacent frame of image perform feature transfer fusion on the spatial dimension features in the current frame of image to obtain the space-time dimension features of the current frame of image.
[0018] The space-time dimension features of the current frame of image are subjected to convolution calculation to obtain the action category of the current frame of image.
[0019] Preferably, the spatial dimension features of the associated previous adjacent frame of image perform feature transfer fusion on the spatial dimension features in the current frame of image to obtain the space-time dimension features of the current frame of image, and specifically includes:
[0020] The current frame of image is subjected to convolution processing to obtain the spatial dimension features of the current frame of image.
[0021] Part of the spatial dimension features in the previous adjacent frame of image are used to replace part of the spatial dimension features of the current frame of image to obtain the space-time dimension features of the current frame of image.
[0022] Preferably, after the part of the spatial dimension features in the previous adjacent frame of image are used to replace part of the spatial dimension features of the current frame of image, the method further includes:
[0023] The replaced part of the spatial dimension features of the current frame of image is stored to replace part of the spatial dimension features of the subsequent adjacent frame of image.
[0024] Preferably, the action category corresponding to the dynamic gesture is determined from the action categories of the plurality of frames of images, and specifically includes:
[0025] A target action category is selected from the action categories of the plurality of frames of images as the action category corresponding to the dynamic gesture, wherein the number of images corresponding to the target action category is above a preset number threshold.
[0026] Preferably, the action category corresponding to the dynamic gesture is determined from the action categories of the plurality of frames of images, and specifically includes:
[0027] A time window is determined, and the time window is a time window corresponding to a set number of frames.
[0028] According to the time window, an action category corresponding to the time window is determined from the action categories of the plurality of frames of images.
[0029] determine the target action category based on the time window, and determine an action category corresponding to the dynamic gesture based on the target action category.
[0030] Preferably, after determining the action category corresponding to the dynamic gesture from the action categories of the plurality of frames of images, the method further comprises:
[0031] transmitting the action instruction corresponding to the dynamic gesture to a downstream object for execution;
[0032] exiting the dynamic gesture recognition mode.
[0033] The present specification provides an intelligent vehicle-mounted device, comprising:
[0034] a camera module configured to acquire a plurality of frames of images containing a dynamic gesture in a dynamic gesture recognition mode;
[0035] a space-time conversion module configured to call an action recognition model to perform feature transfer fusion and classification on the plurality of frames of images to obtain action categories of the plurality of frames of images;
[0036] a determination module configured to determine an action category corresponding to the dynamic gesture from the action categories of the plurality of frames of images.
[0037] Preferably, the intelligent vehicle-mounted device further comprises:
[0038] the camera module is configured to capture a hand action;
[0039] a hand detection module configured to detect whether the hand action meets a triggering condition using a hand detection model;
[0040] if yes, start the dynamic gesture recognition mode.
[0041] Preferably, the hand detection module is specifically configured to:
[0042] detect whether a hovering time of the hand action exceeds a preset time;
[0043] if yes, it indicates that the hand action meets the triggering condition.
[0044] Preferably, the space-time conversion module is specifically configured to:
[0045] perform feature transfer fusion in a time dimension on spatial dimension features of a current frame of image in association with spatial dimension features of a previous adjacent frame of image in the plurality of frames of images to obtain space-time dimension features of the current frame of image;
[0046] perform convolution calculation on the space-time dimension features of the current frame of image to obtain an action category of the current frame of image.
[0047] Preferably, the space-time conversion module is specifically used for:
[0048] performing convolution processing on the current frame image to obtain spatial dimension features of the current frame image;
[0049] replacing part of the spatial dimension features of the current frame image with part of the spatial dimension features in the previous frame image to obtain space-time dimension features of the current frame image.
[0050] Preferably, the intelligent vehicle-mounted device further comprises:
[0051] a storage module configured to store the replaced part of the spatial dimension features in the current frame image to replace part of the spatial dimension features in a subsequent frame image.
[0052] Preferably, the space-time conversion module is specifically used for:
[0053] selecting a target action category as the action category corresponding to the dynamic gesture from the action categories of the plurality of frame images, wherein the target action category corresponds to an image quantity above a preset quantity threshold.
[0054] Preferably, the determination module is specifically used for:
[0055] determining a time window, wherein the time window is a time window corresponding to a set number of frames;
[0056] determining an action category corresponding to the time window from the action categories of the plurality of frame images according to the time window;
[0057] determining the target action category based on the action category corresponding to the time window, and determining the action category corresponding to the dynamic gesture based on the target action category.
[0058] Preferably, the intelligent vehicle-mounted device further comprises:
[0059] a transmission module configured to transmit the action instruction corresponding to the dynamic gesture to a downstream object for execution;
[0060] a quit module configured to quit the dynamic gesture recognition mode.
[0061] The specification discloses a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the steps of the above method.
[0062] The specification discloses an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the above method when executing the program.
[0063] Through one or more embodiments of the present specification, the present specification has the following beneficial effects or advantages:
[0064] The present scheme firstly acquires a plurality of frames of images containing a dynamic gesture in a dynamic gesture recognition mode; then calls an action recognition model to perform feature transmission fusion and classification on the plurality of frames of images, through feature transmission fusion, the feature extraction of time dimension and space dimension (referred to as spatio-temporal dimension feature in the present specification) can be realized at the same time and classification is performed accordingly, so that the spatio-temporal dimension calculation is not needed to be performed multiple times layer by layer by using a dimension transformation layer, a convolution layer, a batch normalization BN layer, a rectified linear unit ReLu layer, a maximum pooling layer and a feature joint layer, etc., thus a large amount of calculation resources can be saved, and then the limited resource scene of the intelligent vehicle-mounted device can be adapted. In addition, the action category corresponding to the dynamic gesture is determined from the action categories of the plurality of frames of images obtained through classification, which can also ensure the accuracy of action recognition.
[0065] The above description is only a summary of the technical scheme of the present specification, in order to enable the technical means of the present specification to be implemented according to the content of the specification, and in order to make the above and other purposes, features and advantages of the present specification more obvious and easy to understand, the following specific embodiments of the present specification are described. BRIEF DESCRIPTION OF DRAWINGS
[0066] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments and are not meant to limit the present specification. Moreover, the same reference numerals are used throughout the various drawings to designate identical parts. In the drawings:
[0067] Figure 1 A process diagram of dynamic gesture recognition according to one embodiment of the present specification is shown;
[0068] Figure 2 An implementation schematic diagram of feature transmission fusion according to one embodiment of the present specification is shown;
[0069] Figure 3 An implementation process diagram according to one example of the present specification is shown;
[0070] Figure 4 A structural schematic diagram of an intelligent vehicle-mounted device according to one embodiment of the present specification is shown;
[0071] Figure 5 A schematic diagram of an electronic device according to one embodiment of the present specification is shown. DETAILED DESCRIPTION
[0072] Exemplary embodiments of the present disclosure will be described in greater detail below with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be thoroughly understood and fully conveyed to those skilled in the art.
[0073] The current dynamic gesture recognition scheme is not suitable for the limited resource scenario of the intelligent vehicle device due to the huge consumption of computing resources. Therefore, the embodiment of the present specification provides a dynamic gesture recognition method, which is mainly suitable for intelligent vehicle devices. In the scheme, first, in the dynamic gesture recognition mode, a plurality of frames of images containing dynamic gestures are obtained; then an action recognition model is called to perform feature transmission fusion and classification on the plurality of frames of images. Through feature transmission fusion, time dimension and space dimension features (referred to as space-time dimension features in the present specification) can be extracted at the same time and classified accordingly, so that the space-time dimension calculation is not performed multiple times layer by layer using the dimension transformation layer, the convolution layer, the batch normalization BN layer, the rectified linear unit ReLu layer, the maximum pooling layer and the feature joint layer, etc. Therefore, a large amount of computing resources can be saved, and the limited resource scenario of the intelligent vehicle device can be adapted. In addition, the action category corresponding to the dynamic gesture is determined from the action category of the plurality of frames of images obtained by classification, which can also ensure the accuracy of action recognition.
[0074] Based on the above inventive concept, the method of the present specification comprises two stages: a trigger stage and a recognition and classification stage.
[0075] In the trigger stage, the hand action is photographed by the camera module, and then the hand detection model is used to detect whether the hand action meets the trigger condition. If the trigger condition is met, the dynamic gesture recognition mode is started, and the recognition and classification stage is entered. Since the computing resources of the intelligent vehicle device are very limited, the embodiment sets a trigger condition, and only when the trigger condition is met can the recognition and classification stage be switched to for action recognition, thereby avoiding a large amount of useless recognition occupying computing resources and reducing the load of the vehicle end to a certain extent. It is worth noting that the hand action of the present embodiment is a hand action in the air, such as a hand hovering action in the car, thereby avoiding the distraction from the road due to the operation of the central control, screen and button, etc., which brings safety hazards.
[0076] In the specific implementation process, the hand action is photographed by using the camera module to obtain a hand action image. Optionally, the hand action can be a hand hovering action, such as a hand hovering action after the palm is unfolded (or clenched). Specifically, the monocular camera module is used to photograph the hand action in this embodiment. By using a monocular camera module as an input source, no additional hardware devices are required, so the cost of the entire device is relatively low, and the image collected by the monocular camera module occupies less resources, and the processing is faster, which can improve the processing efficiency.
[0077] Further, the hand action image is input into the hand detection model for processing. During the processing, there are various processing methods, and two processing methods are provided in this embodiment for illustration, but do not form a limitation. Processing method 1: the hand detection model detects whether the hovering time of the hand action exceeds a preset time, for example, exceeds 0.5s. If it exceeds, it means that the hand action meets the triggering condition, and then enters the recognition and classification stage, thereby reducing the load of the vehicle side to a certain extent. Of course, in addition to detecting the hovering time of the hand, the predetermined action can also be used as a triggering condition to start the dynamic gesture recognition mode. Processing method 2: the hand detection model detects whether the hand action is a predetermined gesture (such as a hand hovering action), and if it is a predetermined gesture, it enters the recognition and classification stage, thereby reducing the load of the vehicle side to a certain extent.
[0078] In addition, the hand detection model is developed using a lightweight backbone framework to adapt to the limited computing resources on the vehicle system, and of course other models can also be used. For example, the hand detection model uses a lightweight slim-net version of yolox. In the training process of the hand detection model, a plurality of video segments labeled with action start frames and action end frames are first obtained, then the video segments are converted into image data, and the image data with label information is used to train the selected initial model, and then the hand detection model is obtained. The training process of the action recognition model is similar, and will not be described in detail. In actual application, the hand detection model can be configured into the palm detector to detect the hand action in the triggering stage, and the action recognition model can be configured into the action recognizer to recognize the dynamic gesture. When the palm detector detects that the hand action meets the triggering condition, the action recognizer is triggered to start, so as to reduce the load of the vehicle side.
[0079] In some optional embodiments, the palm detector and the motion recognizer jointly correspond to a cooling period, and during the cooling period, the palm detector and the motion recognizer do not respond to detection of the hand motion and recognition of the dynamic gesture. If the palm detector and the motion recognizer are not in the cooling period, it indicates that both are in a normal working state, and the following detection process is performed: when the palm detector detects that the hand motion time is 0.5s-1s, the dynamic gesture recognition mode is triggered to start, and the recognition classification stage is entered. At this time, it is first judged whether the motion recognizer is timed out, that is, whether the working time of the motion recognizer exceeds the preset time, and if it is timed out, it is controlled to enter the cooling period. If it is not timed out, the motion recognition model is started for classification and recognition, and after the recognition is completed, the dynamic gesture recognition mode is exited to enter the cooling period, so as to avoid the misrecognition caused by the interference of the remaining motion.
[0080] And in the implementation scheme in the recognition classification stage, refer to Figure 1 , the following steps are performed:
[0081] Step 101, in the dynamic gesture recognition mode, a plurality of frames of images containing a dynamic gesture are acquired.
[0082] In this embodiment, a monocular camera module is used for shooting to acquire the plurality of frames of images. By using a monocular camera module as an input source, no additional hardware device is needed, so the cost of the entire device is low, and the images collected by the monocular camera module occupy less resources and are processed more quickly, which can improve the processing efficiency. The plurality of frames of images can be arranged and merged into a video stream containing a dynamic gesture. The specific form of the dynamic gesture in this embodiment is various, such as grabbing, releasing, clicking, waving forward and backward, sliding left and right, etc., and is not limited to gestures that are easy to recognize. It is worth noting that the dynamic gesture of this embodiment is a hand motion in the air, so as to avoid the safety hazard caused by the distraction of attention from the road due to the operation of the central control, the screen and the button, etc.
[0083] Step 102, calling a motion recognition model to perform feature transfer fusion and classification on the plurality of frames of images to obtain the motion categories of the plurality of frames of images.
[0084] Before calling the motion recognition model for classification, it is necessary to judge in advance whether the motion recognition model classification is timed out, and if it is timed out, it is controlled to enter the cooling period. If it is not timed out, the motion recognition model is started for classification and recognition, and after the recognition is completed, the dynamic gesture recognition mode is exited to enter the cooling period, so as to avoid the misrecognition caused by the interference of the remaining motion.
[0085] The feature transmission and fusion is to transmit and fuse the features of the previous adjacent frame image into the current frame image, so that the features of the current frame image and the features of the previous adjacent frame image have a correlation in the time dimension, and the features in the current frame image can have both the time dimension and the space dimension without a large amount of calculation, thereby realizing the feature extraction of the current frame image in the time dimension and the space dimension.
[0086] The feature transmission and fusion is to transmit and fuse the features of the previous adjacent frame image into the current frame image, so that the features of the current frame image and the features of the previous adjacent frame image have a correlation in the time dimension, and the features in the current frame image can have both the time dimension and the space dimension without a large amount of calculation, thereby realizing the feature extraction of the current frame image in the time dimension and the space dimension.
[0087] In order to adapt to the limited resource scene of the intelligent vehicle-mounted device, the embodiment adopts the feature transmission and fusion manner to process the plurality of frame images to obtain the space-time dimension features of each frame image. In the process of feature transmission and fusion and classification of the plurality of frame images in the embodiment, each frame image is treated as a current frame image during processing, so that the following processing is performed on each frame image: the space dimension features of the previous adjacent frame image are associated with the space dimension features in the current frame image to perform feature transmission and fusion in the time dimension, and the space-time dimension features of the current frame image are obtained. Specifically, the space dimension features of the current frame image can be obtained by performing convolution processing on the current frame image. The number of times of convolution processing is not limited in the embodiment, and the obtained space dimension features include space information of width and height. Further, part of the space dimension features in the previous adjacent frame image is used to replace part of the space dimension features in the current frame image, so that the feature transmission in the time dimension is realized, and the space-time dimension features of the current frame image are obtained. When classifying, convolution calculation is performed on the space-time dimension features of the current frame image to obtain the action category of the current frame image. In this way, the classification of the current frame image can be realized. By performing the above processing on each frame image, the action category of each frame image can be obtained, and the action categories of the plurality of frame images are obtained.
[0088] The replaced part of the space dimension features in the current frame image can be processed as follows: the replaced part of the space dimension features in the current frame image is stored to replace part of the space dimension features in the subsequent adjacent frame image. Of course, this processing process can also be implemented before the replacement, for example, part of the space dimension features of the current frame image is extracted and stored first, and then part of the space dimension features in the previous adjacent frame image is used to replace part of the space dimension features in the current frame image.
[0089] In order to facilitate the description and explanation of the above scheme, the following refers to Figure 2 for illustration.
[0090] In N frames of images, the value of N is not limited in this application. t , the t+1th frame image F t+1 , t+2th frame image as an example. t is any frame image in N frames, and F t+1 It's F t The next adjacent frame image.
[0091] The t-th frame image F t The corresponding spatial dimension features are extracted through multi-layer convolution calculations. The spatial dimension features exist in the form of a matrix. t Some spatial dimension features in Save in the computer cache; use the previously saved partial feature matrix Replace some spatial dimension features To join F t The spatial dimension characteristics of F are obtained based on this. t At this point, multiple layers of convolution are used to calculate F t The spatiotemporal dimension features can be used to obtain the classification result y t .
[0092] When the t+1 frame image F t+1 The corresponding spatial dimension features are extracted through multi-layer convolution calculations. The spatial dimension features exist in the form of a matrix. t+1 Some spatial dimension features in Save in the computer cache, use F t Partial feature matrix Replace some spatial dimension features To join F t+1 The spatial dimension characteristics of F are obtained based on this. t+1 The spatiotemporal dimension features of . Then through multiple layers of convolution to calculate F t+1 The spatiotemporal dimension features can be used to obtain the classification result y t+1 .
[0093] Specifically, convolution of each image frame yields spatial features, including spatial width and height information. By transferring and fusing the temporal features of each frame, we obtain spatiotemporal features for each frame. This allows for simultaneous temporal and spatial feature extraction within the same classification model, enabling dynamic action classification. Compared to traditional spatiotemporal classification methods such as 3DCNN / RNN / LSTM, this example utilizes per-frame feature transfer and fusion, eliminating the need for extensive computation to extract features in both temporal and spatial dimensions. This approach reduces computational complexity, resource usage, and efficiency, making it suitable for resource-constrained scenarios in smart vehicles.
[0094] It is worth noting that for the first frame image in the plurality of frame images, an image is obtained, the spatial dimension features of which exist in the form of a matrix, and part of the spatial dimension features (the value can be 0) in the image are extracted in the aforementioned manner to replace the part of the spatial dimension features at the corresponding position in the first frame image, so as to realize the feature extraction of the first frame image in the time and spatial dimensions.
[0095] Step 103, determining the action category corresponding to the dynamic gesture from the action categories of the plurality of frame images.
[0096] After classification, each frame image in the plurality of frame images corresponds to a respective action category, so the target action category is selected from the action categories of the plurality of frame images as the action category corresponding to the dynamic gesture. Among them, the number of images corresponding to the target action category is above the preset number threshold, and further, the number of images corresponding to the target action category is the most. Specifically, since each frame image in the plurality of frame images has an action category, if most frame images correspond to the same action category, the action category is taken as the target action category. For example, if the number of frame images corresponding to the action category is the most, the action category is taken as the target action category. For example, if 15 frame images out of 20 frame images have the same action category, the action category is taken as the target action category.
[0097] In some optional embodiments, a time window is determined, which is a time window corresponding to a set number of frames. For example, 10 frames correspond to a time window, and of course other frame numbers can also be set. The action category corresponding to the time window is determined from the action categories of the plurality of frame images according to the time window. And the target action category is determined based on the action category corresponding to the time window, for example, the action category with the most frame numbers in the time window is determined as the target action category. And the action category corresponding to the dynamic gesture is determined based on the target action category, for example, the target action category is directly taken as the action category corresponding to the dynamic gesture, so as to improve the accuracy of action recognition.
[0098] In some optional embodiments, after the time window is determined, the action class corresponding to each time window of each frame segment is determined from the action classes of the plurality of frames of images based on the time window. For example, the example has 10 frames of images, and 5 frames correspond to a time window. The action class is determined for each time window of 1-5 frames, 2-6 frames, 3-7 frames, 4-8 frames, and 5-10 frames. Further, the target action class corresponding to each time window of each frame segment is determined based on the action class corresponding to each time window of each frame segment. For example, the action class with the most frames in each time window of each frame segment is determined as the target action class. Thus, each time window of each frame segment has a corresponding target action class. The action class corresponding to the dynamic gesture is determined based on the target action class corresponding to each time window of each frame segment. Specifically, the target action classes corresponding to each time window of each frame segment are compared with each other. If the target action classes corresponding to each time window of each frame segment are consistent, the target action class corresponding to a time window of a frame segment is randomly determined as the action class corresponding to the dynamic gesture. If the target action classes corresponding to each time window of each frame segment are inconsistent, the target action class with the most frames is determined as the action class corresponding to the dynamic gesture, thereby improving the accuracy of action recognition.
[0099] In some optional embodiments, after the action class corresponding to the dynamic gesture is determined from the action classes of the plurality of frames of images, the action instruction corresponding to the dynamic gesture is transmitted to a downstream object for execution, and then the dynamic gesture recognition mode is exited to avoid misrecognition caused by interference of remaining actions.
[0100] It is worth noting that the applicable objects of the above-mentioned solutions include but are not limited to RGB images and IR infrared images.
[0101] In order to further illustrate and explain the solutions of the present specification, the following specific examples are used for illustration. Figure 3
[0102] When the user is driving a vehicle, the intelligent vehicle-mounted device is in a working state and the monocular camera captures the user in real time.
[0103] When the user makes a hand hovering action in the air, S301 is performed, the monocular camera captures the hand hovering action and transmits it to the palm detector.
[0104] S302, determining whether the palm detector is in a cooling time. If yes, no response is made. If no, S303 is executed. Optionally, in this step, it can be simultaneously determined whether the palm detector and the motion recognizer are both in the cooling time. If neither of them is in the cooling time, it indicates that both of them can normally respond, and then S303 is executed.
[0105] S303, detecting whether the hand hovering motion meets a triggering condition.
[0106] If no, no response is made.
[0107] If yes, S304 is executed, and a dynamic gesture recognition mode is started.
[0108] S305, determining whether the motion recognizer is in a timeout. If yes, S306 is executed. If no, S307 is executed.
[0109] S306, controlling the motion recognizer to enter a cooling period.
[0110] S307, transmitting a plurality of frames of images containing a dynamic gesture captured by the monocular camera to the motion recognizer. Specifically, the plurality of frames of images are acquired by the monocular camera when the user makes the dynamic gesture in the air.
[0111] S308, calling a motion recognition model in the motion recognizer to perform feature transfer fusion and classification on the plurality of frames of images, to obtain motion categories of the plurality of frames of images.
[0112] S309, determining a time window, and determining a motion category corresponding to the time window from the motion categories of the plurality of frames of images according to the time window.
[0113] S310, determining a motion category with the most frames in the time window as a dynamic category of the dynamic gesture.
[0114] S311, exiting the dynamic gesture recognition mode, and executing S306 to control the motion recognizer to enter a cooling state.
[0115] The above scheme needs to detect a hand hovering motion first, and then recognize a dynamic gesture. However, for a user, after the user makes a hand hovering motion in the air and a subsequent dynamic gesture, the intelligent vehicle-mounted device can detect, recognize, and timely execute a corresponding instruction in real time, so as to timely respond to the user, and can bring a smooth operation experience.
[0116] The above embodiment introduces a specific implementation process of the dynamic gesture recognition of the present specification. Based on the same inventive concept as in the foregoing embodiment, the present specification also provides an intelligent vehicle-mounted device.
[0117] The intelligent vehicle-mounted device in the present embodiment refers to Figure 4 , which comprises:
[0118] The camera module 401 is configured to acquire a plurality of frames of images containing a dynamic gesture in a dynamic gesture recognition mode.
[0119] The space-time conversion module 402 is configured to perform feature transmission fusion and classification on the plurality of frames of images by calling an action recognition model, to obtain an action category of the plurality of frames of images.
[0120] The determination module 403 is configured to determine an action category corresponding to the dynamic gesture from the action category of the plurality of frames of images.
[0121] In some optional embodiments, the intelligent vehicle-mounted device further includes:
[0122] The camera module 401 is configured to capture a hand action.
[0123] The hand detection module is configured to detect, by using a hand detection model, whether the hand action meets a triggering condition.
[0124] If the hand action meets the triggering condition, the dynamic gesture recognition mode is started.
[0125] In some optional embodiments, the hand detection module is specifically configured to:
[0126] detect whether a hovering time of the hand action exceeds a preset time.
[0127] If the hovering time exceeds the preset time, it indicates that the hand action meets the triggering condition.
[0128] In some optional embodiments, the space-time conversion module 402 is specifically configured to:
[0129] perform feature transmission fusion on spatial dimension features of a previous adjacent frame image in the plurality of frames of images and spatial dimension features in a current frame image in a time dimension, to obtain space-time dimension features of the current frame image.
[0130] perform convolution calculation on the space-time dimension features of the current frame image, to obtain an action category of the current frame image.
[0131] In some optional embodiments, the space-time conversion module 402 is specifically configured to:
[0132] perform convolution processing on the current frame image, to obtain spatial dimension features of the current frame image.
[0133] replace part of the spatial dimension features of the current frame image with part of spatial dimension features in the previous adjacent frame image, to obtain space-time dimension features of the current frame image.
[0134] In some optional implementations, the intelligent vehicle-mounted device further includes:
[0135] The storage module is used to store the replaced part of the spatial dimension features in the current frame image to replace the part of the spatial dimension features of the subsequent adjacent frame image.
[0136] In some optional implementations, the spatiotemporal conversion module 402 is specifically configured to:
[0137] A target action category is selected from the action categories of the plurality of frames of images as the action category corresponding to the dynamic gesture; wherein the number of images corresponding to the target action category is above a preset number threshold.
[0138] In some optional implementations, the determining module 403 is specifically configured to:
[0139] Determine a time window, where the time window is a time window corresponding to a set number of frames;
[0140] Determining, according to the time window, from the action categories of the plurality of frames of image, an action category corresponding to the time window;
[0141] The target action category is determined based on the action category corresponding to the time window, and the action category corresponding to the dynamic gesture is determined based on the target action category.
[0142] In some optional implementations, the intelligent vehicle-mounted device further includes:
[0143] A transmission module, used to transmit the action instructions corresponding to the dynamic gesture to the downstream object for execution;
[0144] Exit module, used to exit dynamic gesture recognition mode.
[0145] Based on the same inventive concept as in the aforementioned embodiment, an embodiment of this specification further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the aforementioned methods when executed by a processor.
[0146] Based on the same inventive concept as in the above embodiments, the embodiments of this specification further provide an electronic device, such as Figure 5 As shown, it includes a memory 504, a processor 502 and a computer program stored in the memory 504 and executable on the processor 502. When the processor 502 executes the program, the steps of any of the above-mentioned methods are implemented.
[0147] Among them, Figure 5In particular embodiments, a bus architecture, represented generally by the bus 500, can include any number of interconnected buses and bridges, the bus 500 linking together various circuits such as the processor 502 represented by one or more processors and the memory 504 represented by the memory. The bus 500 can also link together various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and thus, not described further. The bus interface 505 provides an interface between the bus 500 and the receiver 501 and the transmitter 503. The receiver 501 and the transmitter 503 can be the same component, i.e., a transceiver, providing a means for communicating with various other terminal devices on the transmission medium. The processor 502 is responsible for managing the bus 500 and general processing, while the memory 504 can be used for storing data used by the processor 502 in executing operational processes.
[0148] Through one or more embodiments of the present specification, the present specification has the following beneficial effects or advantages:
[0149] The scheme can realize feature extraction in time dimension and space dimension (referred to as space-time dimension feature in the present specification) at the same time through feature transfer and fusion of a plurality of frame images containing dynamic gestures, and classify according to the feature transfer and fusion, so as to avoid a large amount of calculation of 3DCNN / RNN / LSTM and its variants, and without calculation of three-dimensional point cloud, so it is very friendly to end-side calculation, can save a large amount of calculation resources and ensure calculation efficiency to achieve real-time inference, and further adapt to the limited resource scene of intelligent vehicle equipment.
[0150] Further, the dynamic gestures that can be recognized by the present scheme are more diverse, such as grabbing, releasing, clicking, waving forward and backward, sliding left and right, etc.
[0151] Further, the gestures recognized by the present scheme are all empty-handed gestures, so as to avoid distraction from the road due to operation of the central control, screen and button, etc., and bring safety hazards.
[0152] Further, the camera module of the present scheme adopts a monocular camera module to reduce cost and improve processing efficiency, and the hand detection model adopts a lightweight backbone framework to adapt to the limited computing power of the vehicle-mounted system.
[0153] The algorithms and displays presented herein are not inherently related to any particular computer, virtual system, or other apparatus. Various general purpose systems can be used with programs in accordance with the teachings herein, or it can prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will be apparent from the description above. In addition, the present description is not intended to be limited to any particular programming language. It will be appreciated that a variety of programming languages can be used to implement the teachings of the description herein, and any references below to specific languages are provided for disclosure of enablement of the best mode of the description.
[0154] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the description can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been described in detail in order to not obscure the understanding of this description.
[0155] Similarly, it is to be understood that the description of exemplary embodiments of the description herein is intended to be illustrative, but not limiting, of the scope of the application as set forth in the following claims. Thus, this description is not intended to be complete without the detailed description set forth in the following claims. In particular, although many of the examples presented herein involve specific combinations of method acts or system elements, it should be understood that those acts and those elements can be combined in other ways to accomplish the same objectives.
[0156] Those skilled in the art will appreciate that the modules in the apparatuses in the embodiments can be adapted and placed in one or more apparatuses other than the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and further can be divided into more sub-modules or sub-units or sub-components. Any combination of all the features disclosed in the description (including accompanying claims, abstract and drawings) and any method or apparatus so disclosed can be made, except that at least some of such features and / or processes or units are mutually exclusive. Each feature disclosed in the description (including accompanying claims, abstract and drawings) can be replaced by alternative features serving the same, equivalent or similar purpose, unless expressly stated otherwise.
[0157] Furthermore, those skilled in the art will recognize that, while certain embodiments described herein include certain features that are not included in other embodiments, combinations of features of the different embodiments are meant to be within the scope of the present description and form different embodiments. For example, in the claims below, any of the claimed embodiments can be used in any combination.
[0158] Various component embodiments of the present description can be implemented in hardware, or as software modules running in one or more processors, or combinations thereof. Those skilled in the art will appreciate that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functionality of some or all of the components of the gateway, the proxy server, the system according to the embodiments of the present description. The present description can also be implemented as a program (e.g., computer program and computer program product) for performing part or all of the methods described herein. Such a program implementing the present description can be stored on a computer readable medium, or can have one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier medium, or in any other form.
[0159] It should be noted that the above-mentioned embodiments illustrate rather than limit the present description, and that one skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word 'comprising' does not exclude the presence of elements or steps not listed in a claim. The word 'a' or 'an' preceding an element does not exclude the presence of a plurality of such elements. The disclosure can be implemented by means of both hardware and software, and any combination thereof. In a unit claim, several devices can be listed with a conjunction like 'or', but it is to be understood that a combination of these devices can be used in the embodiments of the present description. The use of the word 'at least' followed by a list of one or more items does not exclude additional such items. The use of the words 'one' or 'only one' with a list of one or more items does not exclude additional such items. The scope of the description is not limited by the embodiments described herein but only by the claims.
Claims
1. A dynamic gesture recognition method, the method being used in an intelligent vehicle-mounted device, the method comprising: In the dynamic gesture recognition mode, a plurality of frames of images containing dynamic gestures are acquired; Calling an action recognition model to perform feature transfer fusion and classification on the plurality of frame images to obtain action categories of the plurality of frame images, specifically comprising: in the plurality of frame images, associating the spatial dimension features of the preceding frame images with the spatial dimension features of the current frame image to perform feature transfer fusion on the time dimension to obtain the spatiotemporal dimension features of the current frame image; performing convolution processing on the current frame image to obtain the spatial dimension features of the current frame image; replacing part of the spatial dimension features of the current frame image with part of the spatial dimension features of the preceding frame image to obtain the spatiotemporal dimension features of the current frame image; performing convolution calculation on the spatiotemporal dimension features of the current frame image to obtain the action category of the current frame image; The action category corresponding to the dynamic gesture is determined from the action categories of the plurality of frames of images.
2. The method according to claim 1, before acquiring the plurality of frames of images containing dynamic gestures, the method further comprises: Use a camera module to capture hand movements; Using a hand detection model to detect whether the hand movement meets the triggering conditions; If so, the dynamic gesture recognition mode is activated.
3. The method according to claim 2, wherein the detecting whether the hand motion meets the triggering condition using the hand detection model specifically comprises: Detecting whether the hovering time of the hand movement exceeds a preset time; If so, it means that the hand movement meets the triggering condition.
4. The method according to claim 1, further comprising: after replacing the partial spatial dimensional features of the current frame image with the partial spatial dimensional features in the previous frame image; The replaced portion of the spatial dimension features in the current frame image is stored to replace the portion of the spatial dimension features in the subsequent adjacent frame image.
5. The method according to claim 1, wherein determining the action category corresponding to the dynamic gesture from the action categories of the plurality of frames of images specifically comprises: A target action category is selected from the action categories of the plurality of frames of images as the action category corresponding to the dynamic gesture; wherein the number of images corresponding to the target action category is above a preset number threshold.
6. The method according to claim 1, wherein determining the action category corresponding to the dynamic gesture from the action categories of the plurality of frames of images specifically comprises: Determine a time window, where the time window is a time window corresponding to a set number of frames; Determining, according to the time window, from the action categories of the plurality of frames of image, an action category corresponding to the time window; A target action category is determined based on the action category corresponding to the time window, and an action category corresponding to the dynamic gesture is determined based on the target action category.
7. The method according to claim 1, after determining the action category corresponding to the dynamic gesture from the action categories of the plurality of frames of images, the method further comprises: Transmitting the action instruction corresponding to the dynamic gesture to the downstream object for execution; Exit dynamic gesture recognition mode.
8. An intelligent vehicle-mounted device comprising: A camera module is used to acquire a plurality of frames of images containing dynamic gestures in a dynamic gesture recognition mode; The spatiotemporal conversion module is used to call the action recognition model to perform feature transfer fusion and classification on the multiple frames of images to obtain the action categories of the multiple frames of images, and is specifically used to: in the multiple frames of images, associate the spatial dimension features of the previous frame images with the spatial dimension features of the current frame image to perform feature transfer fusion on the time dimension to obtain the spatiotemporal dimension features of the current frame image; wherein, convolution processing is performed on the current frame image to obtain the spatial dimension features of the current frame image; partial spatial dimension features of the current frame image are replaced by partial spatial dimension features of the previous frame image to obtain the spatiotemporal dimension features of the current frame image; and convolution calculation is performed on the spatiotemporal dimension features of the current frame image to obtain the action category of the current frame image. The determination module is configured to determine the action category corresponding to the dynamic gesture from the action categories of the plurality of frames of images.
9. The intelligent vehicle-mounted device according to claim 8, further comprising: The camera module is used to capture hand movements; A hand detection module, configured to detect whether the hand action meets the triggering conditions using a hand detection model; If so, the dynamic gesture recognition mode is activated.
10. The intelligent vehicle-mounted device according to claim 9, wherein the hand detection module is specifically configured to: Detecting whether the hovering time of the hand movement exceeds a preset time; If so, it means that the hand movement meets the triggering condition.
11. The intelligent vehicle-mounted device according to claim 8, further comprising: The storage module is used to store the replaced part of the spatial dimension features in the current frame image to replace the part of the spatial dimension features of the subsequent adjacent frame image.
12. The intelligent vehicle-mounted device according to claim 8, wherein the space-time conversion module is specifically configured to: Selecting a target action category from the action categories of the plurality of frames of images as the action category corresponding to the dynamic gesture; wherein, The number of images corresponding to the target action category is above a preset number threshold.
13. The intelligent vehicle-mounted device according to claim 8, wherein the determining module is specifically configured to: Determine a time window, where the time window is a time window corresponding to a set number of frames; Determining, according to the time window, from the action categories of the plurality of frames of image, an action category corresponding to the time window; A target action category is determined based on the action category corresponding to the time window, and an action category corresponding to the dynamic gesture is determined based on the target action category.
14. The intelligent vehicle-mounted device according to claim 8, further comprising: A transmission module, used to transmit the action instructions corresponding to the dynamic gesture to the downstream object for execution; Exit module, used to exit dynamic gesture recognition mode.
15. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
16. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 7 when executing the program.
Citation Information
Patent Citations
Method and system for gesture recognition based on surface electromyogram signals
CN112783327A
Video behavior recognition method and system based on channel attention-oriented time modeling
CN112818843A