Two-handed gesture interaction method and apparatus, electronic device, and storage medium

By acquiring real-time video inside the vehicle cabin and identifying the same hand group and left and right hand matching pairs, the accuracy and response speed issues of in-cabin interaction technology in noisy environments have been solved, achieving efficient two-hand gesture control.

WO2026025648A1PCT designated stage Publication Date: 2026-02-05HANGZHOU RUIJIAN ZHIXING TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/126164
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-31
Filing Date
2024-10-21
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing in-vehicle cabin interaction technologies suffer from reduced recognition accuracy and response speed in noisy environments, lack anti-interference capabilities, and are inconvenient to use.

Method used

The system acquires real-time video of the vehicle cabin using a camera device, uses the vehicle's speed to determine hand information in the target area, identifies the same hand group and left-right hand pairs, determines hand gestures, and controls controllable components of the vehicle.

Benefits of technology

It improves the anti-interference capability and efficiency of interaction in the cabin environment, and enables direct control of vehicle functions through hand gestures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126164_05022026_PF_FP_ABST
    Figure CN2024126164_05022026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a two-handed gesture interaction method and apparatus, an electronic device, and a storage medium. In the present disclosure, by comparing hand information of hands within the current image frame with that within the previous image frame, hand groups belonging to the same hand and matching pairs of left and right hands belonging to the same person are determined; and then, a two-handed gesture of a certain person can be determined by means of a first hand group and a second hand group that correspond to the matching pairs of left and right hands belonging to the same person, and on the basis of the determined two-handed gesture, a controllable component is controlled to perform a corresponding operation. In the described method, control is completed by means of a two-handed gesture. Such a control approach is less affected by the surrounding environment, thereby helping enhance the anti-interference capability during interaction in a cabin environment; moreover, the two-handed gesture-based control approach allows for direct control of functions on a vehicle, thereby helping improve interaction efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

A two-handed gesture interaction method and device, electronic device and storage medium

[0001] Cross Reference to Related Applications

[0002] This application claims priority to the Chinese patent application No. 202411036116.0, filed on July 31, 2024, and entitled "A two-handed gesture interaction method and device, electronic device and storage medium", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] The present disclosure relates to the technical field of image recognition, in particular, to a two-handed gesture interaction method and device, electronic device and storage medium. BACKGROUND

[0004] Currently, in the interaction technology in the vehicle cabin, the interaction is usually realized through physical buttons or touch screens, but this interaction method needs to find the accurate control position before interaction can be performed, and although the voice control system has been supplemented, when there is noise in the vehicle cabin, the accuracy and response speed of voice recognition will be significantly reduced, therefore, there is an urgent need for an interaction method with strong anti-interference ability and convenient interaction.

[0005] SUMMARY

[0006] Therefore, the embodiments of the present disclosure provide a two-handed gesture interaction method and device, electronic device and storage medium to improve the anti-interference ability when interacting in the vehicle cabin environment and to enable efficient interaction.

[0007] In a first aspect, the embodiments of the present disclosure provide a two-handed gesture interaction method, which comprises:

[0008] After obtaining a current frame image of an internal real-time video of a vehicle cabin part of a vehicle, hand information of each hand in a target region corresponding to a driving speed of the vehicle is obtained according to the driving speed;

[0009] According to target hand key points of hand information of each hand in a previous frame image of the current frame image and target hand key points of hand information of each hand in the current frame, a hand group belonging to the same hand in the current frame image and the previous frame image is determined;

[0010] According to the target hand key points of each hand belonging to a first hand in the hand information in the current frame image and the target hand key points of each hand belonging to a second hand in the hand information in the current frame image, a left-right hand matching pair belonging to the same person in the current frame image is determined;

[0011] For each left-right hand matching pair, according to a first hand group and a second hand group corresponding to the left-right hand matching pair in the current frame image and the previous frame image, a double-hand gesture corresponding to the left-right hand matching pair is determined;

[0012] According to the double-hand gesture, a controllable component corresponding to the double-hand gesture in the vehicle is controlled to perform a corresponding operation.

[0013] Optionally, the interior of the vehicle cabin part is provided with a camera device, and the interior real-time video is obtained through the camera device.

[0014] Optionally, the interior of the vehicle cabin part is provided with a camera device close to the top of the head of the vehicle.

[0015] Optionally, the interior of the vehicle cabin part is provided with a camera device equal to the number of seats in the vehicle, and each camera device is used to collect the interior real-time video of the corresponding seat area.

[0016] Optionally, the hand information of each hand in the target area corresponding to the driving speed of the vehicle in the current frame image is obtained according to the driving speed of the vehicle, including:

[0017] When the driving speed of the vehicle is greater than or equal to a preset speed, the hand information of each hand belonging to the co-pilot area and the rear area in the current frame image is obtained;

[0018] When the driving speed of the vehicle is less than the preset speed, the hand information of each hand in all seat areas in the current frame image is obtained.

[0019] Optionally, the hand information of each hand in the target area corresponding to the driving speed of the vehicle in the current frame image is obtained according to the driving speed of the vehicle, including:

[0020] According to a preset resolution, the current frame image is down-sampled to obtain a down-sampled image;

[0021] According to the driving speed of the vehicle, the hand information of each hand in the target area corresponding to the driving speed of the vehicle in the down-sampled image is obtained.

[0022] Optionally, the hand information of each hand in the target area corresponding to the driving speed of the vehicle in the current frame image is obtained, including:

[0023] For each hand in the target area in the current frame image, according to the diagonal coordinates of the hand-enclosing rectangle of the hand, the hand image of the hand is obtained.

[0024] inputting the hand image of the hand into a hand skeleton point detection model to obtain hand key points of the hand and confidence of each hand key point;

[0025] calculating an arithmetic mean of the confidence of each hand key point included in the hand;

[0026] when the corresponding arithmetic mean of the hand is greater than a preset confidence, inputting the hand image of the hand into a hand classifier to determine a handness of the hand;

[0027] The hand information includes hand key points of the hand and a handness of the hand, and the handness includes a left hand and a right hand.

[0028] Optionally, the target hand key points of the hand information of each hand in the previous frame image of the current frame image and the target hand key points of the hand information of each hand in the current frame image are used to determine a hand group belonging to a same hand in the current frame image and the previous frame image, and the method comprises the following steps:

[0029] calculating a first Euclidean distance between the target hand key points of the hand information of each hand in the previous frame image and the target hand key points of the hand information of each hand in the current frame;

[0030] constructing a first Euclidean distance matrix according to the first Euclidean distance;

[0031] performing matching on each hand in the previous frame image and each hand in the current frame by using a Hungarian algorithm according to the first Euclidean distance matrix, to determine a hand group belonging to a same hand in the current frame image and the previous frame image.

[0032] Optionally, the target hand key points of each hand belonging to a first handness in the hand information in the current frame image and the target hand key points of each hand belonging to a second handness in the hand information in the current frame image are used to determine a left-right hand matching pair in the current frame image, and the method comprises the following steps:

[0033] calculating a square value of a second Euclidean distance between the target hand key points of each hand belonging to the first handness in the hand information in the current frame image and the target hand key points of each hand belonging to the second handness in the hand information in the current frame image;

[0034] constructing a second Euclidean distance matrix according to the square value of the second Euclidean distance;

[0035] According to the second Euclidean distance matrix, the Hungarian algorithm is used to match each hand belonging to the first hand type in the current frame image and each hand belonging to the second hand type in the current frame image, so as to determine the left-right hand matching pair belonging to the same person in the current frame image.

[0036] Optionally, the method further comprises:

[0037] When the left-right hand matching pair belonging to the same person in the previous frame image is matched with the same hand group in the current frame image, the hand corresponding to the left-right hand matching pair belonging to the same person in the previous frame image in the current frame image is determined as the left-right hand matching pair belonging to the same person in the current frame image.

[0038] Optionally, the method further comprises:

[0039] For each hand included in the current frame image, when the hand is matched successfully, the hand is marked using the identifier configured for the hand matched in the previous frame image;

[0040] When the hand is not matched successfully, the hand is marked using a target identifier, which is different from the identifier configured for each hand included in the previous frame image.

[0041] Optionally, when the left-right hand matching pair belonging to the same person in the previous frame image is matched with the same hand group in the current frame image, the hand corresponding to the left-right hand matching pair belonging to the same person in the previous frame image in the current frame image is determined as the left-right hand matching pair belonging to the same person in the current frame image, comprising:

[0042] The identifier included in the left-right hand matching pair belonging to the same person in the previous frame image is searched in the identifiers configured for each hand included in the current frame image to determine whether the identifier included in the left-right hand matching pair belonging to the same person in the previous frame image exists;

[0043] If the identifier exists, the hand corresponding to the identifier identical to the identifier included in the left-right hand matching pair belonging to the same person in the previous frame image in the previous frame image is determined as the left-right hand matching pair belonging to the same person in the current frame image.

[0044] Optionally, the determination of the double-hand gesture corresponding to each left-right hand matching pair according to the first hand group and the second hand group corresponding to the left-right hand matching pair in the current frame image and the previous frame image comprises:

[0045] The sum of the first movement distance of the target hand key point of each hand included in the first hand group and the second movement distance of the target hand key point of each hand included in the second hand group is calculated.

[0046] calculate a ratio of the sum and a sampling interval length of the internal real-time video, to take the ratio as a moving speed of the corresponding double-hand gesture of the left-right hand matching pair;

[0047] when the moving speed is less than or equal to a preset moving speed, input the current frame image into a double-hand static gesture classification model to determine a corresponding double-hand static gesture of the left-right hand matching pair;

[0048] when the moving speed is greater than the preset moving speed, input a preset number of image frames containing the left-right hand matching pair into a double-hand dynamic gesture classification model to determine a corresponding double-hand dynamic gesture of the left-right hand matching pair.

[0049] Optionally, the controllable component includes:

[0050] a vehicle-mounted screen, a sunroof, a vehicle-mounted player, a vehicle-mounted camera, and cabin light.

[0051] In a second aspect, the embodiments of the present disclosure provide a double-hand gesture interaction device, and the device includes:

[0052] an acquisition unit configured to, after acquiring a current frame image of an internal real-time video of a cabin part of a vehicle, acquire hand information of each hand in a target region corresponding to a driving speed of the vehicle in the current frame image according to the driving speed of the vehicle;

[0053] a first determination unit configured to determine a hand group belonging to a same hand in the current frame image and a previous frame image of the current frame image according to target hand key points of each hand in the hand information in the current frame image and target hand key points of each hand in the hand information in the previous frame image of the current frame image;

[0054] a second determination unit configured to determine a left-right hand matching pair belonging to a same person in the current frame image according to the target hand key points of each hand belonging to a first hand in the hand information in the current frame image and the target hand key points of each hand belonging to a second hand in the hand information in the current frame image;

[0055] a gesture recognition unit configured to, for each left-right hand matching pair, determine a corresponding double-hand gesture of the left-right hand matching pair according to a first hand group and a second hand group corresponding to the left-right hand matching pair in the current frame image and the previous frame image;

[0056] a control unit configured to control a controllable component corresponding to the double-hand gesture in the vehicle to perform a corresponding operation according to the double-hand gesture.

[0057] Optionally, the interior of the vehicle cabin part is provided with a camera, and the real-time video of the interior is obtained through the camera.

[0058] Optionally, the interior of the vehicle cabin part is provided with a camera near the top of the vehicle head.

[0059] Optionally, the interior of the vehicle cabin part is provided with a camera equal to the number of seats in the vehicle, and each camera is used to collect the real-time video of the corresponding seat area.

[0060] Optionally, when the acquisition unit is used to acquire hand information of each hand in a target area corresponding to the driving speed of the vehicle in the current frame image according to the driving speed of the vehicle, the acquisition unit comprises:

[0061] When the driving speed of the vehicle is greater than or equal to a preset speed, the hand information of each hand belonging to the co-driver area and the rear area in the current frame image is acquired.

[0062] When the driving speed of the vehicle is less than the preset speed, the hand information of each hand in all seat areas in the current frame image is acquired.

[0063] Optionally, when the acquisition unit is used to acquire hand information of each hand in a target area corresponding to the driving speed of the vehicle in the current frame image according to the driving speed of the vehicle, the acquisition unit comprises:

[0064] The current frame image is down-sampled according to a preset resolution to obtain a down-sampled image;

[0065] According to the driving speed of the vehicle, the hand information of each hand in a target area corresponding to the driving speed in the down-sampled image is acquired.

[0066] Optionally, when the acquisition unit is used to acquire hand information of each hand in a target area corresponding to the driving speed of the vehicle in the current frame image, the acquisition unit comprises:

[0067] For each hand in the target area in the current frame image, a hand image of the hand is obtained according to the diagonal coordinates of a hand circumscribed rectangle of the hand.

[0068] The hand image of the hand is input into a hand skeleton point detection model to obtain hand key points of the hand and confidence of each hand key point.

[0069] An arithmetic mean value of the confidence of each hand key point included in the hand is calculated.

[0070] When the arithmetic mean value corresponding to the hand is greater than a preset confidence, the hand image of the hand is input into a hand classifier to determine the chirality of the hand.

[0071] The hand information includes hand key points of the hand and a hand type of the hand, and the hand type includes a left hand and a right hand.

[0072] Optionally, the first determination unit is configured to, when determining that the hand groups in the current frame image and the previous frame image belong to the same hand, include:

[0073] calculate first Euclidean distances between the target hand key points of the hand information of each hand in the previous frame image and the target hand key points of the hand information of each hand in the current frame;

[0074] construct a first Euclidean distance matrix according to the first Euclidean distances;

[0075] match each hand in the previous frame image and each hand in the current frame by using a Hungarian algorithm according to the first Euclidean distance matrix, to determine that the hand groups in the current frame image and the previous frame image belong to the same hand.

[0076] Optionally, the second determination unit is configured to, when determining that the left-right hand matching pairs in the current frame image belong to the same person, include:

[0077] calculate square values of second Euclidean distances between the target hand key points of each hand belonging to a first hand type in the hand information in the current frame image and the target hand key points of each hand belonging to a second hand type in the hand information in the current frame image;

[0078] construct a second Euclidean distance matrix according to the square values of the second Euclidean distances;

[0079] match each hand belonging to the first hand type in the current frame image and each hand belonging to the second hand type in the current frame image by using a Hungarian algorithm according to the second Euclidean distance matrix, to determine that the left-right hand matching pairs in the current frame image belong to the same person.

[0080] Optionally, the second determination unit is further configured to:

[0081] If the left-right hand matching pair belonging to the same person in the previous frame image is matched with the same hand group in the current frame image, the hand corresponding to the same hand matching pair in the previous frame image in the current frame image is determined as the left-right hand matching pair belonging to the same person in the current frame image.

[0082] Optionally, the apparatus further comprises:

[0083] The marking unit is configured to, for each hand included in the current frame image, mark the hand with an identifier configured for the hand matched in the previous frame image if the hand is matched successfully, and mark the hand with a target identifier if the hand is not matched successfully, the target identifier being different from the identifiers configured for each hand included in the previous frame image.

[0084] Optionally, the second determining unit is configured to, when the left-right hand matching pair belonging to the same person in the previous frame image is matched with the same hand group in the current frame image, determine the hand corresponding to the same hand matching pair in the previous frame image in the current frame image as the left-right hand matching pair belonging to the same person in the current frame image, including:

[0085] According to the identifier included in the left-right hand matching pair belonging to the same person in the previous frame image, searching for whether there is the identifier included in the left-right hand matching pair belonging to the same person in the previous frame image in the identifiers configured for each hand included in the current frame image;

[0086] If there is, determining the hand corresponding to the same identifier in the current frame image as the left-right hand matching pair belonging to the same person in the current frame image.

[0087] Optionally, the second determining unit is configured to, for each left-right hand matching pair, according to the first hand group and the second hand group corresponding to the left-right hand matching pair in the current frame image and the previous frame image, determine the double-hand gesture corresponding to the left-right hand matching pair, including:

[0088] Calculating the sum of the first moving distance of the target hand key point of each hand included in the first hand group and the second moving distance of the target hand key point of each hand included in the second hand group;

[0089] Calculating the ratio of the sum and the sampling interval length of the internal real-time video, so as to take the ratio as the moving speed of the double-hand gesture corresponding to the left-right hand matching pair;

[0090] when the moving speed is less than or equal to the preset moving speed, inputting the current frame image into a double-hand static gesture classification model to determine a double-hand static gesture corresponding to the left-hand matching pair and the right-hand matching pair;

[0091] when the moving speed is greater than the preset moving speed, inputting a preset number of image frames containing the left-hand matching pair and the right-hand matching pair into a double-hand dynamic gesture classification model to determine a double-hand dynamic gesture corresponding to the left-hand matching pair and the right-hand matching pair.

[0092] Optionally, the controllable component comprises:

[0093] a vehicle-mounted screen, a sunroof, a vehicle-mounted player, a vehicle-mounted camera, and cabin light.

[0094] In a third aspect, an electronic device is provided, including a processor and a memory, the memory storing machine executable instructions capable of being executed by the processor, and the processor executes the machine executable instructions to implement the double-hand gesture interaction method according to any one of the first aspect.

[0095] In a fourth aspect, a machine readable storage medium is provided, the machine readable storage medium storing machine executable instructions, and when the machine executable instructions are invoked and executed by a processor, the machine executable instructions cause the processor to implement the double-hand gesture interaction method according to any one of the first aspect.

[0096] The technical solution provided by the embodiments of the present disclosure can have the following beneficial effects:

[0097] In the present disclosure, after a current frame image of an internal real-time video of a cabin part of a vehicle is acquired, hand information of each hand in a target region corresponding to a driving speed of the vehicle in the current frame image is acquired according to the driving speed, then a hand group belonging to the same hand in the current frame image and a previous frame image is determined according to the hand information, and a left-hand matching pair and a right-hand matching pair belonging to the same person in the current frame image is determined according to the hand information, finally, a double-hand gesture of a person can be determined through a first hand group and a second hand group corresponding to the left-hand matching pair and the right-hand matching pair belonging to the same person, and the controllable component is controlled to perform a corresponding operation according to the determined double-hand gesture. In the above method, the control is completed through the double-hand gesture, and this control mode is less affected by the surrounding environment, thus being beneficial to improving the anti-interference capability when interacting in the cabin environment, and since the control mode of the double-hand gesture can directly control the functions on the vehicle, the interaction efficiency is improved.

[0098] In order to make the above objectives, characteristics and advantages of the present disclosure more obvious and easy to understand, the following preferred embodiments are described in detail below, and the accompanying drawings are described as follows. BRIEF DESCRIPTION OF DRAWINGS

[0099] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present disclosure, and therefore should not be considered as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor based on these drawings.

[0100] FIG. 1 is a flow diagram of a two-hand gesture interaction method according to an embodiment of the present disclosure;

[0101] FIG. 2 is a flow diagram of another two-hand gesture interaction method according to an embodiment of the present disclosure;

[0102] FIG. 3 is a schematic diagram of a hand key point according to an embodiment of the present disclosure;

[0103] FIG. 4 is a flow diagram of another two-hand gesture interaction method according to an embodiment of the present disclosure;

[0104] FIG. 5 is a flow diagram of another two-hand gesture interaction method according to an embodiment of the present disclosure;

[0105] FIG. 6 is a flow diagram of another two-hand gesture interaction method according to an embodiment of the present disclosure;

[0106] FIG. 7 is a flow diagram of another two-hand gesture interaction method according to an embodiment of the present disclosure;

[0107] FIG. 8 is a structural diagram of a two-hand gesture interaction device according to an embodiment of the present disclosure;

[0108] FIG. 9 is a structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0109] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the following will combine the drawings in the embodiments of the present disclosure to make a clear and complete description of the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. The components of the embodiments of the present disclosure described and shown in the drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the claimed present disclosure, but only represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative labor are within the scope of the present disclosure.

[0110] FIG. 1 is a flow diagram of a two-hand gesture interaction method according to an embodiment of the present disclosure, as shown in FIG. 1, the method includes the following steps:

[0111] In step 101, after obtaining the current frame image of the real-time video of the interior of the vehicle cabin, according to the driving speed of the vehicle, hand information of each hand in a target area corresponding to the driving speed in the current frame image is obtained.

[0112] In step 102, according to the target hand key points of the hand information of each hand in the previous frame image of the current frame image and the target hand key points of the hand information of each hand in the current frame, a hand group belonging to the same hand in the current frame image and the previous frame image is determined.

[0113] In step 103, according to the target hand key points of each hand belonging to a first hand in the hand information in the current frame image and the target hand key points of each hand belonging to a second hand in the hand information in the current frame image, a left-right hand matching pair belonging to the same person in the current frame image is determined.

[0114] In step 104, for each left-right hand matching pair, according to a first hand group and a second hand group corresponding to the left-right hand matching pair in the current frame image and the previous frame image, a double-hand gesture corresponding to the left-right hand matching pair is determined.

[0115] In step 105, according to the double-hand gesture, a controllable component corresponding to the double-hand gesture in the vehicle is controlled to perform a corresponding operation.

[0116] Specifically, the vehicle cabin in the present disclosure refers to a part of the vehicle for carrying a user, and the vehicle cabin includes a main driver seat, a co-driver seat, a rear seat, and various control components for controlling the driving of the vehicle, such as a steering wheel, a brake, a clutch, and the like. In the vehicle cabin, there are also controllable components such as a vehicle-mounted screen, a sunroof, a vehicle-mounted player, and a cabin light, and the like. The controllable components on the vehicle also include a vehicle-mounted camera for shooting the external environment of the vehicle. The video of the external environment shot by the vehicle-mounted camera can be displayed on the vehicle-mounted screen.

[0117] In order to achieve the purpose of controlling the controllable components on the vehicle through the double-hand gesture, the inside real-time video of the vehicle cabin part can be acquired through the camera, when the current frame of the inside real-time video is acquired, the hand information of each hand in the target region corresponding to the current driving speed of the vehicle needs to be acquired according to the driving speed, since the current frame image includes the images of the three parts of the driver area, the co-driver area and the rear area in the vehicle cabin, the users in these three parts can control the controllable components in the vehicle through the double-hand gesture, in order to ensure the safe operation during driving, the hand information in the image of the corresponding region can be acquired according to the driving speed, and then only the user in the region can control the controllable components in the vehicle at the driving speed, and the users in other regions cannot control the gesture.

[0118] The driving speed and the target region can include the following cases:

[0119] 1. Only the hand information of the user in the driver area is recognized during the driving of the vehicle or when the driving speed is 0, so as to avoid the users in other regions and the user in the driver area from competing for the control right of the controllable components in the vehicle;

[0120] 2. Only the hand information of the users in the co-driver area and the rear area is recognized during the driving of the vehicle, so that the user in the driver area cannot realize the gesture control function during the driving, and then the user in the driver area focuses on driving the vehicle during the driving, so as to ensure the driving safety of the vehicle;

[0121] 3. Only the hand information of the user in the co-driver area is recognized during the driving of the vehicle;

[0122] 4. The left region and the right region in the driver area, the co-driver area and the rear area are divided into priorities, when the hand information in multiple regions is recognized in the current frame image during the driving of the vehicle or when the driving speed of the vehicle is 0, only the hand information corresponding to the region with the highest priority is recognized according to the priorities;

[0123] 5. The left region and the right region in the driver area, the co-driver area and the rear area are divided into priorities, when the hand information in multiple regions is recognized in the current frame image during the driving of the vehicle or when the driving speed of the vehicle is 0, the hand information of each region is sequentially recognized according to the priorities, and then the gesture is controlled according to the sequentially recognized gesture;

[0124] It should be noted that the relationship between the driving speed and the corresponding target region can be set according to actual needs, which will not be repeated here.

[0125] For any one hand, the hand is continuous in the time dimension, that is, the relative position or hand shape of the same hand at different times will change, in order to determine the continuity of the hand in time, it is necessary to determine the hand group in the current frame image and the last frame image which belongs to the same hand (that is, the same hand in the current frame image and the last frame image), because the hand information of the same hand is the same, therefore, in determining the hand group in the current frame image and the last frame image which belongs to the same hand, the target hand key points of each hand in the current frame image and the last frame image can be determined according to the target hand key points of each hand in the current frame image and the last frame image.

[0126] After determining the same hand in adjacent image frames, it is also necessary to determine the matching pair of left and right hands in the current frame image which belongs to the same person according to the target hand key points of each hand belonging to the first hand in the hand information in the current frame image and the target hand key points of each hand belonging to the second hand in the hand information in the current frame image, that is, to determine the double hands of the same person.

[0127] After determining the double hands of the same person in the current frame and the hand group belonging to the same hand in the previous and subsequent image frames, the action of the left hand of the person in the current frame image and the last frame image and the action of the right hand of the person in the current frame image and the last frame image can be determined, with the passage of time, the current image frame will become the last frame image at the current time in the next time, and then the double hand gesture of the person can be determined, after determining the double hand gesture, the corresponding controllable component can be controlled to perform the corresponding operation, for example, when the double hand gesture is used to control the volume, the vehicle-mounted player can be controlled to enlarge the volume.

[0128] In the above method, the control is completed by double hand gesture, which is less affected by the surrounding environment, so as to improve the anti-interference ability when interacting in the vehicle cabin environment, and because the control mode of double hand gesture can directly control the function on the vehicle, the interaction efficiency is improved.

[0129] It should be noted that for one hand, if the left and right hand matching pair of the hand does not exist in the current frame image and the last frame image, it proves that the corresponding gesture is not a double hand gesture, at this time, no control operation is performed.

[0130] In a feasible embodiment, the interior of the vehicle cabin part is provided with a camera, and the real-time video of the interior is obtained by the camera.

[0131] In a feasible embodiment, the interior of the vehicle cabin part is provided with a camera, and the real-time video of the interior is obtained through the camera, wherein the camera can be arranged at a rearview mirror in the vehicle cabin or at a position where a reading lamp of a front row in the vehicle cabin is located.

[0132] In a feasible embodiment, the interior of the vehicle cabin part is provided with a camera equal to the number of seats in the vehicle, and each camera is used to collect the real-time video of the interior of the corresponding seat area.

[0133] Specifically, the interior of the vehicle cabin part can be provided with a camera equal to the number of seats in the vehicle, and a camera is arranged in front of each seat, and the camera is used to collect the real-time video of the interior of the corresponding seat area, so that only one user's gesture is included in one real-time video, and the gesture recognition is relatively simple. When a corresponding camera is arranged in front of each seat, it can be determined according to the driving speed of the vehicle which camera to start and stop, and specific cases include the following:

[0134] 1. During the driving process of the vehicle or when the driving speed of the vehicle is 0, only the camera corresponding to the main driving area is started, so as to avoid the users in other areas and the user in the main driving area from competing for the control right of the controllable components in the vehicle;

[0135] 2. During the driving process of the vehicle, only the cameras corresponding to the co-driver area and the rear area are started, so that the user in the main driving area cannot realize gesture control function during the driving process, and the user in the main driving area focuses attention on driving the vehicle during the driving process, so as to ensure the driving safety of the vehicle;

[0136] 3. During the driving process of the vehicle, only the camera corresponding to the co-driver area is started;

[0137] 4. The cameras corresponding to the left and right areas in the main driving area, the co-driver area and the rear area are divided into priorities, and during the driving process of the vehicle or when the driving speed of the vehicle is 0, the cameras corresponding to all areas are started, and when hands exist in the videos collected by the cameras corresponding to multiple areas, the hand information in the video collected by the camera with the highest priority is recognized according to the priorities;

[0138] 5. The cameras corresponding to the left and right areas in the main driving area, the co-driver area and the rear area are divided into priorities, and during the driving process of the vehicle or when the driving speed of the vehicle is 0, the cameras corresponding to all areas are started, and when hands exist in the videos collected by the cameras corresponding to multiple areas, the hand information in the videos collected by the cameras is sequentially recognized according to the priorities, and then the gestures recognized according to the sequence are controlled.

[0139] In an implementation, in the step of acquiring hand information of each hand in the target region corresponding to the driving speed of the vehicle in the current frame image according to the driving speed of the vehicle in step 101, when the driving speed of the vehicle is greater than or equal to a preset speed, the hand information of each hand belonging to the co-driver region and the rear region in the current frame image is acquired; when the driving speed of the vehicle is less than the preset speed, the hand information of each hand in all seat regions in the current frame image is acquired.

[0140] It should be noted that the preset speed can be 0, when the driving speed of the vehicle is greater than 0, the hand information of each hand belonging to the co-driver region and the rear region in the current frame image is acquired; when the driving speed of the vehicle is equal to 0, the hand information of each hand in all seat regions in the current frame image is acquired. Of course, the actual value of the preset speed can also be set according to actual needs, which is not limited here.

[0141] In an implementation, in order to reduce the data processing amount, in the step of acquiring hand information of each hand in the target region corresponding to the driving speed of the vehicle in the current frame image according to the driving speed of the vehicle in step 101, the current frame image can be down-sampled according to a preset resolution to obtain a down-sampled image first; and then the hand information of each hand in the target region corresponding to the driving speed of the vehicle in the down-sampled image is acquired according to the driving speed of the vehicle.

[0142] In an implementation, FIG. 2 is a flowchart of another two-hand gesture interaction method provided by the embodiment of the present disclosure. In the step of acquiring hand information of each hand in the target region corresponding to the driving speed of the vehicle in the current frame image in step 101, the following steps can be implemented:

[0143] Step 201, for each hand in the target region in the current frame image, the hand image of the hand is acquired according to the diagonal coordinates of the hand-enclosing rectangle that can frame the hand.

[0144] Step 202, the hand image of the hand is input into a hand skeleton point detection model to obtain hand key points of the hand and confidence of each hand key point.

[0145] Step 203, the arithmetic mean of the confidence of each hand key point included in the hand is calculated.

[0146] Step 204, when the arithmetic mean value corresponding to the hand is greater than the preset confidence, inputting the hand image of the hand into a hand classifier to determine the chirality of the hand; wherein the hand information comprises: the hand key points of the hand and the chirality of the hand, and the chirality comprises left hand and right hand.

[0147] Specifically, when there are multiple hands in the target region of the current frame image, in order to obtain the hand information (the hand information comprises: the chirality and the hand key points, and the chirality comprises: left hand and right hand) of each hand (the hand of each hand), for the hand of each hand, a minimum circumscribed rectangle capable of framing the hand is used to frame the hand. After obtaining the minimum circumscribed rectangle, the horizontal and vertical coordinates of the four corners of the circumscribed rectangle in the current frame image can be determined. Then, the horizontal and vertical coordinates of the upper left corner and the lower right corner of the circumscribed rectangle can be taken as the diagonal coordinates of the circumscribed rectangle, or the horizontal and vertical coordinates of the upper right corner and the lower left corner can be taken as the diagonal coordinates of the circumscribed rectangle. After obtaining the diagonal coordinates, the hand image corresponding to each hand can be determined.

[0148] When determining the diagonal coordinates of the circumscribed rectangle, a coordinate system can be constructed with the center of the current frame image as the origin, or a coordinate system can be constructed with one of the four corners of the current frame image as the origin. Then, the diagonal coordinates of the circumscribed rectangle can be determined according to the pixel coordinates of the diagonal of the circumscribed rectangle in the coordinate system.

[0149] After obtaining the hand image of each hand, the hand image of each hand is input into a hand skeleton point detection model to obtain each hand key point included in the hand and the confidence of each hand key point. FIG. 3 is a schematic diagram of a hand key point provided by an embodiment of the present disclosure. As shown in FIG. 3, the hand key points of the hand of one hand include 21 points, which are marked as 0-20.

[0150] When determining each hand in the target region, it is determined through the hand shape, that is, as long as the object like the hand shape can be recognized as a hand, in order to determine whether each hand recognized is a real human hand, the arithmetic mean value of the confidence of each hand key point included in the hand can be calculated according to the confidence of each hand key point of the hand. For example: when the hand key points include 21 points, for any hand, 21 confidences can be obtained. The sum of the 21 confidences is calculated, and then the sum is divided by 21. The result obtained is taken as the arithmetic mean value of the hand. When the arithmetic mean value corresponding to the hand is greater than the preset confidence, it indicates that the hand is a real human hand. At this time, the hand image of the hand can be input into a hand classifier to determine the chirality of the hand. When the arithmetic mean value corresponding to the hand is less than the preset confidence, it indicates that the hand is not a real human hand. At this time, the hand can be removed.

[0151] It should be noted that the preset confidence can be 0.2, or can be set to other values according to actual needs, and is not specifically limited here. The hand skeleton point detection model needs to be trained before use. The RLE Loss (Residual Log-likelihood Estimation Loss) loss function is used in the model training, which is used for positioning of the hand skeleton points and the confidence of each hand skeleton point. The hand classifier also needs to be trained, and the CrossEntropyLoss (Cross Entropy Loss Function) loss function is used in the model training, which is used to determine the hand chirality, that is, to determine the left and right hands.

[0152] In a feasible implementation, FIG. 4 is a flow diagram of another two-hand gesture interaction method provided by an embodiment of the present disclosure. As shown in FIG. 4, when step 102 is performed, the following steps can be used to achieve it:

[0153] Step 401: Calculate the first Euclidean distance between the target hand key points of the hand information of each hand in the previous frame image and the target hand key points of the hand information of each hand in the current frame.

[0154] Step 402: According to the first Euclidean distance, a first Euclidean distance matrix is constructed.

[0155] Step 403: According to the first Euclidean distance matrix, the Hungarian algorithm is used to match each hand in the previous frame image and each hand in the current frame, and determine the hand group belonging to the same hand in the current frame image and the previous frame image.

[0156] Specifically, after obtaining the hand information of each hand, a first Euclidean distance between the target hand key point of the hand information of each hand in the previous frame image and the target hand key point of the hand information of each hand in the current frame is calculated. For example, the previous frame image includes four hands: hand 1, hand 2, hand 3 and hand 4, and the current frame image includes three hands: hand 5, hand 6 and hand 7. Then, the first Euclidean distance between the target hand key point of hand 1 and the target hand key point of hand 5, the first Euclidean distance between the target hand key point of hand 1 and the target hand key point of hand 6, the first Euclidean distance between the target hand key point of hand 1 and the target hand key point of hand 7, the first Euclidean distance between the target hand key point of hand 2 and the target hand key point of hand 5, the first Euclidean distance between the target hand key point of hand 2 and the target hand key point of hand 6, and the like are calculated. Twelve first Euclidean distances are obtained. Then, a first Euclidean distance matrix is constructed according to all the obtained first Euclidean distances. Then, the KM (Kuhn-Munkres, Hungarian algorithm) algorithm is used to match each hand in the previous frame image with each hand in the current frame, so as to determine the hand group in the current frame image and the previous frame image that belongs to the same hand, that is, the hand group composed of two hands belonging to the same hand in the current frame image and the previous frame image is determined. In the next moment, the current frame image becomes the previous frame image corresponding to the current frame image in the next moment. In this way, it can be determined which hand in each image frame in a sequence of image frames belongs to the same hand.

[0157] It should be noted that the target hand key point can be a wrist key point, such as the hand key point marked as 0 in FIG. 3. Of course, other hand key points can also be selected as the target hand key point, such as the hand key point marked as 1 or the hand key point marked as 3. The specific selection of the hand key point as the target hand key point can be set according to actual needs, which is not limited herein.

[0158] In a feasible implementation, FIG. 5 is a flow diagram of another two-hand gesture interaction method provided by the embodiment of the present disclosure. As shown in FIG. 5, when step 103 is performed, the following steps can be used to achieve the method:

[0159] Step 501: calculating a square value of a second Euclidean distance between the target hand key point of each hand belonging to a first hand in the hand information in the current frame image and the target hand key point of each hand belonging to a second hand in the hand information in the current frame image.

[0160] Step 502: constructing a second Euclidean distance matrix according to the square value of the second Euclidean distance.

[0161] Step 503, according to the second Euclidean distance matrix, using the Hungarian algorithm to match each hand belonging to the first chirality in the current frame image and each hand belonging to the second chirality in the current frame image, to determine the left and right hand matching pair belonging to the same person in the current frame image.

[0162] Specifically, when the first chirality is left hand and the second chirality is right hand, or when the first chirality is right hand and the second chirality is left hand, in order to determine the left and right hand matching pair belonging to the same person in the current frame image, when the current frame image includes 3 left hands: left hand 1, left hand 2 and left hand 3, and 3 right hands: right hand 1, right hand 2 and right hand 3, the second Euclidean distance between the target key points of any two hands can be calculated, for example, the square value of the second Euclidean distance between the target key points of left hand 1 and right hand 1, and the square value of the second Euclidean distance between the target key points of left hand 1 and right hand 2, and the square value of the second Euclidean distance between the target key points of left hand 1 and right hand 3, and the square value of the second Euclidean distance between the target key points of left hand 2 and right hand 1, and the square value of the second Euclidean distance between the target key points of left hand 2 and right hand 2, and so on, can be calculated. 9 square values of the second Euclidean distance can be obtained, then a second Euclidean distance matrix is constructed according to all the square values of the second Euclidean distance, and finally the left and right hands in the current frame image are matched by using the Hungarian algorithm according to the second Euclidean distance matrix, so as to determine the left and right hands belonging to the same person in the current frame image, i.e. the left and right hand matching pair belonging to the same person.

[0163] Over time, the left and right hand matching pair belonging to the same hand in each image frame included in a video can be obtained, and through steps 102 and 103, each hand in the current frame and other frames included in the internal real-time video can be determined, those hands belong to the same hand, and the left and right hand matching pair belonging to the same person in each frame image, so that the gesture action made by the left and right hands belonging to the same person over time can be determined according to the information in the two dimensions determined above, so that the corresponding two-hand gesture can be recognized.

[0164] In a feasible implementation, when the left and right hand matching pair belonging to the same person in the previous frame image is matched with the same hand in the hand group in the current frame image, the hand corresponding to the left and right hand matching pair belonging to the same person in the previous frame image in the current frame image is determined as the left and right hand matching pair belonging to the same person in the current frame image.

[0165] Specifically, for the left and right hand matching pairs that have appeared in the previous frame image, the left and right hand matching pairs belonging to the same person in the current frame image can be determined by inheritance, for example: when the left hand 1 and the right hand 1 in the previous frame image belong to the first left and right hand matching pair of the same person, and the left hand 1 and the left hand 2 in the current frame image form a hand group, and the right hand 1 and the right hand 2 in the current frame image form a hand group, then the left hand 2 and the right hand 2 in the current frame image are determined as the left and right hand matching pairs belonging to the same person in the current frame image. For a segment of internal real-time video, at different times, the current frame image and the previous frame image of the internal real-time video can be different, but the matching hands of the same person do not change, so through the above method, the left and right hand matching pairs belonging to the same person in the current frame image can be determined by inheritance of the left and right hand matching pairs belonging to the same person matched in the previous frame image, and only when the left and right hands belonging to the same person appear for the first time in the current frame image, the matching is performed according to the method shown in FIG. 5. Through the above method, the calculation amount can be reduced, and the recognition efficiency can be improved.

[0166] After determining the corresponding hand gestures of a certain user in the image frames corresponding to different times, the hand gestures corresponding to different image frames are combined in the order of time to obtain the hand gestures made by the user, thereby completing the corresponding gesture control.

[0167] In a feasible implementation, for each hand included in the current frame image, when the hand is matched successfully, the hand is marked using the identifier configured for the hand matched in the previous frame image; when the hand is not matched successfully, the hand is marked using a target identifier, which is different from the identifiers configured for the hands included in the previous frame image.

[0168] Specifically, in order to quickly determine the left and right hand matching pairs belonging to the same person in the current frame image, different identifiers need to be assigned to the hands included in the previous frame image. After determining which two hands in the previous frame image belong to the left and right hand matching pairs of the same person, a plurality of identifier pairs can be determined, and after determining which two hands in the current frame image and the previous frame image belong to the hand group of the same hand, the same identifier can be used to mark the same hand, so that the identifier pair corresponding to the previous frame image can be used to determine the identifier pair corresponding to the current frame image, thereby quickly completing the matching of the left and right hand matching pairs in the current frame image.

[0169] The hand matching is successful, indicating that the hand exists in the previous frame image, and thus the hand can be marked by inheritance. The hand matching is not successful, indicating that the hand appears for the first time in the current frame image, and the hand does not exist in the previous frame image. At this time, the hand needs to be marked using an identifier that has not been used.

[0170] In a feasible implementation, FIG. 6 is a flowchart of another two-hand gesture interaction method provided by an embodiment of the present disclosure. As shown in FIG. 6, when performing the step of determining, in the current frame image, left and right hand matching pairs belonging to the same person as the left and right hand matching pairs belonging to the same person in the previous frame image, the hand group in which the left and right hand matching pairs belong to the same person in the previous frame image matches the same hand in the current frame image, the left and right hand matching pairs belonging to the same person in the current frame image can be determined by the following steps:

[0171] Step 601: According to the identifiers included in the left and right hand matching pairs belonging to the same person in the previous frame image, whether the identifiers included in the left and right hand matching pairs belonging to the same person in the previous frame image exist in the identifiers configured for each hand in the current frame image is found.

[0172] Step 602: If the identifiers exist, the hands corresponding to the identifiers identical to the identifiers included in the left and right hand matching pairs belonging to the same person in the previous frame image in the current frame image are determined as the left and right hand matching pairs belonging to the same person in the current frame image.

[0173] Specifically, after the hand group belonging to the same hand in the current frame image and the previous frame image is determined, each hand in the current frame image can inherit the identifiers configured for each hand in the previous frame image. When the identifier group corresponding to the left and right hand matching pairs exists in the previous frame image, the hands corresponding to the identifier group identical to the identifier group in the current frame image are determined as the left and right hand matching pairs belonging to the same person. Through the above method, the left and right hand matching pairs belonging to the same person in the current frame image can be determined by inheritance, which is beneficial to reducing the calculation amount.

[0174] In a feasible implementation, FIG. 7 is a flowchart of another two-hand gesture interaction method provided by an embodiment of the present disclosure. As shown in FIG. 7, when performing step 104, the following steps can be implemented:

[0175] Step 701: The sum of the first movement distance of the target hand key point of each hand included in the first hand group and the second movement distance of the target hand key point of each hand included in the second hand group is calculated.

[0176] Step 702: The ratio of the sum to the sampling interval length of the internal real-time video is calculated, so as to take the ratio as the moving speed of the two-hand gesture corresponding to the left and right hand matching pair.

[0177] Step 703, when the moving speed is less than or equal to the preset moving speed, input the current frame image into the double-hand static gesture classification model to determine the corresponding double-hand static gesture of the left and right hand matching pair.

[0178] Step 704, when the moving speed is greater than the preset moving speed, input the preset number of image frames containing the left and right hand matching pair into the double-hand dynamic gesture classification model to determine the corresponding double-hand dynamic gesture of the left and right hand matching pair.

[0179] Specifically, in order to determine whether the double-hand gesture is a double-hand static gesture or a double-hand dynamic gesture, the moving speed of the same person's double hands in the front and rear frames can be determined. When the moving speed is less than or equal to the preset moving speed, it indicates that the double-hand gesture is a double-hand static gesture, and then the image corresponding to the double-hand gesture in the current frame image is input into the double-hand static gesture classification model, so as to identify the specific double-hand static gesture and further complete the control corresponding to the double-hand static gesture. When the moving speed is greater than the preset speed, it indicates that the double-hand gesture is a double-hand dynamic gesture, and then the image corresponding to the double-hand gesture in the preset number of image frames is input into the double-hand dynamic gesture classification model, so as to identify the specific double-hand dynamic gesture and further complete the control corresponding to the double-hand dynamic gesture, for example, input the image corresponding to the double-hand gesture in the 10 frames of images corresponding to the t n-9 -t n instant into the double-hand dynamic gesture classification model, wherein t n is the current time. Moreover, in the determination of the same person's double hands, the identification in the above content can be used for determination.

[0180] It should be noted that the specific value of the preset moving speed can be set according to actual needs, which is not limited here. The specific double-hand static gesture classification model and the double-hand dynamic gesture classification model can be selected according to actual needs, which is also not limited here.

[0181] In a feasible embodiment, the controllable components include: a vehicle-mounted screen, a sunroof, a vehicle-mounted player, a vehicle-mounted camera, and a vehicle cabin light.

[0182] Specifically, the double-hand gesture can adjust the size of the resolution of the vehicle-mounted screen, control the game steering wheel displayed on the vehicle-mounted screen, control the opening and closing of the sunroof, control the volume of the vehicle-mounted player, collect the content currently played by the vehicle-mounted player, control the opening and closing of the vehicle-mounted camera, control the shooting time of the vehicle-mounted camera, control the opening and closing of the vehicle cabin light, and cancel the gesture instruction.

[0183] It should be noted that the specific gesture for controlling the specific function can be set according to actual needs, and is not limited here.

[0184] FIG. 8 is a structural schematic diagram of a two-hand gesture interaction device provided by an embodiment of the present disclosure, as shown in FIG. 8, the device includes:

[0185] The acquisition unit 81, after acquiring the current frame image of the internal real-time video of the vehicle cabin part of the vehicle, acquires hand information of each hand in a target region corresponding to the driving speed in the current frame image according to the driving speed of the vehicle;

[0186] The first determination unit 82 is configured to determine a hand group belonging to the same hand in the current frame image and the previous frame image according to target hand key points of the hand information of each hand in the previous frame image of the current frame image and target hand key points of the hand information of each hand in the current frame.

[0187] The second determination unit 83 is configured to determine a left-right hand matching pair in the current frame image according to the target hand key points of each hand belonging to a first hand in the hand information in the current frame image and the target hand key points of each hand belonging to a second hand in the hand information in the current frame image.

[0188] The gesture recognition unit 84 is configured to determine a two-hand gesture corresponding to each left-right hand matching pair according to a first hand group and a second hand group corresponding to the left-right hand matching pair in the current frame image and the previous frame image.

[0189] The control unit 85 is configured to control a controllable component in the vehicle corresponding to the two-hand gesture to perform a corresponding operation according to the two-hand gesture.

[0190] In a feasible implementation, the interior of the vehicle cabin part is provided with a camera device, and the internal real-time video is acquired by the camera device.

[0191] In a feasible implementation, the interior of the vehicle cabin part is provided with a camera device close to the top of the head of the vehicle.

[0192] In a feasible implementation, the interior of the vehicle cabin part is provided with a camera device equal to the number of seats in the vehicle, and each camera device is used to collect the internal real-time video of the corresponding seat area.

[0193] In a feasible implementation, when the acquisition unit acquires the hand information of each hand in the target region corresponding to the driving speed in the current frame image according to the driving speed of the vehicle, it includes:

[0194] When the driving speed of the vehicle is greater than or equal to a preset speed, hand information of each hand in a target region corresponding to the driving speed in the current frame image is acquired according to the driving speed of the vehicle;

[0195] When the driving speed of the vehicle is less than the preset speed, hand information of each hand in all seat regions in the current frame image is acquired.

[0196] In a feasible implementation, the acquisition unit, when acquiring hand information of each hand in a target region corresponding to the driving speed in the current frame image according to the driving speed of the vehicle, comprises:

[0197] The current frame image is down-sampled according to a preset resolution to obtain a down-sampled image;

[0198] Hand information of each hand in a target region corresponding to the driving speed in the down-sampled image is acquired according to the driving speed of the vehicle.

[0199] In a feasible implementation, the acquisition unit, when acquiring hand information of each hand in a target region corresponding to the driving speed in the current frame image, comprises:

[0200] For each hand in a target region in the current frame image, a hand image of the hand is acquired according to diagonal coordinates of a hand circumscribed rectangle capable of framing the hand;

[0201] The hand image of the hand is input into a hand skeleton point detection model to obtain hand key points of the hand and confidence of each hand key point;

[0202] An arithmetic mean of the confidence of each hand key point included in the hand is calculated;

[0203] When the arithmetic mean corresponding to the hand is greater than a preset confidence, the hand image of the hand is input into a hand classifier to determine a handiness of the hand;

[0204] The hand information comprises: hand key points of the hand and handiness of the hand, and the handiness comprises left hand and right hand.

[0205] In a feasible implementation, the first determination unit, when determining a hand group belonging to a same hand in the current frame image and a previous frame image of the current frame image according to target hand key points of hand information of each hand in the previous frame image and target hand key points of hand information of each hand in the current frame, comprises:

[0206] calculate a first Euclidean distance between the target hand key points of the hand information of each hand in the previous frame image and the target hand key points of the hand information of each hand in the current frame;

[0207] construct a first Euclidean distance matrix according to the first Euclidean distances;

[0208] match each hand in the previous frame image and each hand in the current frame by using a Hungarian algorithm according to the first Euclidean distance matrix, and determine a hand group in the current frame image and the previous frame image that belongs to a same hand.

[0209] In a feasible implementation, the second determining unit is configured to, when determining the left-right hand matching pair in the current frame image that belongs to a same person according to the target hand key points of each hand in the hand information in the current frame image that belongs to a first hand property and the target hand key points of each hand in the hand information in the current frame image that belongs to a second hand property, include:

[0210] calculate a square value of a second Euclidean distance between the target hand key points of each hand in the hand information in the current frame image that belongs to the first hand property and the target hand key points of each hand in the hand information in the current frame image that belongs to the second hand property;

[0211] construct a second Euclidean distance matrix according to the square value of the second Euclidean distance;

[0212] match each hand in the current frame image that belongs to the first hand property and each hand in the current frame image that belongs to the second hand property by using a Hungarian algorithm according to the second Euclidean distance matrix, and determine the left-right hand matching pair in the current frame image that belongs to a same person.

[0213] In a feasible implementation, the second determining unit is further configured to:

[0214] when the left-right hand matching pair in the previous frame image that belongs to a same person is matched with a same hand group in the current frame image, determine the hand in the current frame image corresponding to the left-right hand matching pair in the previous frame image that belongs to a same person as the left-right hand matching pair in the current frame image that belongs to a same person.

[0215] In a feasible implementation, the device further includes:

[0216] The marking unit is configured to, for each hand included in the current frame image, when the hand is matched successfully, marking the hand with an identifier configured for the hand matched in the previous frame image; and when the hand is not matched successfully, marking the hand with a target identifier, which is different from the identifiers configured for each hand included in the previous frame image.

[0217] In an implementation, the second determining unit is configured to, when the left and right hand matching pairs belonging to the same person in the previous frame image are matched with the same hand group in the current frame image, determine, as the left and right hand matching pairs belonging to the same person in the current frame image, the hands corresponding to the identifiers included in the left and right hand matching pairs belonging to the same person in the previous frame image.

[0218] According to the identifiers included in the left and right hand matching pairs belonging to the same person in the previous frame image, searching for, in the identifiers configured for each hand included in the current frame image, whether there are the identifiers included in the left and right hand matching pairs belonging to the same person in the previous frame image.

[0219] If there are, determining, as the left and right hand matching pairs belonging to the same person in the current frame image, the hands corresponding to the identifiers identical to the identifiers included in the left and right hand matching pairs belonging to the same person in the previous frame image.

[0220] In an implementation, the second determining unit is configured to, for each left and right hand matching pair, according to the first hand group and the second hand group corresponding to the left and right hand matching pair in the current frame image and the previous frame image, determine the double-hand gesture corresponding to the left and right hand matching pair, including:

[0221] Calculating the sum of the first moving distance of the target hand key point of each hand included in the first hand group and the second moving distance of the target hand key point of each hand included in the second hand group;

[0222] Calculating the ratio of the sum to the sampling interval length of the internal real-time video, and taking the ratio as the moving speed of the double-hand gesture corresponding to the left and right hand matching pair;

[0223] When the moving speed is less than or equal to a preset moving speed, inputting the current frame image into a double-hand static gesture classification model to determine the double-hand static gesture corresponding to the left and right hand matching pair;

[0224] When the moving speed is greater than the preset moving speed, inputting a preset number of image frames containing the left and right hand matching pair into a double-hand dynamic gesture classification model to determine the double-hand dynamic gesture corresponding to the left and right hand matching pair.

[0225] In one possible implementation, the controllable component includes:

[0226] Vehicle screen, sunroof, vehicle player, vehicle camera, vehicle cabin light.

[0227] For the related explanations of the two-hand gesture interaction device, please refer to the related descriptions of the two-hand gesture interaction method, which will not be described in detail here.

[0228] FIG. 9 is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure, which includes a processor 901, a storage medium 902, and a bus 903. The storage medium 902 stores machine-readable instructions executable by the processor 901. When the electronic device runs a two-hand gesture interaction method as in an embodiment, the processor 901 and the storage medium 902 communicate through the bus 903. The processor 901 executes the machine-readable instructions to perform steps as in an embodiment.

[0229] In an embodiment, the storage medium 902 can also execute other machine-readable instructions to perform other methods as in an embodiment. For the specific method steps and principles, please refer to the descriptions of the embodiments, which will not be described in detail here.

[0230] Embodiment four of the present disclosure also provides a computer-readable storage medium, which stores a computer program. When the computer program is run by a processor, the steps shown in the above embodiments are performed.

[0231] In the embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There can be another division during actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0232] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units. That is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.

[0233] In addition, each function unit in the embodiments provided by the present disclosure can be integrated into one processing unit, or each unit can exist physically, or two or more units can be integrated into one unit.

[0234] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0235] It should be noted that: similar reference numbers and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings, in addition, the terms "first", "second", "third" and the like are only used to distinguish description, and cannot be understood as indicating or implying relative importance.

[0236] Finally, it should be noted that: the above-described embodiments are only specific implementations of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, and are not limited thereto, the protection scope of the present disclosure is not limited thereto, although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any skilled person in the art can modify or easily think of changes to the technical solutions described in the foregoing embodiments within the technical range disclosed by the present disclosure, or make equivalent replacement to some technical features; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure. All should be covered in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims. Industrial applicability

[0237] The present disclosure provides a double-hand gesture interaction method and device, electronic equipment and storage medium, which realizes control of controllable components on a vehicle, guarantees safe operation in the driving process, and makes gesture recognition relatively simple by only collecting real-time video by the camera device inside the vehicle cabin, thereby reducing the data processing amount. The double-hand gesture interaction method provided by the present disclosure completes control through double-hand gestures, is less affected by the surrounding environment, and is conducive to improving the anti-interference capability when interacting in the vehicle cabin environment. The control method can directly control the functions on the vehicle, which is conducive to improving the interaction efficiency. The control method pairs the left and right hands of the same person through inheritance, which is conducive to reducing the calculation amount and improving the recognition efficiency.

Claims

1. A two-handed gesture interaction method, characterized by, The method comprises: After obtaining a current frame image of internal real-time video of a vehicle cabin part, hand information of each hand in a target region corresponding to a driving speed of the vehicle in the current frame image is obtained according to the driving speed of the vehicle; Target hand key points of each hand in the current frame image and target hand key points of each hand in a previous frame image of the current frame image are determined to determine a hand group belonging to the same hand in the current frame image and the previous frame image; The target hand key points of each hand belonging to a first hand in the hand information in the current frame image and the target hand key points of each hand belonging to a second hand in the hand information in the current frame image are determined to determine a left-right hand matching pair belonging to the same person in the current frame image; For each left-right hand matching pair, a first hand group and a second hand group corresponding to the left-right hand matching pair in the current frame image and the previous frame image are determined to determine a double-hand gesture corresponding to the left-right hand matching pair; According to the double-hand gesture, a controllable component corresponding to the double-hand gesture in the vehicle is controlled to perform a corresponding operation.

2. The method of claim 1, wherein, The interior of the vehicle cabin part is provided with a camera device, and the internal real-time video is obtained through the camera device.

3. The method of claim 2, wherein, The interior of the vehicle cabin part is provided with one camera device near the top of the head of the vehicle.

4. The method of claim 2, wherein, The interior of the vehicle cabin part is provided with camera devices equal to the number of seats in the vehicle, and each camera device is used to collect the internal real-time video of the corresponding seat area.

5. The method of claim 1, wherein, The method comprises: When the driving speed of the vehicle is greater than or equal to a preset speed, hand information of each hand belonging to a co-pilot area and a rear area in the current frame image is obtained; When the driving speed of the vehicle is less than the preset speed, hand information of each hand in all seat areas in the current frame image is obtained.

6. The method of claim 1, wherein, The method comprises: The current frame image is down-sampled according to a preset resolution to obtain a down-sampled image; According to the driving speed of the vehicle, hand information of each hand in a target region corresponding to the driving speed in the down-sampled image is obtained.

7. The method of claim 1, wherein, The method comprises: For each hand in the target region in the current frame image, a hand image of the hand is obtained according to diagonal coordinates of a hand circumscribed rectangle of the hand; The hand image of the hand is input into a hand skeleton point detection model to obtain hand key points of the hand and confidence of each hand key point; An arithmetic mean value of the confidence of each hand key point included in the hand is calculated; When the arithmetic mean value corresponding to the hand is greater than a preset confidence, the hand image of the hand is input into a hand classifier to determine the hand of the hand. The hand information includes hand key points of the hand and a hand type of the hand, and the hand type includes a left hand and a right hand.

8. The method of claim 1, wherein, The method further includes: The method further includes: The method further includes: The method further includes:

9. The method of claim 1, wherein, The method further includes: The method further includes: The method further includes: The method further includes:

10. The method of claim 9, wherein, The method further includes: The method further includes:

11. The method of claim 8, wherein, The method further includes: The method further includes: The method further includes:

12. The method of claim 11, wherein, The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method further includes: The method According to the label included in the same left-right hand matching pair in the previous frame image, whether there is the label included in the same left-right hand matching pair in the previous frame image in the labels configured for each hand included in the current frame image is found; If there is, the hand corresponding to the same label pair of the current frame image and the label included in the same left-right hand matching pair in the previous frame image is determined as the left-right hand matching pair belonging to the same person in the current frame image.

13. The method of claim 1, wherein, According to the first hand group and the second hand group corresponding to each left-right hand matching pair in the current frame image and the previous frame image, the double-hand gesture corresponding to the left-right hand matching pair is determined, including: The sum of the first moving distance of the target hand key point of each hand included in the first hand group and the second moving distance of the target hand key point of each hand included in the second hand group is calculated; The ratio of the sum to the sampling interval length of the internal real-time video is calculated, and the ratio is taken as the moving speed of the double-hand gesture corresponding to the left-right hand matching pair; When the moving speed is less than or equal to a preset moving speed, the current frame image is input into a double-hand static gesture classification model to determine the double-hand static gesture corresponding to the left-right hand matching pair; When the moving speed is greater than the preset moving speed, a preset number of image frames containing the left-right hand matching pair are input into a double-hand dynamic gesture classification model to determine the double-hand dynamic gesture corresponding to the left-right hand matching pair.

14. The method of claim 1, wherein, The controllable components include: Vehicle-mounted screen, sunroof, vehicle-mounted player, vehicle-mounted camera, vehicle cabin light.

15. A two-handed gesture interaction device, characterized by The device includes: An acquisition unit, after acquiring a current frame image of an internal real-time video of a vehicle cabin part of a vehicle, according to a driving speed of the vehicle, acquires hand information of each hand in a target area corresponding to the driving speed in the current frame image; A first determination unit configured to determine a hand group belonging to the same hand in the current frame image and the previous frame of the current frame image according to target hand key points of hand information of each hand in the previous frame of the current frame image and target hand key points of hand information of each hand in the current frame; A second determination unit configured to determine a left-right hand matching pair belonging to the same person in the current frame image according to the target hand key points of each hand belonging to a first hand in the hand information in the current frame image and the target hand key points of each hand belonging to a second hand in the hand information in the current frame image; A gesture recognition unit configured to determine a double-hand gesture corresponding to each left-right hand matching pair according to a first hand group and a second hand group corresponding to the left-right hand matching pair in the current frame image and the previous frame image; A control unit configured to control a controllable component corresponding to the double-hand gesture in the vehicle to perform a corresponding operation according to the double-hand gesture.

16. An electronic device, comprising: A processor and a memory, the memory stores machine executable instructions executable by the processor, and the processor executes the machine executable instructions to implement the double-hand gesture interaction method in any one of claims 1-14.

Citation Information

Patent Citations

  • Somatosensory-based natural interaction method for virtual mine

    CN104750397A

  • Automobile gesture interaction method

    CN113076836A

  • Hand tracking method, device, equipment and computer storage medium

    CN114067426A

  • Gesture recognition method and system based on two-dimensional image and electronic equipment

    CN114510142A

  • Two-hand gesture interaction method and device, electronic equipment and storage medium

    CN118567488A