Gesture data labeling method and device, equipment and storage medium
Automatically label gesture data through dynamic gesture recognition model, solving the problem of inefficient manual marking in the existing technology, achieving efficient and automatic gesture data labeling, and ensuring data quality.
Patent Information
- Application Number
- CN202311489522.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-09
- Publication Date
- 2025-05-13
AI Technical Summary
Existing gesture data annotation methods rely on manual annotation, resulting in inefficiency and it is difficult to ensure efficient annotation of large-scale data sets.
By acquiring multiple image groups, each image group contains t-frame images arranged in chronological order, and using dynamic gesture recognition model to identify the key point information of the hand, automatically determine the labeling result of the target gesture recognition result.
On the premise of ensuring data quality, the efficiency of gesture data labeling is significantly improved, the dependence of manual labeling is reduced, and the automation level of data labeling is improved.
Smart Images

Figure CN119992245A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data annotation technology, and in particular, relates to a gesture data annotation method, device, equipment and storage medium. Background Art
[0002] With the rapid development of computer vision technology, controlling smart devices through gestures has become a common way to control smart devices. Specifically, the user's gestures are recognized through gesture recognition technology, and the control instructions corresponding to the recognized gestures are executed to control smart devices. At present, gesture recognition technology is usually implemented based on deep neural networks. In order for deep neural networks to be able to perform gesture recognition, it is necessary to annotate gesture data in advance, and then use the annotated gesture data to train the deep neural network model. Among them, the quality of the annotated data directly affects the training results of the deep neural network. Therefore, it is crucial to ensure data quality.
[0003] At present, the annotation of gesture data is usually done manually. Specifically, the user usually first extracts an image sequence frame containing a complete gesture action from the gesture video, then manually annotates the coordinates of multiple hand key points in each frame, and finally determines the coordinates of multiple key points corresponding to the image sequence frame as the annotation result of the gesture action. Manual annotation can usually ensure the quality of the annotated data. However, manual annotation is time-consuming and labor-intensive, resulting in low annotation efficiency. Summary of the invention
[0004] The embodiments of the present application provide a method, apparatus, device and storage medium for annotating gesture data, which can improve the annotation efficiency while ensuring the data quality.
[0005] In a first aspect, an embodiment of the present application provides a method for labeling gesture data, the method comprising:
[0006] Acquire multiple image groups, each of the image groups includes t frames of first images arranged in time sequence, the first images in the multiple image groups all include the same palm, and t is a positive integer;
[0007] Obtaining first hand key point information of each of the image groups;
[0008] For each of the image groups, using a dynamic gesture recognition model to perform dynamic gesture recognition on the first hand key point information to obtain a gesture recognition result and score for each of the image groups;
[0009] Determine a plurality of target image groups and scores corresponding to target gesture recognition results, wherein the plurality of target image groups are continuous and the target gesture recognition result is any one of the gesture recognition results of the plurality of target image groups;
[0010] The first hand key point information of the target image group corresponding to the highest score among the multiple scores is determined as the labeling result of the target gesture recognition result.
[0011] In a possible implementation manner, acquiring multiple image groups includes:
[0012] Acquire multiple frames of the first image that are arranged in time sequence and include the same palm;
[0013] The sliding window is moved according to a preset step size to divide multiple frames of the first image into multiple image groups, and the size of the sliding window is the size of t frames of the first image.
[0014] In a possible implementation, the acquiring of multiple frames of the first images, which are arranged in time sequence and include the same palm, includes:
[0015] Acquire a plurality of consecutive frames of second images arranged in time sequence, where the second images are images in the gesture video;
[0016] For each frame of the second image, perform palm detection on the second image using a hand detection model to obtain position information of the palm in the second image;
[0017] According to step 1, the position information of the palm in two adjacent frames of the second image is matched in sequence to obtain a matching result corresponding to the first target image, where the first target image is the second image in the latter frame of the two adjacent frames of the second image;
[0018] The second images with successful matching results among the multiple frames of second images are determined as the multiple frames of first images arranged in chronological order.
[0019] In a possible implementation, determining the second images with successful matching results in the multiple frames of second images as the multiple frames of first images arranged in chronological order includes:
[0020] When the matching result is successful, the first target image is added to the tracking sequence where the second target image is located, and the second target image is the second image of the previous frame in two adjacent frames of the second image;
[0021] The multiple frames of images in the tracking sequence are determined as multiple frames of the first images arranged in time sequence.
[0022] In a possible implementation manner, before adding the first target image to the tracking sequence where the second target image is located, the method further includes:
[0023] Determine the first second image in the plurality of second images as the third target image, or, in the process of sequentially matching two adjacent second images, determine the next second image whose matching result with the previous second image is a failed match as the third target image;
[0024] Establishing a tracking sequence corresponding to each frame of the third target image, wherein each frame of the third target image is the first frame of the tracking sequence corresponding to it, and the remaining images in the tracking sequence corresponding to the j-th frame of the third target image include the second image of the next frame of the second image in the two adjacent frames after the j-th frame of the third target image that successfully matches the second image of the previous frame, and j is a positive integer;
[0025] In the case that the palm in the third target image is the same as the palm in the second target image, the tracking sequence corresponding to the third target image is determined as the tracking sequence where the second target image is located.
[0026] In a possible implementation manner, the obtaining the first hand key point information of each of the image groups includes:
[0027] For each frame of the first image in each of the image groups, perform hand key point detection on the first image using a key point detection model to obtain second hand key point information of each frame of the first image;
[0028] The t pieces of the second hand key point information arranged in time sequence are determined as the first hand key point information.
[0029] In a possible implementation, the using a key point detection model to perform hand key point detection on the first image to obtain second hand key point information of each frame of the first image includes:
[0030] Performing hand key point detection and palm recognition on the first image using a key point detection model to obtain the second hand key point information and a palm recognition result, wherein the palm recognition result indicates whether an area corresponding to the second hand key point information is a hand area;
[0031] The step of using a dynamic gesture recognition model to perform dynamic gesture recognition on the first hand key point information includes:
[0032] When the area corresponding to the second hand key point information is a hand area, the dynamic gesture recognition model is used to perform dynamic gesture recognition on the first hand key point information.
[0033] In a possible implementation, each of the second hand key point information includes first coordinates of i hand key points, where i is a positive integer; and performing dynamic gesture recognition on the first hand key point information by using a dynamic gesture recognition model includes:
[0034] Subtracting the first coordinates of each of the image groups from the first coordinates of the target hand key point of the third target image to obtain t×i second coordinates, wherein the third target image is any one of the t frames of the first images, and the target hand key point is any one of the i hand key points;
[0035] The dynamic gesture recognition model is used to perform dynamic gesture recognition on the t×i second coordinates.
[0036] In a possible implementation manner, determining a plurality of target image groups corresponding to the target gesture recognition results includes:
[0037] When a gesture recognition result of the first image group is the same as a gesture recognition result of the second image group and different from a gesture recognition result of the third image group, determining the gesture recognition result of the first image group as a target gesture recognition result, wherein the second image group is an image group previous to the first image group and the third image group is an image group subsequent to the first image group;
[0038] The gesture recognition result is a target gesture recognition result, and a plurality of continuous image groups located before the first image group are determined as the plurality of target image groups.
[0039] In a second aspect, an embodiment of the present application provides a gesture data annotation device, the device comprising:
[0040] A first acquisition module is used to acquire a plurality of image groups, each of which includes t frames of first images arranged in time sequence, the first images in the plurality of image groups all include the same palm, and t is a positive integer;
[0041] A second acquisition module, used for acquiring first hand key point information of each of the image groups;
[0042] a recognition module, configured to perform dynamic gesture recognition on the first hand key point information for each of the image groups using a dynamic gesture recognition model, to obtain a gesture recognition result and score for each of the image groups;
[0043] A first determination module is used to determine a plurality of target image groups and scores corresponding to a target gesture recognition result, wherein the plurality of target image groups are continuous and the target gesture recognition result is any one of the gesture recognition results of the plurality of target image groups;
[0044] The second determination module is used to determine the first hand key point information of the target image group corresponding to the highest score among the multiple scores as the labeling result of the target gesture recognition result.
[0045] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising: a processor and a memory storing computer program instructions;
[0046] When the processor executes the computer program instructions, the processor implements any possible implementation method of the first aspect described above.
[0047] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implements a method in any possible implementation method of the first aspect described above.
[0048] The gesture data annotation method, device, equipment and storage medium of the embodiment of the present application, by obtaining the first hand key point information corresponding to each of the multiple image groups, and using the dynamic gesture recognition model to perform dynamic gesture recognition on the first hand key point information, obtain the gesture recognition result of each image group, and can automatically determine the correspondence between the first hand key point information and the gesture recognition result. By selecting a group of hand key point information corresponding to the target gesture recognition result from multiple groups of hand key point information, instead of manually annotating the hand key point information after determining the target gesture action, the annotation efficiency can be improved. By determining multiple target image groups and scores corresponding to the target gesture recognition result, and determining the first hand key point information of the target image group corresponding to the highest score among the multiple scores as the annotation result of the target gesture recognition result, the key point information that best matches the target gesture recognition result can be selected from multiple groups of key point information, thereby ensuring data quality. In this way, through the embodiment of the present application, the annotation efficiency can be improved while ensuring data quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0050] Figure 1 It is a flowchart of a method for labeling gesture data provided by an embodiment of the present application;
[0051] Figure 2 is a flowchart of another method for labeling gesture data provided in an embodiment of the present application;
[0052] Figure 3 is a structural schematic diagram of a gesture data annotation device provided in an embodiment of the present application;
[0053] Figure 4 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the scheme of the present application will be further described below. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0055] In the following description, many specific details are set forth to facilitate a full understanding of the present application, but the present application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only part of the embodiments of the present application, rather than all of the embodiments.
[0056] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0057] As described in the background technology section, in order to solve the problems of the prior art, the embodiments of the present application provide a method, apparatus, device and storage medium for annotating gesture data.
[0058] The following first introduces the gesture data labeling method provided in the embodiment of the present application.
[0059] Figure 1 FIG. 1 is a flow chart showing a method for labeling gesture data provided by an embodiment of the present application. Figure 1 As shown, the gesture data annotation method provided in the embodiment of the present application includes the following steps:
[0060] S110, acquiring multiple image groups, each image group including t frames of first images arranged in time sequence, the first images in the multiple image groups all include the same palm, and t is a positive integer;
[0061] S120, obtaining first-hand key point information of each image group;
[0062] S130, for each image group, using a dynamic gesture recognition model to perform dynamic gesture recognition on the first hand key point information to obtain a gesture recognition result and score for each image group;
[0063] S140, determining a plurality of target image groups and scores corresponding to a target gesture recognition result, wherein the plurality of target image groups are continuous, and the target gesture recognition result is any one of the gesture recognition results of the plurality of target image groups;
[0064] S150: Determine the first hand key point information of the target image group corresponding to the highest score among the multiple scores as the labeling result of the target gesture recognition result.
[0065] The method for labeling gesture data of the embodiment of the present application obtains the first hand key point information corresponding to each of the multiple image groups, and uses the dynamic gesture recognition model to perform dynamic gesture recognition on the first hand key point information to obtain the gesture recognition result of each image group, and can automatically determine the correspondence between the first hand key point information and the gesture recognition result. By selecting a group of hand key point information corresponding to the target gesture recognition result from multiple groups of hand key point information, rather than manually labeling the hand key point information after determining the target gesture action, the labeling efficiency can be improved. By determining multiple target image groups and scores corresponding to the target gesture recognition result, and determining the first hand key point information of the target image group corresponding to the highest score among the multiple scores as the labeling result of the target gesture recognition result, the key point information that best matches the target gesture recognition result can be selected from multiple groups of key point information, thereby ensuring data quality. In this way, through the embodiment of the present application, the labeling efficiency can be improved while ensuring data quality.
[0066] The specific implementation methods of the above steps are introduced below.
[0067] In some embodiments, in S110, the image group can be used for dynamic gesture recognition. That is, by analyzing multiple frames of first images in the image group, the gesture action corresponding to the image group can be determined. Therefore, on the one hand, the t frames of first images in each image group can be arranged in chronological order. For example, t can be 10, 11, 12, etc. On the other hand, the t frames of first images in the image group can include the same palm. In addition, multiple image groups can be used to jointly annotate the target gesture action, so the first images in multiple image groups can all include the same palm. Among them, the target gesture action can be any one of the gesture actions corresponding to each of the multiple image groups.
[0068] In addition, the gesture actions corresponding to the multiple image groups may be the same or different, which is not limited here. When the gesture actions corresponding to the multiple image groups are the same, the multiple image groups may be arranged in chronological order or not. When the gesture actions corresponding to the multiple image groups are different, the multiple image groups may be arranged in chronological order. Among them, the arrangement of the multiple image groups in chronological order may be the arrangement of the first frames of the multiple image groups in chronological order.
[0069] If multiple image groups are arranged in chronological order, and the first image of frame t in each image group is also arranged in chronological order, the orderliness and coherence of multiple gestures recognized subsequently can be guaranteed, thereby improving the accuracy of gesture annotation. Therefore, in order to improve the accuracy of gesture annotation, in some embodiments, the above S110 may specifically include:
[0070] Acquire a plurality of first image frames that are arranged in time sequence and include the same palm;
[0071] The sliding window is moved according to a preset step size to divide the multiple frames of first images into multiple image groups, and the size of the sliding window is the size of the t frames of first images.
[0072] Here, the multiple frames of first images may be multiple frames of images including the same palm. In addition, the preset step size may be a moving step size of a preset sliding window. The sliding window may be a window of a fixed size. By moving the sliding window, multiple frames of first images may be traversed. In addition, the preset step size may be, for example, 1, 2, 3, etc. If the preset step size is recorded as m, the image is divided by means of a sliding window to obtain multiple image groups, which can ensure that two adjacent image groups include tm frames of the same first image. Therefore, the smaller the value of the preset step size, the more identical images can be included in two adjacent image groups, and the more image groups the multiple frames of first images can be divided into, which is more conducive to subsequent gesture annotation.
[0073] As an example, all the first images of multiple frames may be acquired first, and then the first images may be divided according to the sliding windows to obtain multiple image groups.
[0074] As another example, t frames of first images that are arranged in time order and include the same palm can be first acquired, and the t frames of first images can be determined as the first image group. Then, the t+1th frame of first images that include the same palm as the first t frames of first images can be acquired, and the first images from the 2nd to t+1th frames can be determined as the second image group. The above steps can be repeated to obtain multiple image groups.
[0075] In this way, by dividing multiple frames of first images into multiple image groups in a sliding window manner, it can be ensured that the multiple image groups are arranged in chronological order and the order of t frames of first images included in each image group remains unchanged, thereby ensuring the orderliness and continuity of multiple gesture actions recognized subsequently and improving the accuracy of gesture annotation.
[0076] As an example, if multiple frames of images all include the same palm, the multiple frames of images can be determined as multiple frames of first images. If the multiple frames of images include not only the same palm but also other palms or do not include any palm, multiple frames of first images including the same palm can be obtained from the multiple frames of images.
[0077] Based on this, in some embodiments, the acquisition of multiple frames of first images arranged in time sequence and including the same palm may specifically include:
[0078] Acquire a plurality of consecutive frames of second images arranged in time sequence, where the second images are images in the gesture video;
[0079] For each frame of the second image, use the hand detection model to perform palm detection on the second image to obtain position information of the palm in the second image;
[0080] According to step 1, the position information of the palm in two adjacent frames of the second image is matched in sequence to obtain a matching result corresponding to the first target image, where the first target image is the second image in the latter frame of the two adjacent frames of the second image;
[0081] The second images with successful matching results in the multiple frames of second images are determined as the multiple frames of first images arranged in chronological order.
[0082] Here, the gesture video can be a video specially recorded by the user performing gesture actions. Among them, a gesture video can include one gesture action or multiple gesture actions. Gesture actions can include waving up, waving down, clenching a fist, opening, waving left, waving right, etc. In addition to gesture actions, gesture videos may also include dirty data that needs to be cleaned, such as starting actions, ending actions, changing palms, non-standard gesture actions, etc. If the scene for recording the gesture video is an interior scene of a vehicle, the gesture video may also include changing seats. Among them, the starting action and / or ending action may, for example, include raising your hand, letting go, etc.
[0083] As an example, after obtaining the gesture video, multiple frames of images in the gesture video may be named with timestamps and saved. By obtaining multiple frames of images according to timestamps, multiple frames of second images arranged in time sequence may be obtained.
[0084] In addition, the hand detection model may be a deep neural network model capable of performing palm detection. The hand detection model may include, for example, a residual neural network (ResNet) model. The hand detection model may perform palm detection on an image to obtain a palm detection result. The palm detection result may include the number of palms in the image and the position information of each palm in the image.
[0085] Based on this, as an example, after obtaining multiple frames of the second image, each frame of the second image can be processed in sequence. Specifically, after the second image is input into the hand detection model, the second image can be detected by the hand detection model to obtain the position information of the palm in the second image. Among them, the second image may not include the palm, may include one palm, or may include multiple palms. If the second image does not include the palm, the position information of the palm in the image can be replaced by a null value. If the second image includes multiple palms, the position information corresponding to the multiple palms can be obtained.
[0086] In addition, since the time difference between frames can be ignored, if the palm included in two adjacent frames is the same, the position information of the two palms may not differ much. Specifically, the position information matches successfully when the difference between the two position information is within a preset range.
[0087] Based on this, as an example, after obtaining the position information corresponding to multiple frames of second images, the position information in the current second image can be matched with the position information in the previous frame of the second image in sequence, starting from the second frame of the second image, according to step size 1. If both frames of the second image include a palm, then if the position information of the two frames of the second image is successfully matched, it can be determined that the palms in the two frames of the second image are the same palm, and the two frames of the second image are arranged in chronological order to obtain two frames of the first image including the same palm. Repeat the above steps to obtain multiple frames of the first image.
[0088] In addition, the latter second image in two adjacent second images can be recorded as the first target image. The former second image in two adjacent second images can be recorded as the second target image. Therefore, if the second target image includes a palm and the first target image includes two palms, it is possible to determine which palm in the first target image is the same as the palm in the second target image based on the position information, or both palms in the first target image are different from the palm in the second target image. If palm A in the first target image is the same as the palm in the second target image, the first target image and the second target image can be jointly determined as the first image, and palm A in the first target image can be located using a tracking algorithm in a subsequent process.
[0089] In this way, by acquiring multiple frames of first images from multiple frames of second images, data cleaning can be completed, the validity of the multiple frames of first images is guaranteed, and the labeling efficiency and accuracy are improved.
[0090] Based on this, in order to ensure that the multiple frames of first images are arranged in chronological order, in some embodiments, determining the second image with a successful matching result among the multiple frames of second images as the multiple frames of first images arranged in chronological order may specifically include:
[0091] When the matching result is successful, the first target image is added to the tracking sequence where the second target image is located, and the second target image is the previous second image of two adjacent second images;
[0092] The multiple frames of images in the tracking sequence are determined as multiple frames of first images arranged in time sequence.
[0093] Here, multiple frames of images in the same tracking sequence may include the same palm. In the case where the palm in the second target image is the target palm, since the matching is performed in sequence according to the arrangement order of the multiple frames of the second image, each time a frame of image corresponding to the target palm is matched, the image is added to the tracking sequence corresponding to the target palm, so that the multiple frames of the first image can be arranged in time order.
[0094] Based on this, in some embodiments, before adding the first target image to the tracking sequence where the second target image is located, the method may further include: determining the tracking sequence where the second target image is located.
[0095] Determining the tracking sequence where the second target image is located may specifically include:
[0096] Determine the first second image in the plurality of second image frames as the third target image, or, in the process of sequentially matching two adjacent second image frames, determine the next second image frame whose matching result with the previous second image frame is a failed match as the third target image;
[0097] Establishing a tracking sequence corresponding to each third target image, wherein each third target image is the first image in the corresponding tracking sequence, and the remaining images in the tracking sequence corresponding to the j-th third target image include the second image of the next two adjacent second images after the j-th third target image that successfully matches the previous second image, and j is a positive integer;
[0098] In the case that the palm in the third target image is the same as the palm in the second target image, the tracking sequence corresponding to the third target image is determined as the tracking sequence where the second target image is located.
[0099] Here, the third target image may be the first frame of the second image in the tracking sequence. After acquiring multiple frames of the second image, the first frame of the second image in the multiple frames of the second image may be first determined as the third target image, and a tracking sequence corresponding to the first frame of the second image may be established. Afterwards, the two adjacent frames of the second image are matched in sequence. The second frame of the second image is first matched with the first frame of the second image, and if the match is successful, the second frame of the second image is added to the tracking sequence where the first frame of the second image is located. Repeat the above steps of matching the two adjacent frames of the second image in sequence until the nth frame of the second image fails to match the n-1th frame of the second image, determine the nth frame of the second image as the third target image, and establish a tracking sequence corresponding to the nth frame of the second image. Continue to repeat the above steps of matching the two adjacent frames of the second image in sequence until multiple frames of the second image are matched, and multiple tracking sequences are obtained. Among them, the number of tracking sequences and the number of third target images may be the same. That is, if there are J tracking sequences, there may be J third target images. Among them, J is a positive integer. In addition, in the embodiment of the present application, j may be any one from 1 to J. That is, the jth third target image may be any one of the J third target images.
[0100] For example, it is assumed that a total of 5 consecutive second image frames arranged in chronological order are acquired. The 5 consecutive second image frames arranged in chronological order may include image 1, image 2, image 3, image 4, and image 5. Based on this, a first tracking sequence corresponding to image 1 may be established first, and image 1 may be determined as the first frame image in the first tracking sequence, and then image 2 may be matched with image 1. Assuming that image 2 is successfully matched with image 1, image 2 may be added to the first tracking sequence, and image 3 may be matched with image 2. In particular, image 2 may be the second frame image in the first tracking sequence. Assuming that image 3 is successfully matched with image 2, image 3 may be added to the first tracking sequence, and image 4 may be matched with image 3. In particular, image 3 may be the third frame image in the first tracking sequence. Assuming that image 4 fails to match image 3, a second tracking sequence corresponding to image 4 may be established, and image 4 may be determined as the first frame image in the second tracking sequence. Thereafter, image 5 may be matched with image 4. Assuming that image 5 is successfully matched with image 4, image 5 may be added to the second tracking sequence. In particular, image 5 may be the second frame image in the second tracking sequence. Since image 5 is the last frame of the 5-frame second image, the matching is stopped. So far, the first tracking sequence and the second tracking sequence corresponding to the 5-frame second image can be obtained. The first tracking sequence can be expressed as {image 1, image 2, image 3}, and the second tracking sequence can be expressed as {image 4, image 5}.
[0101] In addition, if the first frame of the second image includes a palm, a tracking sequence corresponding to the first frame of the second image can be established. If the first frame of the second image includes two palms, a tracking sequence corresponding to each of the two palms in the first frame of the second image can be established. Similarly, if the nth frame of the second image includes a palm, a tracking sequence corresponding to the nth frame of the second image can be established. If the nth frame of the second image includes two palms, a tracking sequence corresponding to each of the two palms in the nth frame of the second image can be established.
[0102] In this way, when the palm in the third target image is the same as the palm in the second target image, the tracking sequence corresponding to the third target image can be determined as the tracking sequence where the second target image is located.
[0103] In some embodiments, in S120, since each image group includes t frames of first images, the first hand key point information of the image group may be a collection of hand key point information corresponding to each of the t frames of first images. In addition, since the t frames of first images are arranged in time sequence, the hand key point information corresponding to each of the t frames of first images may be arranged in time sequence.
[0104] As an example, obtaining the first hand key point information may include obtaining the hand key point information corresponding to each of the t-frame first images.
[0105] Based on this, in some embodiments, the above S120 may specifically include:
[0106] For each first image frame in each image group, use the key point detection model to perform hand key point detection on the first image to obtain second hand key point information of each first image frame;
[0107] The t second hand key point information arranged in time sequence is determined as the first hand key point information.
[0108] Here, the key point detection model may be a deep neural network model capable of performing key point detection. The key point detection model may, for example, include a simple decoupled coordinate representation (Simple Disentagled coordinateRepresentation, SimDR) network model and a nested network (pixel-in-pixel net, PipNet) model. The key point detection model may perform hand key point detection on the palm in the first image to obtain the first coordinates of i hand key points. That is, the second hand key point information may include the first coordinates of i hand key points. Among them, i can specifically be 21. Based on this, the first hand key point information may be a first matrix of (t, 42) dimensions.
[0109] Based on this, in order to ensure the accuracy of the second hand key point information, in some embodiments, the above-mentioned key point detection model is used to perform hand key point detection on the first image to obtain the second hand key point information of each frame of the first image, which may specifically include:
[0110] The key point detection model is used to perform hand key point detection and palm recognition on each frame of the first image to obtain second hand key point information and palm recognition results. The palm recognition result represents whether the area corresponding to the second hand key point information is a hand area.
[0111] Here, the palm recognition result may be a binary classification result. For example, if the palm recognition result is 1, it can be determined that the area corresponding to the second hand key point information is the hand area. If the palm recognition result is 0, it can be determined that the area corresponding to the second hand key point information is not the hand area.
[0112] As an example, by adding a classifier to the open source key point detection model, the hand key point detection model can output not only the second hand key point information but also the palm recognition result. The classifier can be used to determine whether the area corresponding to the second hand key point information is the hand area.
[0113] In this way, by outputting the palm recognition result, it is possible to determine again whether the first image includes a palm. If the first image includes a palm, it is possible to avoid key point detection errors caused by palm detection errors, thereby ensuring the accuracy of the second hand key point information.
[0114] In some embodiments, in S130, the dynamic gesture recognition model may be a deep neural network model implemented based on a dynamic gesture recognition algorithm. The dynamic gesture recognition algorithm may include, for example, a dynamic time warping (Dynamic TimeWraping, DTW) algorithm and a multiple dimensions dynamic time warping (Multiple Dimensions Dynamic Time Wraping, ND-DTW) algorithm. The dynamic gesture recognition model may perform dynamic gesture recognition on the first hand key point information, and output a gesture recognition result and a score. Among them, the gesture recognition result may include a gesture action. The gesture recognition result may include, for example, an upward swing, a downward swing, a fist, an opening, a left swing, a right swing, and others. In addition, the score may characterize the credibility of the gesture recognition result. In the case where the first hand key point information is a first matrix of (t, 42) dimensions, the first matrix corresponding to the first hand key point information may be input into the dynamic gesture recognition model so that the dynamic gesture recognition model processes the first matrix to obtain a gesture recognition result and a score.
[0115] As an example, when using a dynamic gesture recognition model to perform dynamic gesture recognition, dynamic gesture recognition can be performed once for each image group obtained, or after obtaining multiple image groups, dynamic gesture recognition can be performed on multiple image groups in sequence, which is not limited here.
[0116] Based on this, in order to improve the gesture recognition efficiency, in some embodiments, the above-mentioned dynamic gesture recognition of the first hand key point information using the dynamic gesture recognition model may specifically include:
[0117] Subtract the t×i first coordinates in each image group from the first coordinates of the target hand key point in the third target image to obtain t×i second coordinates, where the third target image is any one of the t frames of the first image, and the target hand key point is any one of the i hand key points;
[0118] The dynamic gesture recognition model is used to perform dynamic gesture recognition on t×i second coordinates.
[0119] Here, the third target image may be, for example, the first frame image in the t-frame first image. The target hand key point may be, for example, the wrist key point. In addition, by processing the t×i second coordinates, a second matrix of (t, 42) dimensions may be obtained. The complexity of the second matrix may be less than the complexity of the first matrix.
[0120] In this way, by using the dynamic gesture recognition model to perform dynamic gesture recognition on the second matrix, the processing efficiency of the dynamic gesture recognition model can be improved, thereby improving the gesture recognition efficiency.
[0121] Based on this, in order to ensure the accuracy of the gesture recognition result, in some embodiments, the above S130 may specifically include:
[0122] When the area corresponding to the second hand key point information is a hand area, a dynamic gesture recognition model is used to perform dynamic gesture recognition on the first hand key point information.
[0123] Here, by performing dynamic gesture recognition when the second hand key point information is determined to be accurate, the accuracy of the gesture recognition result can be ensured.
[0124] In some embodiments, in S140, each image group may correspond to a gesture recognition result. The gesture recognition results corresponding to different image groups may be the same or different. If the gesture recognition results of multiple image groups are the same, the target gesture recognition result may be any one of the multiple gesture recognition results. If the gesture recognition results of multiple image groups are not exactly the same, but there are multiple consecutive image groups with the same gesture recognition results, the target gesture recognition result may be any one of the multiple consecutive and identical gesture recognition results. If there are no consecutive multiple image groups with the same gesture recognition results, it is possible to return to performing palm detection on the second image.
[0125] As an example, after obtaining multiple gesture recognition results, if there are multiple consecutive and identical gesture recognition results, the gesture recognition result can be determined as a target gesture recognition result, and the image group corresponding to the target gesture recognition result can be determined as a target image group.
[0126] Based on this, in order to accurately determine the target gesture action and the target image group, in some embodiments, the above S140 may specifically include:
[0127] When the gesture recognition result of the first image group is the same as the gesture recognition result of the second image group and different from the gesture recognition result of the third image group, determining the gesture recognition result of the first image group as the target gesture recognition result, wherein the second image group is an image group previous to the first image group and the third image group is an image group subsequent to the first image group;
[0128] The gesture recognition result is a target gesture recognition result, and a plurality of continuous image groups located before the first image group are determined as a plurality of target image groups.
[0129] Here, since the image group is obtained by intercepting the video sequence frames, the gesture action corresponding to one image group may not be complete. Therefore, by comprehensively analyzing the gesture recognition results of multiple image groups, the complete gesture action can be determined. Specifically, if multiple consecutive gesture recognition results are the same, it can be determined that the gesture action has not ended. If after multiple consecutive gesture recognition results are the same, the next gesture recognition result is different from the previous gesture recognition result, the previous gesture recognition result can be determined as the target gesture recognition result, and the target gesture recognition result can be stored.
[0130] Since the multiple image groups may include two identical but discontinuous gesture actions, the multiple target image groups may be multiple continuous image groups whose gesture recognition results are target gesture recognition results and are located before the first image group.
[0131] In some embodiments, in S150, the score may represent the credibility of the gesture recognition result. The higher the score, the higher the accuracy of the gesture recognition result.
[0132] As an example, after determining multiple target image groups and scores, multiple scores can be traversed to determine the highest score and the first hand key point information of the corresponding target image group, and the first hand key point information corresponding to the highest score is determined as the annotation result of the target gesture recognition result. After obtaining the annotation result, the target gesture recognition result and the corresponding first hand key point information can be stored.
[0133] In addition, in order to further prevent key point detection errors and / or dynamic gesture recognition errors caused by algorithm misdetection, the annotation results after storage can be manually inspected to further ensure the data quality of the annotation data.
[0134] In the embodiment of the present application, since the user only needs to accept the labeling results, the labor cost of labeling can be reduced compared with manual labeling.
[0135] In order to better describe the entire solution, some specific examples are given based on the above embodiments.
[0136] For example, Figure 2 As shown, a gesture data annotation method provided in an embodiment of the present application may include the following steps:
[0137] S21, obtaining a gesture video;
[0138] S22, acquiring a plurality of frames of second images arranged in time sequence from the gesture video;
[0139] S23, performing palm detection on the target image, determining that the palm in the target image is the target palm, and the target image is the first frame image in the multiple frames of second images;
[0140] S24, establishing a tracking sequence corresponding to the target palm;
[0141] S25, performing palm detection on the next frame image of the target image, and matching the palm detection result with the palm detection result of the target image;
[0142] S26, determine whether the match is successful, if so, execute S27, if not, execute S28;
[0143] S27, adding the next frame of the target image to the tracking sequence, and taking the next frame of the target image as the target image, returning to execute S25, until the detection of the multiple frames of the second image is completed, and determining the multiple frames of the image in the tracking sequence as the multiple frames of the first image;
[0144] S28, taking the next frame image of the target image as the target image, and returning to execute S24, until the detection of the multiple frames of the second image is completed, and determining the multiple frames of the image in the tracking sequence as the multiple frames of the first image;
[0145] S29, after the detection of the multiple frames of the second image is completed, the multiple frames of the image in the tracking sequence are determined as the multiple frames of the first image;
[0146] S210, moving the sliding window according to a preset step length to divide the multiple frames of first images into multiple image groups;
[0147] S211, performing hand key point detection on each image group in turn to obtain first key point information;
[0148] S212, subtracting the t×i first coordinates in the first key point information from the first coordinates of the target hand key point of the third target image respectively to obtain t×i second coordinates;
[0149] S213, using a dynamic gesture recognition model to perform dynamic gesture recognition on the t×i second coordinates, to obtain a gesture recognition result and score corresponding to the image group;
[0150] S214, determining whether there are multiple continuous and identical gesture recognition results, if so, executing S215, if not, executing S23;
[0151] S215, determining a plurality of continuous and identical gesture recognition results as target gesture recognition results, and determining a plurality of target image groups and scores corresponding to the target gesture recognition results;
[0152] S216, determining the first hand key point information of the target image group corresponding to the highest score among the multiple scores as the labeling result of the target gesture recognition result;
[0153] S217, storing the labeling result, that is, storing the target gesture recognition result and the first hand key point information of the corresponding target image group;
[0154] S218: Accept the stored annotation results.
[0155] Therefore, by using algorithms such as hand detection, key point detection, hand tracking and dynamic gesture recognition, we can automatically clean and label key points from gesture data, and finally generate a key point data set. Compared with manual cleaning and labeling, it not only reduces labor costs, but also greatly improves the efficiency of data cleaning and labeling.
[0156] Based on the gesture data annotation method provided in the above embodiment, the present application also provides a specific implementation of the gesture data annotation device. Please refer to the following embodiment.
[0157] like Figure 3 As shown, the gesture data annotation device 300 provided in the embodiment of the present application includes the following modules:
[0158] A first acquisition module 310 is used to acquire multiple image groups, each of which includes t frames of first images arranged in time sequence, the first images in the multiple image groups all include the same palm, and t is a positive integer;
[0159] A second acquisition module 320, used to acquire first hand key point information of each image group;
[0160] A recognition module 330 is used to perform dynamic gesture recognition on the first hand key point information for each image group using a dynamic gesture recognition model to obtain a gesture recognition result and score for each image group;
[0161] A first determination module 340 is used to determine a plurality of target image groups and scores corresponding to a target gesture recognition result, wherein the plurality of target image groups are continuous and the target gesture recognition result is any one of the gesture recognition results of the plurality of target image groups;
[0162] The second determination module 350 is used to determine the first hand key point information of the target image group corresponding to the highest score among the multiple scores as the labeling result of the target gesture recognition result.
[0163] The gesture data annotation device 300 is described in detail below, as follows:
[0164] In some embodiments, the first acquisition module 310 may specifically include:
[0165] An acquisition submodule, used for acquiring a plurality of first images arranged in time sequence and including a same palm;
[0166] The moving submodule is used to move the sliding window according to a preset step size to divide the multiple frames of first images into multiple image groups, and the size of the sliding window is the size of t frames of the first image.
[0167] In some embodiments, the acquisition submodule may specifically include:
[0168] An acquisition unit, used to acquire a plurality of consecutive frames of second images arranged in time sequence, where the second images are images in the gesture video;
[0169] A detection unit, configured to perform palm detection on the second image using a hand detection model for each frame of the second image, and obtain position information of the palm in the second image;
[0170] A matching unit, used to match the position information of the palm in two adjacent frames of the second image in sequence according to a step size of 1, to obtain a matching result corresponding to the first target image, where the first target image is the second image in the latter frame of the two adjacent frames of the second image;
[0171] The determining unit is used to determine the second images with successful matching results in the multiple frames of second images as the multiple frames of first images arranged in time sequence.
[0172] In some embodiments, the determining unit may specifically include:
[0173] An adding subunit, used for adding the first target image to the tracking sequence where the second target image is located when the matching result is successful, and the second target image is the second image of the previous frame in two adjacent frames of second images;
[0174] The first determining subunit is used to determine the multiple frames of images in the tracking sequence as multiple frames of first images arranged in time sequence.
[0175] In some embodiments, the determining unit may further include:
[0176] A second determination subunit is used to determine the first frame of the second image among the multiple frames of the second image as the third target image, or, in the process of sequentially matching two adjacent frames of the second image, determine the next frame of the second image whose matching result with the previous frame of the second image is a failed match as the third target image;
[0177] Establishing a subunit, used to establish a tracking sequence corresponding to each frame of the third target image, wherein each frame of the third target image is the first frame of the tracking sequence corresponding to it, and the remaining images in the tracking sequence corresponding to the j-th frame of the third target image include the second image of the next frame of the second image in the two adjacent frames of the second image after the j-th frame of the third target image that successfully matches the previous frame of the second image, and j is a positive integer;
[0178] The third determining subunit is used to determine the tracking sequence corresponding to the third target image as the tracking sequence of the second target image when the palm in the third target image is the same as the palm in the second target image.
[0179] In some embodiments, the second acquisition module 320 may specifically include:
[0180] A detection submodule, for performing hand key point detection on each first image frame in each image group using a key point detection model to obtain second hand key point information of each first image frame;
[0181] The first determining submodule is used to determine t second hand key point information arranged in time sequence as first hand key point information.
[0182] In some embodiments, the detection submodule may specifically include:
[0183] The recognition unit is used to perform hand key point detection and palm recognition on each frame of the first image using a key point detection model to obtain second hand key point information and a palm recognition result, wherein the palm recognition result indicates whether the area corresponding to the second hand key point information is a hand area.
[0184] Based on this, the identification module 330 may specifically include:
[0185] The first recognition submodule is used to perform dynamic gesture recognition using the first hand key point information of the dynamic gesture recognition model when the area corresponding to the second hand key point information is the hand area.
[0186] In some of the embodiments, each second hand key point information includes first coordinates of i hand key points, where i is a positive integer.
[0187] Based on this, the identification module 330 may specifically include:
[0188] a calculation submodule, used for subtracting the first coordinates of each image group from the first coordinates of the target hand key point of the third target image to obtain t×i second coordinates, where the third target image is any one of the t frames of the first image, and the target hand key point is any one of the i hand key points;
[0189] The second recognition submodule is used to perform dynamic gesture recognition on t×i second coordinates by using a dynamic gesture recognition model.
[0190] In some embodiments, the first determining module 340 may specifically include:
[0191] a second determination submodule, configured to determine the gesture recognition result of the first image group as a target gesture recognition result when the gesture recognition result of the first image group is the same as the gesture recognition result of the second image group and different from the gesture recognition result of the third image group, wherein the second image group is an image group previous to the first image group and the third image group is an image group subsequent to the first image group;
[0192] The third determination submodule is configured to determine a plurality of consecutive image groups whose gesture recognition results are target gesture recognition results and are located before the first image group as a plurality of target image groups.
[0193] The gesture data annotation device of the embodiment of the present application obtains the first hand key point information corresponding to each of the multiple image groups, and uses the dynamic gesture recognition model to perform dynamic gesture recognition on the first hand key point information to obtain the gesture recognition result of each image group, and can automatically determine the correspondence between the first hand key point information and the gesture recognition result. By selecting a group of hand key point information corresponding to the target gesture recognition result from multiple groups of hand key point information, instead of manually annotating the hand key point information after determining the target gesture action, the annotation efficiency can be improved. By determining multiple target image groups and scores corresponding to the target gesture recognition result, and determining the first hand key point information of the target image group corresponding to the highest score among the multiple scores as the annotation result of the target gesture recognition result, the key point information that best matches the target gesture recognition result can be selected from multiple groups of key point information, thereby ensuring data quality. In this way, through the embodiment of the present application, the annotation efficiency can be improved while ensuring data quality.
[0194] Based on the gesture data annotation method provided in the above embodiment, the embodiment of the present application also provides a specific implementation of the electronic device. Figure 4 A schematic diagram of an electronic device 400 provided in an embodiment of the present application is shown.
[0195] The electronic device 400 may include a processor 410 and a memory 420 storing computer program instructions.
[0196] Specifically, the processor 410 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0197] The memory 420 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 420 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 420 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 420 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 420 is a non-volatile solid-state memory.
[0198] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of the present application.
[0199] The processor 410 implements any one of the gesture data labeling methods in the above embodiments by reading and executing computer program instructions stored in the memory 420 .
[0200] In one example, the electronic device 400 may further include a communication interface 430 and a bus 440. Figure 4 As shown, the processor 410, the memory 420, and the communication interface 430 are connected via a bus 440 and communicate with each other.
[0201] The communication interface 430 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0202] Bus 440 includes hardware, software or both, and the parts of electronic equipment are coupled to each other.For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 440 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.
[0203] Exemplarily, the electronic device 400 may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA).
[0204] The electronic device can execute the gesture data annotation method in the embodiment of the present application, thereby realizing the combination Figures 1 to 3 The invention describes a method and device for labeling gesture data.
[0205] In addition, in combination with the gesture data annotation method in the above embodiment, the embodiment of the present application can provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by the processor, any one of the gesture data annotation methods in the above embodiment is implemented.
[0206] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0207] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0208] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0209] The above reference is according to the method of the embodiment of the present application, the flow chart of the device (system) and the computer program product and / or the block diagram described various aspects of the present application.It should be understood that each square box in the flow chart and / or the block diagram and the combination of each square box in the flow chart and / or the block diagram can be realized by computer program instructions.These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the realization of the function / action specified in one or more square boxes of the flow chart and / or the block diagram.Such a processor can be but is not limited to a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit.It can also be understood that each square box in the block diagram and / or the flow chart and the combination of the square boxes in the block diagram and / or the flow chart can also be realized by the dedicated hardware that performs the specified function or action, or can be realized by the combination of dedicated hardware and computer instructions.
[0210] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. A method for labeling gesture data, characterized in that: include: Acquire multiple image groups, each of the image groups includes t frames of first images arranged in time sequence, the first images in the multiple image groups all include the same palm, and t is a positive integer; Obtaining first hand key point information of each of the image groups; For each of the image groups, using a dynamic gesture recognition model to perform dynamic gesture recognition on the first hand key point information to obtain a gesture recognition result and score for each of the image groups; Determine a plurality of target image groups and scores corresponding to target gesture recognition results, wherein the plurality of target image groups are continuous and the target gesture recognition result is any one of the gesture recognition results of the plurality of target image groups; The first hand key point information of the target image group corresponding to the highest score among the multiple scores is determined as the labeling result of the target gesture recognition result.
2. The method according to claim 1, characterized in that The acquiring of multiple image groups comprises: Acquire multiple frames of the first image that are arranged in time sequence and include the same palm; The sliding window is moved according to a preset step size to divide multiple frames of the first image into multiple image groups, and the size of the sliding window is the size of t frames of the first image.
3. The method according to claim 2, characterized in that The acquiring of the plurality of frames of the first images, which are arranged in time sequence and include the same palm, comprises: Acquire a plurality of consecutive frames of second images arranged in time sequence, where the second images are images in the gesture video; For each frame of the second image, perform palm detection on the second image using a hand detection model to obtain position information of the palm in the second image; According to step 1, the position information of the palm in two adjacent frames of the second image is matched in sequence to obtain a matching result corresponding to the first target image, where the first target image is the second image in the latter frame of the two adjacent frames of the second image; The second images with successful matching results among the multiple frames of second images are determined as the multiple frames of first images arranged in chronological order.
4. The method according to claim 3, characterized in that The step of determining the second images with successful matching results among the multiple frames of second images as the multiple frames of the first images arranged in chronological order includes: When the matching result is successful, the first target image is added to the tracking sequence where the second target image is located, and the second target image is the second image of the previous frame in two adjacent frames of the second image; The multiple frames of images in the tracking sequence are determined as multiple frames of the first images arranged in time sequence.
5. The method according to claim 4, characterized in that Before adding the first target image to the tracking sequence where the second target image is located, the method further includes: Determine the first second image in the plurality of second images as the third target image, or, in the process of sequentially matching two adjacent second images, determine the next second image whose matching result with the previous second image is a failed match as the third target image; Establishing a tracking sequence corresponding to each frame of the third target image, wherein each frame of the third target image is the first frame of the tracking sequence corresponding to it, and the remaining images in the tracking sequence corresponding to the j-th frame of the third target image include the second image of the next frame of the second image in the two adjacent frames after the j-th frame of the third target image that successfully matches the second image of the previous frame, and j is a positive integer; In the case that the palm in the third target image is the same as the palm in the second target image, the tracking sequence corresponding to the third target image is determined as the tracking sequence where the second target image is located.
6. The method according to claim 1, characterized in that The obtaining of the first hand key point information of each of the image groups comprises: For each frame of the first image in each of the image groups, perform hand key point detection on the first image using a key point detection model to obtain second hand key point information of each frame of the first image; The t pieces of the second hand key point information arranged in time sequence are determined as the first hand key point information.
7. The method according to claim 6, characterized in that The step of performing hand key point detection on the first image using a key point detection model to obtain second hand key point information of each frame of the first image includes: Using a key point detection model to perform hand key point detection and palm recognition on each frame of the first image, to obtain the second hand key point information and a palm recognition result, wherein the palm recognition result indicates whether an area corresponding to the second hand key point information is a hand area; The step of using a dynamic gesture recognition model to perform dynamic gesture recognition on the first hand key point information includes: When the area corresponding to the second hand key point information is a hand area, the dynamic gesture recognition model is used to perform dynamic gesture recognition on the first hand key point information.
8. The method according to claim 7, characterized in that Each of the second hand key point information includes first coordinates of i hand key points, where i is a positive integer; The step of using a dynamic gesture recognition model to perform dynamic gesture recognition on the first hand key point information includes: Subtracting the first coordinates of each of the image groups from the first coordinates of the target hand key point of the third target image to obtain t×i second coordinates, wherein the third target image is any one of the t frames of the first images, and the target hand key point is any one of the i hand key points; The dynamic gesture recognition model is used to perform dynamic gesture recognition on the t×i second coordinates.
9. The method according to claim 1, characterized in that: The determining of a plurality of target image groups corresponding to the target gesture recognition results includes: When a gesture recognition result of the first image group is the same as a gesture recognition result of the second image group and different from a gesture recognition result of the third image group, determining the gesture recognition result of the first image group as a target gesture recognition result, wherein the second image group is an image group previous to the first image group and the third image group is an image group subsequent to the first image group; The gesture recognition result is a target gesture recognition result, and a plurality of continuous image groups located before the first image group are determined as the plurality of target image groups.
10. A gesture data labeling device, characterized in that: The device comprises: A first acquisition module is used to acquire a plurality of image groups, each of which includes t frames of first images arranged in time sequence, the first images in the plurality of image groups all include the same palm, and t is a positive integer; A second acquisition module, used for acquiring first hand key point information of each of the image groups; a recognition module, configured to perform dynamic gesture recognition on the first hand key point information for each of the image groups using a dynamic gesture recognition model, to obtain a gesture recognition result and score for each of the image groups; A first determination module is used to determine a plurality of target image groups and scores corresponding to a target gesture recognition result, wherein the plurality of target image groups are continuous and the target gesture recognition result is any one of the gesture recognition results of the plurality of target image groups; The second determination module is used to determine the first hand key point information of the target image group corresponding to the highest score among the multiple scores as the labeling result of the target gesture recognition result.
11. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the method for labeling gesture data as described in any one of claims 1-9 is implemented.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the method for labeling gesture data according to any one of claims 1 to 9 is implemented.