Key point detection method and device, equipment and storage medium

By using palmar skeleton information in the key point detection model for feature extraction, the problem of low key point detection accuracy when the hand is small or blocked in the prior art is solved, and a higher accuracy of key point detection of hand is achieved.

CN119963854APending Publication Date: 2025-05-09BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311489709.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, in the case where the hand is small or there is occlusion in the hand image, the accuracy of hand key point detection based on the PipNet model is low.

Method used

By acquiring hand image and palm skeleton information, the feature extraction layer and graph convolution neural network layer in the key point detection model perform feature extraction of palm skeleton information and feature information, and convert it into the target coordinates of hand key points in the hand image.

Benefits of technology

The accuracy of hand key points detection is improved, especially when the hand is small or there is occlusion in the hand image, the relative positions between multiple hand key points can be reasonably determined to obtain target coordinates with high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963854A_ABST
    Figure CN119963854A_ABST
Patent Text Reader

Abstract

The invention discloses a key point detection method and device, equipment and a storage medium. The method comprises the following steps: acquiring a hand image and palm skeleton information; using a feature extraction layer in the key point detection model to perform feature extraction of hand key points on the hand image to obtain first feature information, the first feature information being feature information of the hand key points; performing feature extraction on the palm skeleton information and the first feature information by using a graph convolutional neural network layer in the key point detection model to obtain second feature information, the second feature information being feature information of a hand key point including the palm skeleton information; and converting the second feature information into a target coordinate of the hand key point in the hand image. According to the key point detection method provided by the embodiment of the invention, the accuracy of key point detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of key point detection, and in particular, relates to a key point detection method, device, equipment and storage medium. Background Art

[0002] With the development of artificial intelligence, the application of dynamic gesture recognition technology is becoming more and more extensive. And key point detection technology is the basis of dynamic gesture recognition technology.

[0003] At present, key point detection technology can be implemented based on a nested network (pixel-in-pixel net, PipNet). The PipNet model can first divide the image into a larger grid, and then convert the detection of key points into grid heat map regression and prediction of horizontal and vertical offsets, and use the constraints of neighboring points to correct the predicted key points to obtain the predicted coordinates of the key points. Specifically, the PipNet model can only detect key points based on hand images. Moreover, the coordinates of the key points can only be predicted when the key points in the hand image are clearly visible. If the key points in the hand image are not obvious or are obscured, the PipNet model cannot predict the coordinates. If the predicted target includes unclear or obscured key points, the deviation between the predicted coordinates and the actual coordinates will be large.

[0004] Thus, when the hand in the image is small and / or the hand in the image is occluded, the accuracy of hand key point detection based on the PipNet model is low. Summary of the invention

[0005] The embodiments of the present application provide a key point detection method, apparatus, device and storage medium, which can improve the accuracy of hand key point detection.

[0006] In a first aspect, an embodiment of the present application provides a key point detection method, the method comprising:

[0007] Get hand image and palm skeleton information;

[0008] Extracting features of the hand key points from the hand image using a feature extraction layer in a key point detection model to obtain first feature information, where the first feature information is feature information of the hand key points;

[0009] Using the graph convolutional neural network layer in the key point detection model to perform feature extraction on the palm skeleton information and the first feature information to obtain second feature information, where the second feature information is feature information of the hand key points including the palm skeleton information;

[0010] The second feature information is converted into target coordinates of the hand key points in the hand image.

[0011] In a possible implementation, the feature extraction layer includes a hand feature extraction layer and a nested network PipNet layer, and the feature extraction layer in the key point detection model is used to extract the features of the hand key points of the hand image to obtain the first feature information, including:

[0012] Using the hand feature extraction layer to extract hand features from the hand image, to obtain third feature information, where the third feature information is hand feature information;

[0013] The PipNet layer is used to extract the feature information of the hand key points from the third feature information to obtain the first feature information.

[0014] In a possible implementation, the graph convolutional neural network layer includes multiple convolutional layers, and the graph convolutional neural network layer in the key point detection model is used to extract features from the palm skeleton information and the first feature information to obtain the second feature information, including:

[0015] Using a first convolutional layer among the multiple convolutional layers to perform feature extraction on the palm skeleton information and the first feature information to obtain fourth feature information;

[0016] Using a target convolution layer to perform feature extraction on the palm skeleton information, the first feature information and the first target feature information to obtain second target feature information, the target convolution layer is the i-th convolution layer among the multiple convolution layers, the first target feature information is the feature information output by the (i-1)-th convolution layer, and the first target feature information includes the fourth feature information, wherein i is a positive integer greater than or equal to 2;

[0017] Let i=i+1, and repeat the step of obtaining the second target feature information until the target convolution layer is the last convolution layer among the multiple convolution layers, and determine the second target feature information obtained most recently as the second feature information.

[0018] In a possible implementation manner, converting the second feature information into target coordinates of the hand key point in the hand image includes:

[0019] For a target hand key point, converting the second feature information into a first coordinate of the target hand key point in the hand image, wherein the target hand key point is any one of a plurality of hand key points;

[0020] Acquire a second coordinate of the target hand key point in the hand image, where the second coordinate is a coordinate of the target hand key point when it is a neighboring point of other hand key points, where the other hand key points are hand key points other than the target hand key point among the multiple hand key points;

[0021] An average value of the first coordinate and the second coordinate is calculated to obtain the target coordinate of the target hand key point in the hand image.

[0022] In a possible implementation manner, obtaining the second coordinate of the target hand key point in the hand image includes:

[0023] For each hand key point among the multiple hand key points, obtaining a third coordinate of a neighboring point of the hand key point in the hand image;

[0024] Determine the target hand key point among the neighboring points corresponding to each of the plurality of hand key points;

[0025] The third coordinate corresponding to the target hand key point is determined as the second coordinate.

[0026] In a possible implementation, the first feature information includes feature information corresponding to the neighboring point, and acquiring, for each hand key point among the multiple hand key points, a third coordinate of a neighboring point of the hand key point in the hand image includes:

[0027] For each hand key point in the plurality of hand key points, determining a neighboring point of the hand key point in the plurality of hand key points;

[0028] Determining feature information corresponding to the neighboring point in the first feature information;

[0029] The feature information corresponding to the neighboring point is converted into the third coordinate of the neighboring point in the hand image.

[0030] In a possible implementation, before extracting features of hand key points from the hand image using the feature extraction layer in the key point detection model, the method further includes:

[0031] Acquire a hand sample image, label coordinates corresponding to each hand key point in the hand sample image, and the palm skeleton information, wherein the label coordinates are obtained by manual annotation;

[0032] Extracting features of hand key points from the hand sample image using the initial feature extraction layer in the initial key point detection model to obtain first predicted feature information of each hand key point;

[0033] Using the initial graph convolutional neural network layer in the initial key point detection model to perform feature extraction on the palm skeleton information and the first predicted feature information to obtain second predicted feature information of each of the hand key points, wherein the second predicted feature information includes the palm skeleton information;

[0034] Converting the second predicted feature information of each of the hand key points into predicted coordinates of each of the hand key points in the hand sample image;

[0035] Determine a loss function value according to a difference between the predicted coordinates and the label coordinates of each of the hand key points;

[0036] The model parameters of the initial feature extraction layer and the initial graph convolutional neural network layer are adjusted according to the loss function value, and the key point detection model is obtained by training.

[0037] In a second aspect, an embodiment of the present application provides a key point detection device, the device comprising:

[0038] A first acquisition module is used to acquire a hand image and palm skeleton information;

[0039] A first extraction module, configured to extract features of hand key points from the hand image using a feature extraction layer in a key point detection model to obtain first feature information, where the first feature information is feature information of the hand key points;

[0040] A second extraction module is used to perform feature extraction on the palm skeleton information and the first feature information by using the graph convolutional neural network layer in the key point detection model to obtain second feature information, where the second feature information is feature information of the hand key points including the palm skeleton information;

[0041] The first conversion module is used to convert the second feature information into the target coordinates of the hand key point in the hand image.

[0042] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising: a processor and a memory storing computer program instructions;

[0043] When the processor executes the computer program instructions, the processor implements any possible implementation method of the first aspect described above.

[0044] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implements a method in any possible implementation method of the first aspect described above.

[0045] The key point detection method, apparatus, device and storage medium of the embodiments of the present application extract features from palm skeleton information and first feature information by using the graph convolutional neural network layer in the key point detection model to obtain second feature information, and convert the second feature information into the target coordinates of the hand key points in the hand image. That is, compared with key point detection based only on the hand image, by adding palm skeleton information in the process of using the key point detection model to detect hand key points, the key point detection model can be constrained to be optimized in a direction consistent with the human hand skeleton, so that when the hand key points in the hand image are not obvious or are blocked, the key point detection model can also reasonably determine the relative positions between multiple hand key points. In this way, when the hand in the hand image is small and not obvious, and / or the hand in the hand image is blocked, the hand key point detection can be reasonably and accurately performed to obtain the target coordinates with higher accuracy, thereby improving the accuracy of hand key point detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] Figure 1 It is a flowchart of a key point detection method provided in an embodiment of the present application;

[0048] Figure 2 is a palm skeleton diagram provided by an embodiment of the present application;

[0049] Figure 3 is a schematic diagram of feature extraction using GCN provided in an embodiment of the present application;

[0050] Figure 4 is a schematic diagram of a key point detection method provided in an embodiment of the present application;

[0051] Figure 5 It is a structural schematic diagram of a key point detection device provided in an embodiment of the present application;

[0052] Figure 6 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the scheme of the present application will be further described below. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0054] In the following description, many specific details are set forth to facilitate a full understanding of the present application, but the present application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only part of the embodiments of the present application, rather than all of the embodiments.

[0055] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprises" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0056] As described in the background technology section, in order to solve the problems of the prior art, the embodiments of the present application provide a key point detection method, apparatus, device and storage medium. The key point detection method can be applied to the scene of dynamic gesture recognition.

[0057] The following first introduces the key point detection method provided in the embodiment of the present application.

[0058] Figure 1 FIG. 1 is a flow chart of a key point detection method provided in an embodiment of the present application. The key point detection method can be executed by any module including a key point detection model. Figure 1 As shown, the key point detection method provided in the embodiment of the present application includes the following steps:

[0059] S110, obtaining a hand image and palm skeleton information;

[0060] S120, extracting features of the hand key points from the hand image using the feature extraction layer in the key point detection model to obtain first feature information, where the first feature information is feature information of the hand key points;

[0061] S130, using the graph convolutional neural network layer in the key point detection model to perform feature extraction on the palm skeleton information and the first feature information to obtain second feature information, where the second feature information is feature information of the hand key points including the palm skeleton information;

[0062] S140: Convert the second feature information into target coordinates of the hand key points in the hand image.

[0063] The key point detection method of the embodiment of the present application uses the graph convolutional neural network layer in the key point detection model to extract features from the palm skeleton information and the first feature information to obtain the second feature information, and converts the second feature information into the target coordinates of the hand key points in the hand image. That is, compared with key point detection based only on the hand image, by adding palm skeleton information in the process of using the key point detection model to detect hand key points, the key point detection model can be constrained to be optimized in a direction that conforms to the human hand skeleton, so that when the hand key points in the hand image are not obvious or are obscured, the key point detection model can also reasonably determine the relative positions between multiple hand key points. In this way, when the hand in the hand image is small and not obvious, and / or the hand in the hand image is obscured, it is possible to reasonably and accurately perform hand key point detection, obtain target coordinates with higher accuracy, and improve the accuracy of hand key point detection.

[0064] The specific implementation methods of the above steps are introduced below.

[0065] In some embodiments, in S110, the hand image may be an image including a hand. The state of the hand may be an open state, a partially open state, a clenched fist state, a partially clenched fist state, etc. When the hand is partially open or partially clenched, some fingers may be blocked. In addition, when there are multiple hand images, the sizes of the multiple hand images may be the same. The pixels of the hand image may be, for example, 192x192 pixels.

[0066] In addition, the palm skeleton information may include multiple key points in the palm and association information between the multiple key points. That is, the palm skeleton information may actually be the topological structure information between the multiple key points in the palm. In the case where the palm includes 21 key points, the palm skeleton diagram may be as follows: Figure 2 As shown in Figure 2, by extracting and encoding the topological structure between multiple key points in the palm skeleton graph, an adjacency matrix can be obtained. Figure 2 The corresponding adjacency matrix A can be shown as follows:

[0067]

[0068] In the adjacency matrix A, the elements on the diagonal are all 1, indicating that each key point forms a self-loop with itself, so that the key point information of each hand key point itself will not be lost in the subsequent calculation process using the adjacency matrix A.

[0069] In some embodiments, in S120, the first feature information may be feature information of a key point of the hand.

[0070] As an example, the key point detection model may include a feature extraction layer. The feature extraction layer may be used to extract features of the hand key points of the hand image. After the hand image is input into the key point detection model, it may first pass through the feature extraction layer, and use the feature extraction layer to extract features of the hand key points of the hand image to obtain first feature information.

[0071] In some embodiments, the feature extraction layer may include a hand feature extraction layer and a nested network (pixel-in-pixel net, PipNet) layer. The hand feature extraction layer may be used to extract hand features from a hand image. The PipNet layer may be used to extract key point features from hand features.

[0072] Based on this, the above S120 may specifically include:

[0073] Using the hand feature extraction layer to extract hand features from the hand image, to obtain third feature information, where the third feature information is hand feature information;

[0074] The PipNet layer is used to extract the feature information of the key points of the hand from the third feature information to obtain the first feature information.

[0075] Here, the hand feature information may include feature information of the key points of the hand and feature information of the non-key points of the hand. In addition, the hand feature extraction layer may be implemented based on a residual neural network (Residual Neural Network, ResNet). For example, ResNet18 may be used as the backbone network for hand feature extraction. The extracted third feature information may be, for example, (512, 6, 6).

[0076] After obtaining the third feature information, the third feature information can be directly regressed by convolution operation using the grid division and horizontal offset and vertical offset of PipNet to obtain the first feature information. The first feature information may include Map, Offset_x and Offset_y, for example. The processing logic of PipNet may be to first divide the hand image into a plurality of large grids, and then divide each large grid into a plurality of small grids, and determine the first feature information by the position of the key points in the large grid and the small grid. Based on this, Map may be, for example, a tensor of (21, 6, 6), indicating the position of each key point in a 6×6 image grid. In (21, 6, 6), 21 may be the number of key points. In addition, Offset_x may represent the horizontal offset of the current key point in the current small grid for the target small grid. The current small grid and the target small grid may be in the same large grid, and the target small grid may be any small grid in the large grid. Specifically, the target small grid may be, for example, the small grid in the upper left corner of the large grid. Offset_y can represent the vertical offset of the current key point in the current small grid relative to the target small grid.

[0077] In some embodiments, in S130, the second feature information may be feature information of the key points of the hand after skeleton constraints are performed, wherein the skeleton constraints can optimize the key points predicted by the model in a direction that conforms to the human palm skeleton.

[0078] As an example, the key point detection model may also include a graph convolutional network (GCN). After obtaining the first feature information, the graph convolutional neural network layer may perform feature extraction on the palm skeleton information and the first feature information to obtain the second feature information. The graph convolutional neural network layer may include one convolutional layer or multiple convolutional layers, which is not limited here.

[0079] As a more specific example, after obtaining Map, Offset_x and Offset_y, the graph convolutional neural network layer can obtain Map' after skeleton constraint by performing feature extraction on the adjacency matrix A and Map. The graph convolutional neural network layer can obtain Offset_x' after skeleton constraint by performing feature extraction on the adjacency matrix A and Offset_x. The graph convolutional neural network layer can obtain Offset_y' after skeleton constraint by performing feature extraction on the adjacency matrix A and Offset_y. When GCN includes one convolutional layer, the formula for feature extraction through GCN can be shown as formula (1):

[0080]

[0081] In formula (1), X may represent the input of GCN, and X may be any one of Map, Offset_x, and Offset_y; H(X,A) may represent the output of GCN, i.e., the second feature information; A may represent an adjacency matrix with self-loops; It can represent the symmetric normalized matrix of A; W can represent the weight matrix of GCN; σ can be a coefficient.

[0082] When multiple convolutional layers are included in GCN, the formula for feature extraction through GCN can be shown as formula (2):

[0083]

[0084] In formula (2), H i+1 (X, A) can represent the output of the i-th convolutional layer, W i It can represent the weight matrix of the first convolutional layer. The meanings of other symbols can be found in formula (1) and will not be repeated here.

[0085] When multiple convolutional layers are included in the GCN, the increased complexity of the model may lead to network degradation, affecting the accuracy of feature extraction. Therefore, in order to avoid the network degradation problem caused by the complexity of the model and ensure the accuracy of feature extraction, in some embodiments, the above S130 may specifically include:

[0086] Using a first convolution layer among the multiple convolution layers to extract features from the palm skeleton information and the first feature information, to obtain fourth feature information;

[0087] Using a target convolution layer to perform feature extraction on the palm skeleton information, the first feature information and the first target feature information to obtain second target feature information, the target convolution layer is the i-th convolution layer among the multiple convolution layers, the first target feature information is the feature information output by the (i-1)-th convolution layer, and the first target feature information includes the fourth feature information, wherein i is a positive integer greater than or equal to 2;

[0088] Let i=i+1, and repeat the step of obtaining the second target feature information until the target convolution layer is the last convolution layer among the multiple convolution layers, and the second target feature information obtained most recently is determined as the second feature information.

[0089] Here, the feature extraction process of the first convolutional layer can be implemented by formula (1), and the fourth feature information obtained can be recorded as H 1 (X,A). After getting H 1 (X, A) After that, the residual F between the first convolutional layer and the second convolutional layer can be calculated by the following formula (3): 1 (X,A):

[0090] F i (X, A)= H i (X, A)+X (3)

[0091] In the residual F 1 (X,A) After that, F 1 (X, A) and A are input into the second convolutional layer together, and feature extraction is performed using formula (2) to obtain H 2 (X, A). Repeat the above steps until the last convolutional layer is calculated and the second feature information is obtained.

[0092] In addition, in formula (3), in order to ensure H i (X,A) and X can be added together, so that H i The dimensions of (X, A) are the same as those of X. The number of nodes in each convolutional layer can be, for example, 108 (6×6×3).

[0093] In the case where GCN includes three convolutional layers, the schematic diagram of feature extraction using GCN can be as follows Figure 3 shown.

[0094] In this way, by adding residual links between the GCNs in each layer, the original information of the key points of the hand can be effectively retained, and the convergence of the accelerated model can be avoided, thus ensuring the accuracy of feature extraction.

[0095] In some embodiments, in S140, the target coordinates may be specific coordinates of the hand key points in the hand image predicted by the model, wherein the target coordinates may include horizontal coordinates and vertical coordinates.

[0096] As an example, by combining and calculating Map', Offset_x' and Offset_y' in the second feature information, the specific coordinates of the hand key points in the hand image can be obtained.

[0097] Based on this, in order to further improve the accuracy of the target coordinates, in some embodiments, the above S140 may specifically include:

[0098] For a target hand key point, convert the second feature information into a first coordinate of the target hand key point in the hand image, where the target hand key point is any one of the multiple hand key points;

[0099] Obtaining a second coordinate of the target hand key point in the hand image, where the second coordinate is the coordinate of the target hand key point when it is a neighboring point of other hand key points, where the other hand key points are hand key points other than the target hand key point among the multiple hand key points;

[0100] The average value of the first coordinate and the second coordinate is calculated to obtain the target coordinate of the target hand key point in the hand image.

[0101] Here, the neighboring point of the target hand key point may be a hand key point whose distance to the target hand key point is less than a preset threshold. Each hand key point may correspond to at least one neighboring point. The number of neighboring points may be preset. For example, the number of neighboring points of each hand key point may be 4.

[0102] As an example, a target hand keypoint may have multiple neighboring points, and may also serve as a neighboring point of other hand keypoints.

[0103] As an example, the key point detection model can predict the specific coordinates of the target hand key point and the specific coordinates of the neighboring points. Therefore, the specific coordinates of the target hand key point include the first coordinate and the second coordinate.

[0104] In this way, by calculating the average value of the first coordinate and the second coordinate, the target coordinate of the target hand key point in the hand image is obtained, which can further improve the accuracy of the target coordinate.

[0105] Based on this, in some embodiments, the above-mentioned obtaining the second coordinate of the target hand key point in the hand image may specifically include:

[0106] For each hand key point among the multiple hand key points, obtain the third coordinate of a neighboring point of the hand key point in the hand image;

[0107] Determine a target hand key point among neighboring points corresponding to each of the plurality of hand key points;

[0108] The third coordinate corresponding to the target hand key point is determined as the second coordinate.

[0109] Here, each hand key point may correspond to at least one neighboring point, and may also serve as a neighboring point of other hand key points. For example, the neighboring points of key point A may include key point B, key point C, and key point D. The neighboring points of key point B may include key point A, key point C, and key point E.

[0110] After obtaining the neighboring points corresponding to the plurality of hand key points, the plurality of neighboring points may be determined as a neighboring point set. The neighboring point set may not include the target hand key point, may include one target hand key point, or may include multiple target hand key points. For example, the neighboring point set may include one A key point, one B key point, two C key points, one D key point, and one E key point.

[0111] In the case where the target hand key point is included in the neighboring point set, at least one target hand key point may be determined in the neighboring point set, and the third coordinate corresponding to the target hand key point in the neighboring point set may be determined as the second coordinate. For example, if the coordinate of key point A in the neighboring point set is the third coordinate, then in the case where key point A is the target key point, the third coordinate may be determined as the second coordinate of the target key point.

[0112] In some embodiments, the first feature information may include feature information corresponding to the neighboring points. Based on this, for each of the multiple hand key points, obtaining the third coordinate of the neighboring point of the hand key point in the hand image may specifically include:

[0113] For each hand key point among the plurality of hand key points, determining a neighboring point of the hand key point among the plurality of hand key points;

[0114] Determining feature information corresponding to the neighboring point in the first feature information;

[0115] The feature information corresponding to the neighboring points is converted into the third coordinates of the neighboring points in the hand image.

[0116] Here, the feature information corresponding to the neighboring point may include Neighbor_x and Neighbor_y. Neighbor_x may, for example, represent the horizontal offset of the neighboring point of the current key point in the current small grid relative to the target small grid. Neighbor_y may, for example, represent the vertical offset of the neighboring point of the current key point in the current small grid relative to the target small grid.

[0117] As an example, by combining and calculating Map, Neighbor_x and Neighbor_y in the first feature information, the specific coordinates of the neighboring points in the hand image, ie, the third coordinates, can be obtained.

[0118] Based on this, in order to improve the accuracy of the key point detection model and thus improve the accuracy of key point detection, in some embodiments, before the above S120, the following may also be included:

[0119] Obtain hand sample images, label coordinates corresponding to each hand key point in the hand sample images, and palm skeleton information, where the label coordinates are obtained through manual annotation;

[0120] Extracting features of hand key points from the hand sample image using the initial feature extraction layer in the initial key point detection model to obtain first predicted feature information of each hand key point;

[0121] Using the initial graph convolutional neural network layer in the initial key point detection model to perform feature extraction on the palm skeleton information and the first predicted feature information, obtaining second predicted feature information of each hand key point, wherein the second predicted feature information includes the palm skeleton information;

[0122] Convert the second predicted feature information of each hand key point into the predicted coordinates of each hand key point in the hand sample image;

[0123] Determine the loss function value based on the difference between the predicted coordinates and the label coordinates of each hand key point;

[0124] The model parameters of the initial feature extraction layer and the initial graph convolutional neural network layer are adjusted according to the loss function value, and the key point detection model is trained.

[0125] Here, the hand sample image may include multiple hand key points. The label coordinates may be the actual coordinates of the hand key points in the hand sample image obtained by pre-marking. In addition, the first predicted feature information may be the predicted feature information of the hand key points excluding the palm skeleton information. The second predicted feature information may be the predicted feature information of the hand key points including the palm skeleton information.

[0126] In addition, the initial key point detection model may include an initial feature extraction layer and an initial graph convolutional neural network layer. The key point detection model may include a feature extraction layer and a graph convolutional neural network layer. Among them, the initial key point detection model may be a neural network model that has not been trained for the detection task of hand key points. The key point detection model may be a neural network model that can detect hand key points after training. Similarly, the initial feature extraction layer may be a neural network structure that has not been trained for the feature extraction task of hand key points. The feature extraction layer may be a neural network structure that can extract features of hand key points after training. The initial graph convolutional neural network layer may be a neural network structure that has not been trained for the target task. Among them, the target task may be to further extract features from the palm skeleton information and the feature information of the hand key points. Based on this, the graph convolutional neural network layer may be a neural network structure that can further extract features from the palm skeleton information and the feature information of the hand key points after training.

[0127] In addition, the loss function corresponding to the initial feature extraction layer can be expressed as L R , the loss function corresponding to the initial graph convolutional neural network layer can be expressed as L G Therefore, the loss function corresponding to the initial key point detection model can be expressed by the following formula (4):

[0128] L=L R +L G (4)

[0129] In formula (4), since GCN can correct the skeleton constraints of the three branches of Map, Offset_x and Offset_y, G These three branches can be included in L R The five branches may include Map, Offset_x, Offset_y, Neighbor_x and Neighbor_y.

[0130] As an example, determining the loss function value according to the difference between the predicted coordinates and the label coordinates may specifically include:

[0131] Upsample the label coordinates and convert them into label features using the PipNet conversion idea;

[0132] After using the initial key point detection model to perform key point detection to obtain the features of the hand key points, the label features and the features of the hand key points are input into formula (4), and the loss function value is calculated using formula (4).

[0133] In this way, by adding the GCN network containing palm skeleton information to the loss function, and then guiding the key point detection model to optimize in a direction that is more in line with the human palm structure, the accuracy of the key point detection model can be improved, and thus the accuracy of key point detection can be improved.

[0134] In order to better describe the entire solution, some specific examples are given based on the above embodiments.

[0135] For example, Figure 4 As shown, the feature extraction layer can extract hand features from the hand image to obtain first feature information. GCN can perform constraint correction on the first feature information based on the palm skeleton information to obtain second feature information. By post-processing the second feature information, that is, converting the second feature information into specific coordinates of the hand key points in the hand image, a hand image including the hand key points can be obtained.

[0136] Therefore, by detecting the key points of the hand based on the palm skeleton information, the accuracy of key point detection can be improved, and difficult data under occlusion can be predicted.

[0137] Based on the key point detection method provided in the above embodiment, the present application also provides a specific implementation of the key point detection device. Please refer to the following embodiment.

[0138] like Figure 5 As shown, the key point detection device 500 provided in the embodiment of the present application includes the following modules:

[0139] A first acquisition module 510 is used to acquire a hand image and palm skeleton information;

[0140] A first extraction module 520 is used to extract the features of the hand key points from the hand image using the feature extraction layer in the key point detection model to obtain first feature information, where the first feature information is the feature information of the hand key points;

[0141] A second extraction module 530 is used to perform feature extraction on the palm skeleton information and the first feature information using the graph convolutional neural network layer in the key point detection model to obtain second feature information, where the second feature information is feature information of the hand key points including the palm skeleton information;

[0142] The first conversion module 540 is used to convert the second feature information into target coordinates of the hand key points in the hand image.

[0143] The key point detection device 500 is described in detail below, as shown below:

[0144] In some embodiments, the feature extraction layer includes a hand feature extraction layer and a nested network PipNet layer. Based on this, the first extraction module 520 may specifically include:

[0145] A first extraction submodule is used to extract hand features from the hand image using the hand feature extraction layer to obtain third feature information, where the third feature information is hand feature information;

[0146] The second extraction submodule is used to extract the feature information of the hand key points from the third feature information by using the PipNet layer to obtain the first feature information.

[0147] In some embodiments, the graph convolutional neural network layer includes multiple convolutional layers. Based on this, the second extraction module 530 may specifically include:

[0148] A third extraction submodule is used to extract features from the palm skeleton information and the first feature information using the first convolution layer among the multiple convolution layers to obtain fourth feature information;

[0149] a fourth extraction submodule, configured to perform feature extraction on the palm skeleton information, the first feature information and the first target feature information by using a target convolution layer to obtain second target feature information, wherein the target convolution layer is the i-th convolution layer among the multiple convolution layers, the first target feature information is the feature information output by the (i-1)-th convolution layer, and the first target feature information includes the fourth feature information, wherein i is a positive integer greater than or equal to 2;

[0150] A determination submodule is used to set i=i+1 and repeat the step of obtaining the second target feature information until the target convolution layer is the last convolution layer among the multiple convolution layers, and the second target feature information obtained most recently is determined as the second feature information.

[0151] In some embodiments, the first conversion module 540 may specifically include:

[0152] A conversion submodule, for converting the second feature information into a first coordinate of the target hand key point in the hand image, with respect to the target hand key point, where the target hand key point is any one of the multiple hand key points;

[0153] An acquisition submodule, used to acquire a second coordinate of a target hand key point in the hand image, the second coordinate being a coordinate of the target hand key point when it is a neighboring point of other hand key points, and the other hand key points are hand key points other than the target hand key point among the multiple hand key points;

[0154] The calculation submodule is used to calculate the average value of the first coordinate and the second coordinate to obtain the target coordinate of the target hand key point in the hand image.

[0155] In some embodiments, the acquisition submodule may specifically include:

[0156] An acquisition unit, configured to acquire, for each of the plurality of hand key points, a third coordinate of a neighboring point of the hand key point in the hand image;

[0157] A first determining unit, configured to determine a target hand key point from adjacent points corresponding to each of the plurality of hand key points;

[0158] The second determining unit is used to determine the third coordinate corresponding to the target hand key point as the second coordinate.

[0159] In some embodiments, the first feature information includes feature information corresponding to the neighboring points. Based on this, the acquisition unit may specifically include:

[0160] A first determining subunit is used to determine, for each hand key point among the multiple hand key points, a neighboring point of the hand key point among the multiple hand key points;

[0161] A second determining subunit is used to determine feature information corresponding to the neighboring point in the first feature information;

[0162] The conversion subunit is used to convert the feature information corresponding to the neighboring points into the third coordinates of the neighboring points in the hand image.

[0163] In some of the embodiments, the key point detection device 500 may further include:

[0164] The second acquisition module is used to obtain the hand sample image, the label coordinates corresponding to each hand key point in the hand sample image, and the palm skeleton information, where the label coordinates are obtained by manual annotation;

[0165] A third extraction module is used to extract the features of the hand key points of the hand sample image using the initial feature extraction layer in the initial key point detection model to obtain the first predicted feature information of each hand key point;

[0166] A fourth extraction module is used to perform feature extraction on the palm skeleton information and the first predicted feature information by using the initial graph convolutional neural network layer in the initial key point detection model to obtain second predicted feature information of each hand key point, where the second predicted feature information includes the palm skeleton information;

[0167] A second conversion module, used for converting the second predicted feature information of each hand key point into the predicted coordinates of each hand key point in the hand sample image;

[0168] A determination module, for determining a loss function value according to a difference between a predicted coordinate and a label coordinate of each hand key point;

[0169] The adjustment module is used to adjust the model parameters of the initial feature extraction layer and the initial graph convolutional neural network layer according to the loss function value, and train the key point detection model.

[0170] The key point detection device of the embodiment of the present application performs feature extraction on the palm skeleton information and the first feature information by using the graph convolutional neural network layer in the key point detection model to obtain the second feature information, and converts the second feature information into the target coordinates of the hand key points in the hand image. That is, compared with key point detection based only on the hand image, by adding palm skeleton information in the process of using the key point detection model to perform hand key point detection, the key point detection model can be constrained to be optimized in a direction consistent with the human hand skeleton, so that when the hand key points in the hand image are not obvious or are blocked, the key point detection model can also reasonably determine the relative positions between multiple hand key points. In this way, when the hand in the hand image is small and not obvious, and / or the hand in the hand image is blocked, it is possible to reasonably and accurately perform hand key point detection, obtain target coordinates with higher accuracy, and improve the accuracy of hand key point detection.

[0171] Based on the key point detection method provided in the above embodiment, the embodiment of the present application also provides a specific implementation of the electronic device. Figure 6 A schematic diagram of an electronic device 600 provided in an embodiment of the present application is shown.

[0172] The electronic device 600 may include a processor 610 and a memory 620 storing computer program instructions.

[0173] Specifically, the processor 610 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0174] The memory 620 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 620 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 620 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 620 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 620 is a non-volatile solid-state memory.

[0175] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of the present application.

[0176] The processor 610 implements any key point detection method in the above embodiments by reading and executing computer program instructions stored in the memory 620 .

[0177] In one example, the electronic device 600 may further include a communication interface 630 and a bus 640. Figure 6 As shown, the processor 610, the memory 620, and the communication interface 630 are connected via a bus 640 and communicate with each other.

[0178] The communication interface 630 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0179] Bus 640 includes hardware, software or both, and the parts of electronic equipment are coupled to each other. For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 640 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.

[0180] Exemplarily, the electronic device 600 may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA).

[0181] The electronic device can execute the key point detection method in the embodiment of the present application, thereby realizing the combination Figures 1 to 5 Described is a key point detection method and apparatus.

[0182] In addition, in combination with the key point detection method in the above embodiment, the embodiment of the present application can provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any key point detection method in the above embodiment is implemented.

[0183] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.

[0184] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0185] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.

[0186] The above reference is according to the method of the embodiment of the present application, the flow chart of the device (system) and the computer program product and / or the block diagram described various aspects of the present application.It should be understood that each square box in the flow chart and / or the block diagram and the combination of each square box in the flow chart and / or the block diagram can be realized by computer program instructions.These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the realization of the function / action specified in one or more square boxes of the flow chart and / or the block diagram.Such a processor can be but is not limited to a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit.It can also be understood that each square box in the block diagram and / or the flow chart and the combination of the square boxes in the block diagram and / or the flow chart can also be realized by the dedicated hardware that performs the specified function or action, or can be realized by the combination of dedicated hardware and computer instructions.

[0187] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.

Claims

1. A key point detection method, characterized in that: include: Get hand image and palm skeleton information; Extracting features of the hand key points from the hand image using a feature extraction layer in a key point detection model to obtain first feature information, where the first feature information is feature information of the hand key points; Using the graph convolutional neural network layer in the key point detection model to perform feature extraction on the palm skeleton information and the first feature information to obtain second feature information, where the second feature information is feature information of the hand key points including the palm skeleton information; The second feature information is converted into target coordinates of the hand key points in the hand image.

2. The method according to claim 1, characterized in that The feature extraction layer includes a hand feature extraction layer and a nested network PipNet layer. The feature extraction layer in the key point detection model is used to extract the features of the hand key points of the hand image to obtain the first feature information, including: Using the hand feature extraction layer to extract hand features from the hand image, to obtain third feature information, where the third feature information is hand feature information; The PipNet layer is used to extract the feature information of the hand key points from the third feature information to obtain the first feature information.

3. The method according to claim 1, characterized in that The graph convolutional neural network layer includes a plurality of convolutional layers, and the graph convolutional neural network layer in the key point detection model is used to extract features from the palm skeleton information and the first feature information to obtain the second feature information, including: Using a first convolutional layer among the multiple convolutional layers to perform feature extraction on the palm skeleton information and the first feature information to obtain fourth feature information; Using a target convolution layer to perform feature extraction on the palm skeleton information, the first feature information and the first target feature information to obtain second target feature information, the target convolution layer is the i-th convolution layer among the multiple convolution layers, the first target feature information is the feature information output by the (i-1)-th convolution layer, and the first target feature information includes the fourth feature information, wherein i is a positive integer greater than or equal to 2; Let i=i+1, and repeat the step of obtaining the second target feature information until the target convolution layer is the last convolution layer among the multiple convolution layers, and determine the second target feature information obtained most recently as the second feature information.

4. The method according to claim 1, characterized in that The converting the second feature information into target coordinates of the hand key point in the hand image includes: For a target hand key point, converting the second feature information into a first coordinate of the target hand key point in the hand image, wherein the target hand key point is any one of a plurality of hand key points; Acquire a second coordinate of the target hand key point in the hand image, where the second coordinate is a coordinate of the target hand key point when it is a neighboring point of other hand key points, where the other hand key points are hand key points other than the target hand key point among the multiple hand key points; An average value of the first coordinate and the second coordinate is calculated to obtain the target coordinate of the target hand key point in the hand image.

5. The method according to claim 4, characterized in that The obtaining of the second coordinate of the target hand key point in the hand image includes: For each hand key point among the multiple hand key points, obtaining a third coordinate of a neighboring point of the hand key point in the hand image; Determine the target hand key point among the neighboring points corresponding to each of the plurality of hand key points; The third coordinate corresponding to the target hand key point is determined as the second coordinate.

6. The method according to claim 5, characterized in that The first feature information includes feature information corresponding to the neighboring point, and the acquiring, for each of the plurality of hand key points, a third coordinate of a neighboring point of the hand key point in the hand image includes: For each hand key point in the plurality of hand key points, determining a neighboring point of the hand key point in the plurality of hand key points; Determining feature information corresponding to the neighboring point in the first feature information; The feature information corresponding to the neighboring point is converted into the third coordinate of the neighboring point in the hand image.

7. The method according to claim 1, characterized in that Before extracting the features of the hand key points from the hand image using the feature extraction layer in the key point detection model, the method further includes: Acquire a hand sample image, label coordinates corresponding to each hand key point in the hand sample image, and the palm skeleton information, wherein the label coordinates are obtained by manual annotation; Extracting features of hand key points from the hand sample image using the initial feature extraction layer in the initial key point detection model to obtain first predicted feature information of each hand key point; Using the initial graph convolutional neural network layer in the initial key point detection model to perform feature extraction on the palm skeleton information and the first predicted feature information to obtain second predicted feature information of each of the hand key points, wherein the second predicted feature information includes the palm skeleton information; Converting the second predicted feature information of each of the hand key points into predicted coordinates of each of the hand key points in the hand sample image; Determine a loss function value according to a difference between the predicted coordinates and the label coordinates of each of the hand key points; The model parameters of the initial feature extraction layer and the initial graph convolutional neural network layer are adjusted according to the loss function value, and the key point detection model is obtained by training.

8. A key point detection device, characterized in that: The device comprises: A first acquisition module is used to acquire a hand image and palm skeleton information; A first extraction module, configured to extract features of hand key points from the hand image using a feature extraction layer in a key point detection model to obtain first feature information, where the first feature information is feature information of the hand key points; A second extraction module is used to perform feature extraction on the palm skeleton information and the first feature information by using the graph convolutional neural network layer in the key point detection model to obtain second feature information, where the second feature information is feature information of the hand key points including the palm skeleton information; The first conversion module is used to convert the second feature information into the target coordinates of the hand key point in the hand image.

9. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the key point detection method as described in any one of claims 1-7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the key point detection method according to any one of claims 1 to 7 is implemented.