Gesture recognition method, gesture recognition model training method, and related device

By detecting key points and topological structures of target parts in pose recognition technology and combining this with pose recognition model training, the problem of low pose classification accuracy is solved, achieving higher pose recognition accuracy and faster recognition speed.

CN117095053BActive Publication Date: 2026-04-17HANGZHOU HUACHENG SOFTWARE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU HUACHENG SOFTWARE TECH CO LTD
Filing Date
2023-07-19
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing vision-based pose recognition technologies suffer from low accuracy in pose classification due to the high similarity between poses.

Method used

By acquiring the target part in the image to be identified, key point detection is performed. The pose category is determined by using the position coordinates and topology map of multiple key points. The pose recognition model is then trained and the parameters are optimized to improve the accuracy of pose recognition.

Benefits of technology

It improves the accuracy of pose recognition, reduces the inference time required for pose recognition, and reduces the dependence on the bounding box of the target part.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117095053B_ABST
    Figure CN117095053B_ABST
Patent Text Reader

Abstract

This application discloses a pose recognition method, a training method for a pose recognition model, and related equipment. The method includes: acquiring an image to be recognized, wherein the image contains a target region; performing keypoint detection on the image to be recognized to obtain the position coordinates of multiple keypoints of the target region; and using the position coordinates of the multiple keypoints and a topological graph representing the connections between the multiple keypoints to determine the pose category of the target region in the image to be recognized; wherein the topological graph represents the connection relationships between the multiple keypoints. This approach can improve the accuracy of pose recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a pose recognition method, a pose recognition model training method, a computer device, and a computer-readable storage medium. Background Technology

[0002] With the rapid development of artificial intelligence technology, its application in the field of human-computer interaction is becoming increasingly widespread. In various fields such as intelligent robots, games, simulation, and manufacturing, gesture recognition technology is being applied and developed as a means of communication between individuals and computers.

[0003] Currently, vision-based pose recognition technology refers to using computer vision technology to process images or image sequences containing poses to achieve pose recognition. However, this type of method is affected by the high similarity between poses in the image, resulting in low pose classification accuracy. Summary of the Invention

[0004] The main technical problem addressed in this application is to provide a pose recognition method, a pose recognition model training method, and related equipment, which can improve the accuracy of pose recognition.

[0005] To address the aforementioned issues, the first aspect of this application provides a pose recognition method, comprising: acquiring an image to be recognized, wherein the image to be recognized contains a target part; performing keypoint detection on the image to be recognized to obtain the position coordinates of multiple keypoints of the target part; and determining the pose category to which the target part in the image to be recognized belongs by using the position coordinates of the multiple keypoints and a topological structure graph between the multiple keypoints; wherein the topological structure graph represents the connection relationship between the multiple keypoints.

[0006] To address the aforementioned issues, a second aspect of this application provides a method for training a pose recognition model. This method includes: acquiring a sample recognition image, wherein the sample recognition image contains a target part; processing the sample recognition image using the pose recognition model to implement the pose recognition method described above, thereby obtaining a sample loss function; and adjusting the parameters of the pose recognition model based on the sample loss function to obtain a trained pose recognition model.

[0007] To address the aforementioned problems, a third aspect of this application provides a computer device comprising a memory and a processor coupled to each other. The memory stores program data, and the processor executes the program data to implement any step of any of the aforementioned pose recognition method or pose recognition model training method.

[0008] To address the aforementioned problems, a fourth aspect of this application provides a computer-readable storage medium storing program data executable by a processor, the program data being used to implement any step of any of the above-described pose recognition method and pose recognition model training method.

[0009] The above scheme acquires an image to be recognized, which contains a target part; performs keypoint detection on the image to be recognized to obtain the position coordinates of multiple keypoints of the target part; and uses the position coordinates of multiple keypoints and the topological structure diagram between multiple keypoints to determine the pose category of the target part in the image to be recognized. The topological structure diagram represents the connection relationship between multiple keypoints. Keypoints can not only provide the position information of each keypoint, but also obtain the topological structure diagram between keypoints by arranging the connection order relationship between keypoints. Combining the topological relationship between keypoints for pose recognition can improve the accuracy of pose recognition. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them:

[0011] Figure 1 This is a flowchart illustrating an embodiment of the posture recognition method of this application;

[0012] Figure 2 This application Figure 1 A flowchart illustrating an embodiment of step S12;

[0013] Figure 3 This is a schematic diagram of the structure of an embodiment of the posture recognition model of this application;

[0014] Figure 4 This application Figure 2 A flowchart illustrating an embodiment of step S122;

[0015] Figure 5 This is a flowchart illustrating an embodiment of the feature sampling process of the feature sampler in this application;

[0016] Figure 6 This is the key point and topology of this application. Figure 1 Example schematic diagram of the embodiment;

[0017] Figure 7 This is an example schematic diagram of an embodiment of the image to be identified in this application;

[0018] Figure 8 This is a flowchart illustrating an embodiment of the training method for the pose recognition model of this application;

[0019] Figure 9 This is the sample identification of this application. Figure 1 Example schematic diagram of the embodiment;

[0020] Figure 10 This is an example schematic diagram of another embodiment of the sample recognition diagram in this application;

[0021] Figure 11 This is a schematic diagram of the structure of an embodiment of the computer device of this application;

[0022] Figure 12 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0024] The terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0025] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0026] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0027] This application provides the following embodiments, and each embodiment is described in detail below.

[0028] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the posture recognition method of this application. The method may include the following steps:

[0029] S11: Obtain the image to be identified, wherein the image to be identified contains the target part.

[0030] Images or videos (containing multiple image sequences) can be captured using camera equipment, and the captured images or image sequences can be used as images to be identified, that is, images that need to be used for pose recognition.

[0031] In some embodiments, the image to be identified includes a target part, such as a hand, fingers, arm, leg, the entire body of a person or animal, or an object, etc., and this application does not impose any limitations on this. The following description uses a hand as an example of a target part.

[0032] In some implementations, target region detection can be performed on the acquired image, thereby using the area where the target region is located as the image to be identified.

[0033] S12: Perform key point detection on the image to be recognized to obtain the position coordinates of multiple key points of the target area.

[0034] The system can detect the position coordinates of key points in a target area of ​​an image. Taking the hand as an example, the number of key points can be 21 or 16. For example, 21 key points can be distributed across four joints of each finger, the fingertip, and one point in the palm. Alternatively, 16 key points can be distributed across three joints of each finger and one point in the palm. Taking the entire human body as an example, multiple key points can be distributed across the head, shoulders, hips, hands, legs, etc. In this application, the key points and their positions can be determined according to the specific application scenario. This application does not limit the number of key points obtained or their positions within the target area.

[0035] In some embodiments, please refer to Figure 2 This embodiment can further extend step S12 of the above embodiment. To perform key point detection on the image to be recognized and obtain the position coordinates of multiple key points of the target area, this embodiment may include the following steps:

[0036] S121: Extract features from the image to be recognized to obtain image recognition features.

[0037] Feature extraction networks can be used to extract features from the image to be recognized, so as to obtain image recognition features.

[0038] In some implementations, please refer to Figure 3 Pose recognition can be performed using a pose recognition model. This model can include a feature extraction module, a keypoint detection module, and a topology recognition module. As an example, a pose recognition model can be based on the one-stage object detection method CenterNet. This method allows the image to obtain the position and category of a gesture in the image by passing the image through the network only once.

[0039] The feature extraction module can be the backbone network or it can be composed of a deep aggregation network, such as DLA-34. The image to be recognized is... By processing the input into the backbone network, image recognition features can be obtained. Where H0 and W0 are the height and width of the input image to be recognized, respectively, and H, W, and C are the height, width, and number of channels of the feature map, respectively. The image recognition features have a multiple relationship with the size of the image to be recognized, as shown by the sampling multiple. As an example, H = H0 / 4, W = W0 / 4. The backbone network in this step can also be a network structure with other feature extraction capabilities, which can be flexibly adjusted according to actual needs. This application does not impose any restrictions on this. This application does not restrict the backbone network, and the size and complexity of the feature extraction module can be adjusted according to the actual application scenario, so that the feature extraction of this application has good scalability.

[0040] S122: Utilize image recognition features to obtain the total offset corresponding to multiple key points; where the total offset represents the offset of the key point relative to the grid points of the image to be recognized.

[0041] The key point detection module can be used to process image recognition features to obtain the total offset corresponding to multiple key points.

[0042] In some implementations, a first offset and a second offset can be obtained first, and then the total offset can be obtained based on the first offset and the second offset. The first offset represents the offset between the group representation coordinates corresponding to the group and the grid points of the image to be identified (e.g., the target representation coordinates of the target part), and the second offset represents the offset between the group representation coordinates and the key points. Grid points can refer to pixels in the image to be identified, or each grid point divided according to a set size; this application does not limit this.

[0043] In some embodiments, please refer to Figure 4 This embodiment can be further extended to step S122 of the above embodiment. Utilizing image recognition features to obtain the total offset corresponding to multiple key points, this embodiment may include the following steps:

[0044] S1221: Using image recognition features, obtain the first offset of the first quantity group.

[0045] Here, the first offset represents the offset between the group representation coordinates corresponding to the group and the grid points of the image to be recognized, and the group representation coordinates represent the coordinates of the grid points contained in the group. For each grid point in the image to be recognized, the first offset of a first number of groups can be obtained respectively.

[0046] In some implementations, when the grid points of the image to be identified are the target representation coordinates of the target part, the first offset represents the offset between the group representation coordinates corresponding to the group and the target representation coordinates of the target part. The target representation coordinates are obtained using the coordinates of the grid points contained in the area where the target part is located. Target representation coordinates can be, for example, the center position of the hand area, the center position of the gesture, or the center position of the grid. As an example, the area or position of the hand in the image can be detected first to obtain the center position of the gesture. This application is not limited to this.

[0047] In some implementations, group representation coordinates represent the coordinates of the grid points contained within a group. Group representation coordinates can represent the coordinate center of the grid points contained in the group; for example, the center position of the grid points within the region where each keypoint is located can be used as the group representation coordinates. As an example, the number of groups can be determined according to the number of keypoints; for example, 21 keypoints can be divided into 21 groups, and the center of each group can be used as the 21 group representation coordinates. The center position of each group can be the average of the horizontal and vertical coordinates of the grid points in that group. Alternatively, group representation coordinates can be determined in other ways, and this application does not limit this method.

[0048] In some implementations, grid points may contain keypoints, and group representation coordinates represent the coordinates of the keypoints contained in a group. The center of the coordinates of the keypoints contained in a group can be used as the group representation coordinates. For example, the group representation coordinates can be the average of the horizontal and vertical coordinates of the keypoints in a group. The group center position of a group of keypoints can be learned through the network training process. Taking the hand as an example, the first number of groups can be divided into two groups: one for each finger's keypoints and another for the palm's keypoints; or the keypoints of each finger can be divided into one group, and the keypoints of the palm can be divided into a separate group. This application does not limit the method of dividing the first number of groups.

[0049] Please see Figure 3 The keypoint detection module includes a first convolutional layer, a partitioning convolutional layer, a feature sampler, and a second convolutional layer. The first convolutional layer is connected to both the partitioning convolutional layer and the feature sampler; the partitioning convolutional layer is connected to the feature sampler; and the feature sampler is connected to the second convolutional layer. The output of the first convolutional layer serves as the input to both the partitioning convolutional layer and the feature sampler. The outputs of the partitioning convolutional layer and the first convolutional layer serve as the input to the feature sampler, and the output of the feature sampler serves as the input to the second convolutional layer.

[0050] By processing image recognition features using the first convolutional layer, image convolutional features can be obtained. For example, the first convolutional layer can be a 3×3 Conv layer with a 3×3 kernel and a stride of 1. By processing image recognition features F using the first convolutional layer, the image recognition features F from the backbone network can be integrated and learned. After activation processing by an activation layer ReLU, the image convolutional features can be obtained, which can be represented as convolved image recognition features. In some application scenarios, the following embodiments use this as the image recognition feature.

[0051] The image convolution features are processed by dividing the convolutional layers to obtain the first offset of the first number of groups, wherein the dividing convolutional layers include convolutional layers with the number of convolutional kernels being a preset multiple of the first number.

[0052] For example, if the first quantity is n, where n is a natural number, and the preset multiplier is 2, then the number of convolutional kernels that can be included in the convolutional layer is (n×2). The convolutional layer can include a 3×3 Conv convolutional layer with a 3×3 kernel and a stride of 1, and a 1×1 Conv convolutional layer with a 1×1 kernel, a stride of 1, and (n×2) kernels. Using the 3×3 Conv convolutional layer to convolve the image convolutional features, more explicit keypoint features can be learned. Then, convolution is performed using the 1×1 Conv convolutional layer with (n×2) kernels to obtain the first offset of the first quantity n groups.

[0053] In Offset1, the (n×2) values ​​at position (i,j) represent the first offset values ​​of the n group centers relative to the current position (i,j) in the x and y directions, where (i,j) represents the grid point in the i-th row and j-th column. Adding the coordinates of the current position (i,j) to the n first offset values ​​yields the coordinates of each group center of the gesture at position (i,j), which are also known as the group representation coordinates.

[0054] S1222: Using image recognition features and the first offset of the first quantity group, obtain the second offset of the second quantity group, where the second offset represents the offset between the group characterization coordinates and the key point.

[0055] Please see Figure 3 A feature sampler can be used to sample the first offset of the first set of data based on image recognition features, thereby obtaining the group representation features corresponding to each grid point. When the group representation coordinates are needed to express the position of the group center, the group representation features can also be called group center features.

[0056] Using image recognition feature F1, feature sampling processing is performed on the first offset of the first set of features to obtain the first set of sampled features corresponding to the grid points of the image to be recognized; wherein, each sampled feature corresponds to a first offset, and each grid point can correspond to the first set of sampled features. By concatenating the first set of sampled features corresponding to each grid point, the set of representation coordinates corresponding to each grid point can be obtained.

[0057] As an example, the first offset can be used. We can obtain the positions of the n group centers corresponding to each grid point (i,j), which are the group representation coordinates. Then, we use bilinear interpolation to sample the features of the image recognition feature map F1 at these n group centers. By concatenating them according to the grouping type, we can obtain the group representation features, which are the group center features.

[0058] Please see Figure 5 Taking the first offset of the three groups as an example, the feature sampling process of the feature sampler is explained below.

[0059] The first offset (Offset1) can contain three sets of first offsets, each containing offsets in both the x and y directions. Similarly, the position (i,j) of a grid point within Offset1 can contain three sets of first offsets, each containing offsets in both the x and y directions. Using these three sets of first offsets and the grid point's position (i,j), the positions of the three group centers in image recognition feature F1 can be obtained. Then, bilinear interpolation is used to sample the positions of these three group centers in image recognition feature F1, resulting in three sampled features. These three sampled features are then concatenated according to grouping type or order, yielding the group center features (group representation features).

[0060] Then, after obtaining the group representation features F2, the group representation features can be processed using a second convolutional layer to obtain the second offset of the second number of groups. The number of convolutional kernels in the second convolutional layer is either the number of keypoints contained in the group or the total number of keypoints. The second offset can be the same as the first offset, or it can be different from the first offset.

[0061] Please see Figure 3 As an example, the second convolutional layer can be a 1×1 Conv layer with a kernel size of 1×1, a stride of 1, and m kernels, where m is a positive integer. By convolving the group representation features F2 using the second convolutional layer, we can obtain the second offset of each group's m keypoints relative to the group representation coordinates (or the position of the group center). Taking 21 keypoints as an example, we can obtain the second offset...

[0062] S1223: Add the first offset of the first quantity group and the second offset of the second quantity group to obtain the total offset corresponding to multiple key points.

[0063] Adding the second offset at the current position (i,j) of the grid point to the first offset of the corresponding keypoint group yields the total offset of the keypoint relative to the current position (i,j) of the grid point. Thus, the total offset for each keypoint can be obtained.

[0064] In some implementations, the above process of this embodiment can obtain a first offset, a second offset, and a total offset for each grid point of the image to be identified.

[0065] In some implementations, when the current position (i,j) of a grid point is the target representation coordinate (e.g., the center position of a gesture), the total offset is the total offset of the keypoint relative to the target representation coordinate. Please see Figure 3 The total offset in this case can be output, which is output ③.

[0066] S123: Using the total offset corresponding to multiple key points and the coordinates of the grid points in the image to be identified, the position coordinates of multiple key points of the target part are obtained.

[0067] For each grid point, the total offset of multiple key points corresponding to the current position (i,j) of the grid point can be obtained. By adding the coordinates of the corresponding grid point to its current position (i,j), the coordinates of the 21 key points can be obtained. Please see Figure 3 This can output the position coordinates of the 21 key points, which is also known as output ②.

[0068] In this embodiment, the first offset represents the offset between the target representation coordinates and the group representation coordinates, and the second offset represents the offset between the group representation coordinates and the key points within the group. Taking the target part, the hand, as an example, the target representation coordinates can be the center position of the gesture, and the group representation coordinates can be the center position of the group. The key point detection module realizes a coarse-to-fine detection process for 21 key points of the gesture. First, a coarse group center position of the key point group relative to the center position of the gesture is obtained through convolutional layers. Then, the features of each group center are extracted and further learned using convolutional layers to obtain the offset of each key point in each group relative to the group center. Finally, accurate key point detection is achieved through additive processing. Features near the group center have a stronger correlation with the key points, and the key point position obtained using the group center will be more accurate. The grouping method of the key points can be flexibly adjusted according to the actual situation.

[0069] Furthermore, the keypoint detection module first focuses on the central region of each group of keypoints, and then finds the precise location of each keypoint within the group by searching for the vicinity of the central region. This coarse-to-fine learning approach is more conducive to keypoint detection and reduces the inaccuracy caused by a single offset from the current position to the keypoint. Moreover, this method does not require precise supervision of the initial offset; the coarse center position of each group of keypoints can be adaptively adjusted during the continuous learning of the network.

[0070] S13: Using the position coordinates of multiple key points and the topological structure diagram between multiple key points, determine the pose category of the target part in the image to be identified; wherein, the topological structure diagram represents the connection relationship between multiple key points.

[0071] In some implementations, the topology graph represents the connections between multiple key points.

[0072] In some implementations, the topology graph represents the connections between key points and the adjacency relationships between key points and other key points. For example, an adjacency matrix A can be used to represent the topology graph.

[0073] In some implementations, a topology recognition module can be used to comprehensively determine the pose category of the target part in the image to be recognized by using the position coordinates of multiple key points and the topological structure diagram A between the multiple key points.

[0074] In this embodiment, an image to be identified is acquired, which contains a target part; key point detection is performed on the image to be identified to obtain the position coordinates of multiple key points of the target part; using the position coordinates of multiple key points and the topological structure diagram between multiple key points, the pose category to which the target part in the image to be identified belongs is determined; wherein, the topological structure diagram represents the connection relationship between multiple key points, and key points can not only provide the position information of each key point, but also obtain the topological structure diagram between key points by arranging the connection order relationship between key points. Combining the topological relationship between key points for pose recognition can improve the accuracy of pose recognition.

[0075] Please see Figure 6 For example, if the target part is the hand, an adjacency matrix can be constructed based on the connection relationships of multiple key points. The first channel in matrix A represents the self-connection relationships of the 21 key points, forming a diagonal matrix. The second and third channels in A represent the adjacency relationships of the key points connected in ascending order of their numbers (0-20), and descending order of their numbers, respectively; both can be upper triangular or lower triangular matrices. The construction method of the topology diagram in this application can be selected according to the specific application scenario, and more complex or different adjacency matrices A can be constructed in different ways according to actual application requirements.

[0076] A graph convolutional network can be constructed using the topological graph A between multiple key points. See also... Figure 3 The topology recognition module may include a graph convolutional network (GCN), which can consist of three cascaded graph convolutional layers. These layers can have the same structure as those used in the spatiotemporal graph convolutional network (ST-GCN), or they may differ in feature dimensions. For example, the last layer of the GCN could be a softmax layer. The number of graph convolutional layers and the structure of the GCN can be customized according to the specific application requirements; this application does not impose any limitations on these aspects.

[0077] The pose recognition method based on topological relationships allows the model to learn the relationships between multiple key points of a target area, which is more reasonable than focusing only on the topological relationships of local key points. Taking the hand as an example, a gesture is a semantic expression composed of the position and shape of the palm and fingers. The key point relationships of a specific gesture are specific. Graph convolutional networks jointly learn the global topological relationships between key points of the palm and fingers to parse the meaning of the gesture. The features obtained in this way can recognize gestures more concisely and effectively than those obtained through image regions.

[0078] The coordinates of multiple key points can be preprocessed before being input into a graph convolutional network. The graph convolutional network then processes the coordinates of the key points to obtain the pose category Cls∈ of the target region in the image to be recognized. cls×H×W And target representation coordinates. Please refer to... Figure 3 It can output the pose category and target representation coordinates, which can be output as ①.

[0079] As an example, please refer to Figure 7 Following the above method, based on the positional coordinates of multiple key points and the topological structure diagram between these key points, the pose category of the target part (hand) in the image to be identified can be determined as "like". In the above scheme, the hand key point information not only provides the positional information of each hand joint, but also, by arranging the joint information according to the bone connection order, the topological structure between key points can be obtained. Applying the topological relationship between hand key points to gesture recognition helps improve the accuracy of gesture recognition. A model based on a one-stage object detection framework can be used to realize the category recognition of gestures in images and the detection of hand key points, and the topological relationship between gesture key points can be used to enhance the model's performance in gesture recognition.

[0080] Furthermore, the pose recognition model can obtain the position, category, and multiple keypoint locations of the pose in the image in just one pass, reducing the inference time required for pose recognition. Moreover, the pose recognition result is not limited to the bounding box of the target region. Secondly, in multi-stage methods, if the keypoint-based recognition model is trained with labeled data, the keypoint prediction results input to the model during testing will differ significantly from the labeled data, leading to a decrease in overall recognition performance. If the keypoint prediction results are used for training, complex post-processing is required before inputting them into the recognition model. Therefore, compared to multi-stage methods, the pose recognition method proposed in this application is more convenient and effective.

[0081] In some embodiments, before using the pose recognition model for pose recognition, the pose recognition model can be trained to obtain a trained pose recognition model. The pose recognition model may include a feature extraction module, a keypoint detection module, and a topology recognition module.

[0082] Please see Figure 8 , Figure 8 This is a flowchart illustrating an embodiment of the training method for the pose recognition model of this application. The method may include the following steps:

[0083] S21: Obtain a sample identification image, wherein the sample identification image contains the target area.

[0084] Images can be acquired as sample recognition maps, which contain target parts. The corresponding data labels for the sample recognition maps include labels such as pose category and key point coordinates. This application does not impose any limitations on this.

[0085] Please see Figure 9 Taking the hand as the target area as an example, a dataset can be obtained. For example, the dataset is the public dataset HaGRID, from which sample recognition images can be obtained. The sample recognition images can be taken by a mobile phone or camera at a distance of 0.5 meters to 4 meters. The data labels contain the type of gesture, bounding box, and 2D coordinates of 21 key points (such as x and y axis coordinates).

[0086] Please see Figure 10 Taking the human body as the target part as an example, a dataset containing the human body can be obtained to obtain a sample recognition image. The sample recognition image contains the entire human body, and the data labels contain the human posture category, bounding box, coordinates of 17 key points, etc.

[0087] S22: The steps of implementing the above pose recognition method using the pose recognition model are to process the sample recognition image to obtain the sample loss function.

[0088] Please see Figure 3 The loss function can be obtained separately for the outputs ①, ②, and ③ of the above pose recognition model, and then the three loss functions can be added together as the sample loss function. Alternatively, the sample loss function can be obtained by comprehensively obtaining the outputs ①, ②, and ③. This application does not impose any restrictions on this approach.

[0089] For example, Focal Loss is used for output ①, OKS loss function for keypoint detection is used for output ②, and L1 loss function for keypoint detection is used for output ③. The above three loss functions are added together as the total sample loss function for network training, so as to continuously optimize the network during the training process.

[0090] S23: Based on the sample loss function, adjust the parameters of the pose recognition model to obtain a trained pose recognition model.

[0091] Using the sample loss function described above, the parameters of the pose recognition model can be adjusted. Depending on the accuracy requirements of different application scenarios, the parameters of the pose recognition model can be adjusted. For example, the parameters of any one of the feature extraction module, key point detection module, and topology recognition module can be adjusted to obtain a trained pose recognition model.

[0092] In some implementations, the image to be tested can be input into a trained network, and post-processing can be used to obtain the pose category and keypoint detection results. The post-processing may include: using a Non-Maximum Suppression (NMS) operation to filter out overlapping detection results in output ①, obtaining the number of poses, their categories, and center point coordinates in the image. The NMS operation can be the same as that used in CenterNet. The coordinates of the 21 keypoints corresponding to output ② are obtained at the same locations where poses exist. Then, the center coordinates and keypoint coordinates of the poses are scaled back to obtain their coordinate values ​​on the original image. At this point, the image to be tested has completed the entire pose recognition process. It should be noted that the output ③ of the pose recognition model is only used during training to calculate the sample loss function to optimize network weights; this output result may not be generated during testing.

[0093] By directly integrating the topological relationships of key points into the detection network of the pose recognition module, taking the hand as an example, the adjacency matrix A represents the topological structure relationship of 21 key points of the gesture. During training, the graph convolutional network continuously adjusts and aggregates the features between related key points, thereby achieving a more accurate judgment of the pose category.

[0094] Regarding the above embodiments, this application provides a computer device; please refer to [link / reference]. Figure 11 , Figure 11 This is a schematic diagram of the structure of a computer device according to an embodiment of the present application. The computer device 30 includes a memory 31 and a processor 32, wherein the memory 31 and the processor 32 are coupled to each other. The memory 31 stores program data, and the processor 32 is used to execute the program data to implement the steps of any embodiment of any of the above-described pose recognition method and pose recognition model training method.

[0095] In this embodiment, processor 32 can also be referred to as CPU (Central Processing Unit). Processor 32 may be an integrated circuit chip with signal processing capabilities. Processor 32 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor can be a microprocessor, or processor 32 can be any conventional processor.

[0096] The methods described in the above embodiments can be implemented as computer programs; therefore, this application proposes a computer-readable storage medium. Please refer to [link to relevant documentation]. Figure 12 , Figure 12 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 40 stores program data 41 that can be executed by a processor. The program data 41 can be executed by the processor to implement the steps of any embodiment of the above-described pose recognition method and pose recognition model training method.

[0097] In this embodiment, the computer-readable storage medium 40 can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or a medium that can store program data 41. Alternatively, it can be a server that stores the program data 41. The server can send the stored program data 41 to other devices for execution, or it can run the stored program data 41 itself.

[0098] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0099] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0100] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0101] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application.

[0102] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, and thus stored in a computer-readable storage medium for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Therefore, this application is not limited to any particular hardware and software combination.

[0103] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A pose recognition method, characterized in that, The method includes: Acquire an image to be identified, wherein the image to be identified contains a target region; Key point detection is performed on the image to be identified to obtain the position coordinates of multiple key points of the target part; By using the position coordinates of the multiple key points and the topological structure diagram between the multiple key points, the pose category of the target part in the image to be identified is determined; wherein, the topological structure diagram represents the connection relationship between the multiple key points; The step of performing key point detection on the image to be identified to obtain the position coordinates of multiple key points of the target region includes: Feature extraction is performed on the image to be identified to obtain image recognition features; Using the image recognition features, a first offset of a first quantity group is obtained; wherein, the first offset represents the offset between the group representation coordinates corresponding to the group and the grid points of the image to be recognized, and the group representation coordinates represent the coordinates of the grid points contained in the group. Using the image recognition features and the first offset of the first quantity group, a second offset of the second quantity group is obtained, wherein the second offset represents the offset between the group characterization coordinates and the key point; The first offset of the first quantity group and the second offset of the second quantity group are added together to obtain the total offset corresponding to the multiple key points; wherein, the total offset represents the offset of the key point relative to the grid points of the image to be identified; By using the total offset corresponding to multiple key points and the coordinates of grid points in the image to be identified, the position coordinates of multiple key points of the target part are obtained.

2. The method according to claim 1, characterized in that, The step of obtaining the first offset of the first quantity group using the image recognition features includes: The image recognition features are processed using the first convolutional layer to obtain image convolutional features; The image convolutional features are processed by dividing the convolutional layers to obtain the first offset of the first number group, wherein the dividing convolutional layers include convolutional layers with the number of convolutional kernels being a preset multiple of the first number.

3. The method according to claim 1, characterized in that, The step of obtaining the second offset of the second quantity group using the image recognition features and the first offset of the first quantity group includes: The first offset of the first quantity group is sampled using the image recognition features to obtain the group characterization features; The group representation features are processed using a second convolutional layer to obtain a second offset for a second number of groups, wherein the number of convolutional kernels in the second convolutional layer is equal to the number of keypoints contained in the group.

4. The method according to claim 3, characterized in that, The step of sampling the first offset of the first quantity group using the image recognition features to obtain group characterization features includes: Using the image recognition features, feature sampling processing is performed on the first offset of the first quantity group to obtain a first quantity of sampled features corresponding to the grid points of the image to be recognized; wherein, each sampled feature corresponds to the first offset; The first number of sampled features corresponding to the grid points are concatenated to obtain the group of characterization features.

5. The method according to claim 1, characterized in that, The step of determining the pose category of the target part in the image to be identified by utilizing the position coordinates of the multiple key points and the topological structure diagram between the multiple key points includes: A graph convolutional network is constructed using the topological structure graph between the multiple key points; wherein, the topological structure graph represents the connection relationship between key points and themselves, and the adjacency relationship between key points and other key points. The graph convolutional network is used to process the position coordinates of the multiple key points to obtain the pose category of the target part in the image to be identified.

6. A method for training a pose recognition model, characterized in that, The method includes: Obtain a sample identification image, wherein the sample identification image contains a target region; The sample recognition map is processed using a pose recognition model to achieve the steps of the method described in any one of claims 1 to 5, thereby obtaining a sample loss function; Based on the sample loss function, the parameters of the pose recognition model are adjusted to obtain a trained pose recognition model.

7. A computer device, characterized in that, The method includes a memory and a processor coupled to each other, the memory storing program data and the processor executing the program data to implement the steps of the method according to any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that, The system stores program data that can be executed by a processor, the program data being used to implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for detecting interaction between 3D hand and unknown object in RGB video in real time

    CN112199994A