Pose Recognition Method, Device, Storage Medium and Electronic Device
By using the central trend vector field to characterize the correlation between non-central key points and central key points, the problem of low accuracy in key point matching in multi-person pose recognition is solved, and efficient and accurate pose recognition is achieved.
Patent Information
- Application Number
- CN202210473734.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The existing multi-person pose recognition technology takes too much time or has complex processing logic, making it difficult to ensure the correct matching of key points.
A central trend vector field that characterizes the correlation between non-center key points and central key points is used to determine the object to be identified to the target central key point, and the posture of the object to be identified is determined based on the non-center key points and the central key points.
While ensuring speed advantages, the accuracy of key points matching is improved, thereby improving the accuracy of posture recognition and reducing the dependence between non-central key points.
Smart Images

Figure CN114782994B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a posture recognition method, apparatus, storage medium, and electronic device. Background Art
[0002] Currently, human posture recognition, as an important reference basis for behavior monitoring, is widely used in fields such as intelligent teaching, smart home, autonomous driving, and intelligent monitoring. For example, in an intelligent teaching scenario, posture correction, posture answering, etc. can be performed through human posture recognition.
[0003] In the related art, for multi-person posture recognition, the positions of all people can be detected through a human detector, and then the human key points of each person can be recognized separately. It is also possible to directly recognize the key points and then match the recognized key points to the corresponding human body. After that, the postures of each human body are determined through the recognized key points. However, the former method depends on the human detector and takes a long time when there are too many people. The latter method is fast, but the processing logic is complex, and it is difficult to correctly match the key points to the corresponding human body when there are too many people. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a posture recognition method, apparatus, storage medium, and electronic device to solve the problems in the related art.
[0005] To achieve the above purpose, according to the first aspect of the embodiments of the present disclosure, a posture recognition method is provided, and the method includes:
[0006] Obtain an image to be recognized, where the image to be recognized includes multiple objects to be recognized;
[0007] Input the image to be recognized into a trained key point recognition model to obtain the central key points and non-central key points of each object to be recognized, as well as a central tendency vector field, where the central tendency vector field represents the association degree between each non-central key point and each central key point;
[0008] For a target central key point, determine a target non-central key point that belongs to the same target object to be recognized as the target central key point according to the central tendency vector field, where the target central key point is any one of the central key points;
[0009] Determine the posture of the target object to be recognized according to the target non-central key point and the target central key point.
[0010] Optionally, the central tendency vector field includes multiple central tendency vector sub-fields, different central tendency vector sub-fields correspond to different non-central key point sets, and different non-central key point sets correspond to different key point categories;
[0011] Determining, for the target central key point, target non-central key points that belong to the same target object to be recognized as the target central key point according to the central tendency vector field includes:
[0012] For the target central key point, determining, from each of the non-central key point sets, a non-central key point with the largest matching degree value with the target central key point, to obtain the target non-central key points of multiple categories.
[0013] Optionally, the method further includes:
[0014] Determining the matching degree between the non-central key point and the target central key point through the following formula:
[0015]
[0016] where M is the value of the matching degree, CTV x is the central tendency vector from any non-central key point (x, y) in the non-central key point set to the target central key point (x c , y c ), K is a preset value, and (x j , y j ) are the coordinates of a preset key point on the connection between the target central key point and the non-central key point, and the preset key point is obtained by dividing the connection into K segments.
[0017] Optionally, after inputting the image to be recognized into the trained key point recognition model, the key point recognition model is used to:
[0018] Perform key point recognition on the image to be recognized to obtain a central key point heat map and a non-central key point heat map;
[0019] Determine the central key points from the central key point heat map and determine the non-central key points from the non-central key point heat map.
[0020] Optionally, the training process of the key point recognition model includes:
[0021] Obtaining an image sample to be recognized, where the image sample to be recognized corresponds to a central key point label, a non-central key point label, and a central tendency vector field label. Among them, the image sample to be recognized includes multiple object samples to be recognized, and each object sample to be recognized corresponds to a sample central key point labeled with a central key point label and multiple sample non-central key points labeled with non-central key point labels;
[0022] Train a trained key point recognition model based on the to-be-recognized image samples.
[0023] Optionally, the sample center key point corresponding to the to-be-recognized object sample is the mean value of multiple sample non-center key points corresponding to the to-be-recognized object sample.
[0024] Optionally, the central tendency vector field label represents the correlation degree between each sample non-center key point and each sample center key point. When the sample non-center key point and the sample center key point are on the same to-be-recognized object sample, the correlation degree is calculated by the following formula:
[0025]
[0026] where CTV x ′ is the value of the correlation degree, (x c ′, y c ′) is the coordinate of the sample center key point, and (x′, y′) is the coordinate of the sample non-center key point;
[0027] In the case where the sample non-center key point and the sample center key point correspond to different to-be-recognized object samples, set the value of the correlation degree to 0.
[0028] According to a second aspect of the embodiments of the present disclosure, there is provided a pose recognition device, the device includes:
[0029] An acquisition module, configured to acquire a to-be-recognized image, where the to-be-recognized image includes multiple to-be-recognized objects;
[0030] An input module, configured to input the to-be-recognized image into the trained key point recognition model, and obtain the center key points and non-center key points of each to-be-recognized object, and a central tendency vector field, where the central tendency vector field represents the correlation degree between each non-center key point and each center key point;
[0031] A first determination module, configured to, for a target center key point, determine a target non-center key point that belongs to the same target to-be-recognized object as the target center key point according to the central tendency vector field, where the target center key point is any one of the center key points;
[0032] A second determination module, configured to determine the pose of the target to-be-recognized object according to the target non-center key point and the target center key point.
[0033] According to a third aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any item of the first aspect are implemented.
[0034] According to a fourth aspect of the embodiments of the present disclosure, an electronic device is provided, including:
[0035] A memory on which a computer program is stored;
[0036] A processor configured to execute the computer program in the memory to implement the steps of the method according to any one of the above first aspects.
[0037] Through the above technical solution, by using the central tendency vector field representing the correlation degree between non-central key points and central key points, for a target central key point, target non-central key points belonging to the same target object to be recognized can be determined, so that the posture of the target object to be recognized can be determined according to the target non-central key points and the target central key point. In this process, by introducing the central key point and the central tendency vector field representing the correlation degree between non-central key points and central key points to perform the matching between key points and the object to be recognized, the accuracy of key point matching can be improved while ensuring the speed advantage, thereby improving the accuracy of posture recognition.
[0038] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification. Together with the following specific implementation, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure. In the drawings:
[0040] Figure 1 is a flowchart of a posture recognition method shown according to an exemplary embodiment of the present disclosure.
[0041] Figure 2 is a schematic diagram of a key point recognition model shown according to an exemplary embodiment of the present disclosure.
[0042] Figure 3 is a block diagram of a posture recognition device shown according to an exemplary embodiment of the present disclosure.
[0043] Figure 4 is a block diagram of an electronic device shown according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The following will describe in detail the specific implementation of the present disclosure with reference to the drawings. It should be understood that the specific implementation described herein is only for explaining and understanding the present disclosure, and is not used to limit the present disclosure.
[0045] In related technologies, for multi-person pose recognition, the key points of each person can be determined by directly recognizing the key points and then matching the recognized key points to the corresponding human body, so as to determine the pose of each human body according to the recognized key points.
[0046] For example, in the paper "Associative Embedding: End-to-End Learning for Joint Detection and Grouping", the author proposed that while predicting the key points of each human body, the TAG value corresponding to each key point is also predicted, and the TAG values corresponding to the key points belonging to the same human body are as consistent as possible. In this way, the human body corresponding to each key point can be determined according to the TAG value corresponding to the key point. However, there is only supervision information at the key points in this way, and too little supervision information is not conducive to the learning of the intelligent agent. Moreover, there are significant differences between the key points at different positions on the same human body (such as the head key point and the foot key point). In this case, it is very difficult for the TAG values corresponding to the key points at different positions of the same human body predicted by the intelligent agent to be consistent. Especially when the number of people is too large, the problem of confusion of the TAGs corresponding to the human body key points is likely to occur, resulting in a low accuracy of matching the key points to the human body.
[0047] Also for example, in the solution described in the patent document CN111310625A, a skeleton composed of two key points for each human body is defined (such as the skeleton from the elbow key point to the wrist key point). While predicting the key points of each human body, the affinity of each skeleton connection is also predicted. The connection affinity of each skeleton is represented by a vector field (PAF). The unit vector of the vector field corresponding to any skeleton on the same human body is equal to the unit vector corresponding to the connection of the two key points on the skeleton. In this way, the skeleton to which the key point belongs can be determined according to the unit vector corresponding to the connection of the two key points, and thus the human body corresponding to the key point can be determined. However, this method has a high computational complexity, and there is a dependence between the key points: if there is a situation where a key point on a certain skeleton does not exist (such as the elbow key point connecting the shoulder key point or the wrist key point is not visible), then it is impossible to know whether the other visible key points (such as the shoulder key point and the wrist key point) connected to the invisible key point belong to the same human body.
[0048] In view of this, the embodiments of the present disclosure provide a pose recognition method. By introducing a central key point and a central tendency vector field representing the correlation between non-central key points and the central key point to perform the matching of the key points and the object to be recognized, it is possible to eliminate the dependence between non-central key points and improve the accuracy of key point matching while ensuring the speed advantage, thereby improving the accuracy of pose recognition.
[0049] Figure 1is a flowchart of a gesture recognition method shown according to an exemplary embodiment of the present disclosure. As Figure 1 shown, the method includes:
[0050] S101, obtaining an image to be recognized.
[0051] Among them, the image to be recognized includes a plurality of objects to be recognized. It should be understood that the object to be recognized can be a human body or an animal.
[0052] S102, inputting the image to be recognized into a trained key point recognition model to obtain the central key points and non-central key points of each object to be recognized, as well as the central tendency vector field.
[0053] Among them, the central tendency vector field characterizes the correlation degree between each non-central key point and each central key point.
[0054] It can be understood that the key points to be recognized on the object to be recognized can be preset, and the number of key points to be recognized on each object to be recognized is the same. When the object to be recognized is a human body, the non-central key points can be the key points corresponding to each part of the human body. For example, the nose key point, the right shoulder key point, and the left knee key point, etc. When the object to be recognized is an animal, the non-central key points can be the key points corresponding to each part of the animal. For example, the eye key point, the throat key point, and the paw key point, etc. On this basis, the central key point of the object to be recognized can be the average value of the non-central key points of each part of the object to be recognized. The central tendency vector field can be a set of central tendency vectors from each non-central key point to each central key point in each object to be recognized.
[0055] S103, for a target central key point, determining a target non-central key point belonging to the same target object to be recognized as the target central key point according to the central tendency vector field.
[0056] Among them, the target central key point is any central key point.
[0057] Since the central tendency vector can be used to characterize the correlation degree between each non-central key point and each central key point. Therefore, for a target central key point, the correlation degree between the target central key point and each non-central key point can be determined according to the central tendency vector field, and based on this correlation degree, a target non-central key point belonging to the same target object to be recognized as the target central key point can be determined.
[0058] S104, determining the gesture of the target object to be recognized according to the target non-central key point and the target central key point.
[0059] It is not difficult to understand that after determining the target non - central key points and the target central key point belonging to the same target object to be recognized, the posture of the target object to be recognized corresponding to the target central key point can be determined according to the target non - central key points.
[0060] In addition, it should be noted that in the related art, the method of predicting multiple bone connections and matching key points with bones has a high computational complexity. In the case of the absence of a certain key point, it will also cause the key points belonging to the same bone as the missing key point to be unable to be classified, that is, there is a dependency between key points. However, the method provided by the embodiments of the present disclosure introduces a central key point and calculates the central tendency vector field of all non - central key points and the central key point, so that the target non - central key points and the target central key point belonging to the same target object to be recognized can be determined according to the central tendency vector field. In this process, there is no dependency between non - central key points. Even if a certain non - central key point is missing, it does not affect the attribution of other key points.
[0061] Through the above technical solution, by using the central tendency vector field representing the correlation degree between the non - central key point and the central key point, for the target central key point, the target non - central key points belonging to the same target object to be recognized can be determined, so that the posture of the target object to be recognized can be determined according to the target non - central key points and the target central key point. In this process, by introducing the central key point and the central tendency vector field representing the correlation degree between the non - central key point and the central key point to match the key points with the object to be recognized, the accuracy of key point matching can be improved while ensuring the speed advantage, thereby improving the accuracy of posture recognition.
[0062] Optionally, the central tendency vector field includes multiple central tendency vector sub - fields. Different central tendency vector sub - fields correspond to different non - central key point sets, and different non - central key point sets correspond to different key point categories. On this basis, step S103 above may include:
[0063] For the target central key point, determine a non - central key point with the largest matching degree value with the target central key point from each non - central key point set, and obtain target non - central key points of multiple categories.
[0064] It should be noted that different key point categories can correspond to different parts of the object to be recognized. For example, the key point category of the right shoulder of a human corresponds to the right shoulder of the human, and the key point category of a cat's paw corresponds to the cat's paw. The non - central key point set corresponding to the key point category includes all non - central key points belonging to the key point category in each object to be recognized. For example, for a to - be - recognized image including 3 humans (each human includes 1 right - shoulder key point), the non - central key point set corresponding to the right - shoulder key point category includes 3 right - shoulder key points.
[0065] On this basis, for the target central key point, according to multiple central tendency vector sub-fields, a non-central key point with the largest matching degree with the target central key point can be determined from each non-central key point set, so as to obtain target non-central key points of multiple categories.
[0066] For example, for the left shoulder key point set, according to the central tendency vector sub-field corresponding to the left shoulder key point, the matching degree between each left shoulder key point in the left shoulder key point set and the target central key point can be determined, and the left shoulder key point with the largest matching degree among them is used as the target left shoulder key point. On this basis, by traversing each non-central key point set, target non-central key points of multiple categories can be obtained.
[0067] Optionally, the method provided in the embodiments of the present disclosure can determine the matching degree between the non-central key point and the target central key point through the following formula:
[0068]
[0069] where M is the value of the matching degree, CTV x is the central tendency vector from any non-central key point (x, y) in the non-central key point set to the target central key point (x c , y c ), K is a preset value, and (x j , y j ) are the coordinates of the preset key point on the connection between the target central key point and the non-central key point, and the preset key point is obtained by dividing the connection into K segments.
[0070] For example, if the coordinates of the target central key point are (2, 4) and the coordinates of the non-central key point are (6, 6), then when K = 2, the coordinates of the preset key points on the connection between the target central key point and the non-central key point may include (2, 4), (4, 5), and (6, 6).
[0071] Optionally, after inputting the image to be recognized into the trained key point recognition model, the key point recognition model can be used to:
[0072] Perform key point recognition on the image to be recognized to obtain a central key point heat map and a non-central key point heat map;
[0073] Determine the central key point from the central key point heat map and determine the non-central key point from the non-central key point heat map.
[0074] Among them, a heat map is a statistical chart that displays data by coloring color blocks (i.e., pixel points). In the center key point heat map, a region is predicted for each center key point of each object to be recognized in the image to be recognized. For each color block within the region corresponding to any center key point, the larger the probability that the color block is the center key point, the larger the value of the color block. On this basis, a pixel point with the largest color block value can be determined from the region corresponding to any center key point, and this pixel point is determined as the center key point. Similarly, non-center key points can be determined from the non-center key point heat map.
[0075] Referring to Figure 2 , Figure 2 is a schematic diagram of a key point recognition model shown by the present disclosure according to an exemplary embodiment. As Figure 2 shown, after inputting the image to be recognized (this image can be an RGB three-channel image) into the key point recognition model 400, the features of the image to be recognized can be extracted through the feature extraction network to obtain a key point heat map (including a center key point heat map and a non-center key point heat map) and a center tendency vector field. Among them, the number of channels of the key point heat map is the same as the number of key points to be recognized on the object to be recognized, and the key points to be recognized on the object to be recognized and the number thereof can both be preset. Since the center tendency vector can be represented by a two-channel feature map (the two channels represent the components in the x direction and the y direction respectively), the number of channels of the center tendency vector field is twice the number of non-center key points to be recognized on the object to be recognized.
[0076] On this basis, the matching degree between each non-center key point and the center key points of each object to be recognized can be determined through the matching module, so as to determine the non-center key points (i.e., the key point detection results) that belong to the same object to be recognized as each center key point. After that, the pose of the object to be recognized can be determined according to the key point detection results.
[0077] Optionally, the training process of the key point recognition model may include:
[0078] Obtain an image sample to be recognized, and train a trained key point recognition model according to the image sample to be recognized.
[0079] Among them, the image sample to be recognized corresponds to a center key point label, a non-center key point label, and a center tendency vector field label, and the image sample to be recognized includes multiple object samples to be recognized, and each object sample to be recognized corresponds to a sample center key point marked with a center key point label and multiple sample non-center key points marked with non-center key point labels.
[0080] It should be noted that the key point recognition model can be trained using image samples to be recognized that are labeled with true labels (i.e., central key point labels, non - central key point labels, and central tendency vector field labels). During the training process, the key point recognition model to be trained makes predictions on the image samples to be recognized, obtaining predicted central key points, predicted non - central key points, and a predicted central tendency vector field. Based on this, the loss value can be calculated according to the prediction results and the true labels, and the training parameters of the key point recognition model to be trained can be adjusted through the loss value, thereby obtaining a trained key point recognition model.
[0081] Optionally, the sample central key point corresponding to the object sample to be recognized is the mean of the multiple sample non - central key points corresponding to the object sample to be recognized. In a possible implementation, the mean can be calculated using the following formula:
[0082]
[0083] where (x c ′, y c ′) are the coordinates of the sample central key point, N is the number of key points on the same object to be recognized, and (x i , y i ) are the coordinates of the i - th sample non - central key point on the same object to be recognized.
[0084] Optionally, the central tendency vector field label represents the degree of association between each sample non - central key point and each sample central key point. When the sample non - central key point and the sample central key point are on the same object sample to be recognized, the degree of association is calculated using the following formula:
[0085]
[0086] where CTV x ′ is the sample central tendency vector from the sample non - central key point to the sample central key point (i.e., the value of the degree of association), used to represent the degree of association between the two, (x c ′, y c ′) are the coordinates of the sample central key point, and (x′, y′) are the coordinates of the sample non - central key point;
[0087] In the case where the sample non - central key point and the sample central key point correspond to different object samples to be recognized, the value of the degree of association is set to 0.
[0088] With the above technical solution, by using the central tendency vector field that characterizes the correlation between non-central key points and central key points, for a target central key point, the target non-central key points belonging to the same target object to be recognized can be determined, so that the pose of the target object to be recognized can be determined according to the target non-central key points and the target central key point. In this process, by introducing the central key points and the central tendency vector field that characterizes the correlation between non-central key points and central key points to perform the matching between key points and the object to be recognized, the accuracy of key point matching can be improved while ensuring the speed advantage, thereby improving the accuracy of pose recognition.
[0089] Based on the same inventive concept, the present disclosure also provides a pose recognition device. Refer to Figure 3 , Figure 3 which is a block diagram of a pose recognition device shown according to an exemplary embodiment of the present disclosure. As Figure 3 shown, the pose recognition device 100 includes:
[0090] An acquisition module 101, configured to acquire an image to be recognized, where the image to be recognized includes a plurality of objects to be recognized;
[0091] An input module 102, configured to input the image to be recognized into a trained key point recognition model to obtain the central key points and non-central key points of each object to be recognized, and a central tendency vector field, where the central tendency vector field characterizes the correlation between each non-central key point and each central key point;
[0092] A first determination module 103, configured to, for a target central key point, determine, according to the central tendency vector field, the target non-central key points that belong to the same target object to be recognized as the target central key point, where the target central key point is any one of the central key points;
[0093] A second determination module 104, configured to determine the pose of the target object to be recognized according to the target non-central key points and the target central key point.
[0094] With the above device, by using the central tendency vector field that characterizes the correlation between non-central key points and central key points, for a target central key point, the target non-central key points belonging to the same target object to be recognized can be determined, so that the pose of the target object to be recognized can be determined according to the target non-central key points and the target central key point. In this process, by introducing the central key points and the central tendency vector field that characterizes the correlation between non-central key points and central key points to perform the matching between key points and the object to be recognized, the accuracy of key point matching can be improved while ensuring the speed advantage, thereby improving the accuracy of pose recognition.
[0095] Optionally, the center tendency vector field includes a plurality of center tendency vector sub-fields. Different center tendency vector sub-fields correspond to different non-center key point sets, and different non-center key point sets correspond to different key point categories;
[0096] The first determination module 103 is further configured to:
[0097] For the target center key point, determine a non-center key point with the largest matching degree value with the target center key point from each non-center key point set, and obtain the target non-center key points of multiple categories.
[0098] Optionally, the apparatus 100 further includes a third determination module, and the third determination module is configured to determine the matching degree between the non-center key point and the target center key point through the following formula:
[0099]
[0100] where M is the value of the matching degree, CTV x is the center tendency vector from any non-center key point (x, y) in the non-center key point set to the target center key point (x c , y c ), K is a preset value, and (x j , y j ) are the coordinates of a preset key point on the connection between the target center key point and the non-center key point. The preset key point is obtained by dividing the connection into K segments.
[0101] Optionally, after inputting the image to be recognized into the trained key point recognition model, the key point recognition model is used to:
[0102] Perform key point recognition on the image to be recognized to obtain a center key point heat map and a non-center key point heat map;
[0103] Determine the center key points from the center key point heat map and determine the non-center key points from the non-center key point heat map.
[0104] Optionally, the apparatus 100 further includes a training module, and the training module is configured to train the key point recognition model. The training process of the key point model includes:
[0105] Obtain an image sample to be recognized, where the image sample to be recognized corresponds to a central key point label, a non - central key point label, and a central tendency vector field label. Among them, the image sample to be recognized includes multiple object samples to be recognized, and each object sample to be recognized corresponds to a sample central key point labeled with a central key point label and multiple sample non - central key points labeled with non - central key point labels;
[0106] Train a trained key point recognition model based on the image sample to be recognized.
[0107] Optionally, the sample central key point corresponding to the object sample to be recognized is the mean value of the multiple sample non - central key points corresponding to the object sample to be recognized.
[0108] Optionally, the central tendency vector field label represents the correlation degree between each sample non - central key point and each sample central key point. The device 100 further includes a calculation module, and this calculation module is used for:
[0109] When the sample non - central key point and the sample central key point are on the same object sample to be recognized, calculate the correlation degree through the following formula:
[0110]
[0111] where, CTV x ′ is the value of the correlation degree, (x c ′, y c ′) is the coordinate of the sample central key point, and (x′, y′) is the coordinate of the sample non - central key point;
[0112] When the sample non - central key point and the sample central key point correspond to different object samples to be recognized, set the value of the correlation degree to 0.
[0113] Regarding the device in the above - mentioned embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here in detail.
[0114] Based on the same inventive concept, an embodiment of the present disclosure further provides an electronic device, and this electronic device includes:
[0115] A memory, on which a computer program is stored;
[0116] A processor, configured to execute the computer program in the memory to implement the steps of the above - mentioned posture recognition method.
[0117] Figure 4 is a block diagram of an electronic device 700 shown according to an exemplary embodiment. As Figure 4As shown, the electronic device 200 may include: a processor 201 and a memory 202. The electronic device 200 may also include one or more of a multimedia component 203, an input / output (I / O) interface 204, and a communication component 205.
[0118] Among them, the processor 201 is used to control the overall operation of the electronic device 200 to complete all or part of the steps in the above-mentioned gesture recognition method. The memory 202 is used to store various types of data to support the operation of the electronic device 200. These data may include, for example, instructions for any application or method operating on the electronic device 200, as well as application-related data, such as contact data, received and sent messages, pictures, audio, video, and so on. The memory 202 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The multimedia component 203 may include a screen and an audio component. The screen may be a touch screen, for example, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 202 or sent through the communication component 205. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 204 provides an interface between the processor 201 and other interface modules, and the above other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 205 is used for wired or wireless communication between the electronic device 200 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, 4G, or 5G, NB-IoT (Narrow Band Internet of Things), or a combination of one or more of them. Accordingly, the communication component 205 may include: a Wi-Fi module, a Bluetooth module, and an NFC module.
[0119] In an exemplary embodiment, the electronic device 200 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above-mentioned gesture recognition method.
[0120] Based on the same inventive concept, an embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above-mentioned gesture recognition method are implemented. For example, the non-transitory computer-readable storage medium may be the above-mentioned memory 202 including program instructions, and the above-mentioned program instructions may be executed by the processor 201 of the electronic device 200 to complete the above-mentioned gesture recognition method.
[0121] Specifically, the computer-readable storage medium may be a flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, a server, a public cloud server, and so on.
[0122] Regarding the non-transitory computer-readable storage medium in the above embodiment, the implementation of the steps of the gesture recognition method when the computer program stored thereon is executed has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0123] In another exemplary embodiment, a computer program product is further provided, and the computer program product includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above-mentioned gesture recognition method when executed by the programmable device.
[0124] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all belong to the protection scope of the present disclosure.
[0125] In addition, it should be noted that, among the various specific technical features described in the above specific embodiments, they can be combined in any suitable manner without conflict.
[0126] Furthermore, any combinations can be made among the various different embodiments of the present disclosure, as long as they do not violate the idea of the present disclosure, and they should equally be regarded as the content disclosed by the present disclosure.
Claims
1. A gesture recognition method, characterized in that, The method includes: Obtaining an image to be recognized, where the image to be recognized includes multiple objects to be recognized; Inputting the image to be recognized into a trained key point recognition model to obtain the central key points, non-central key points, and central tendency vector field of each object to be recognized, where the central tendency vector field represents the correlation degree between each non-central key point and each central key point; For a target central key point, determining a target non-central key point that belongs to the same target object to be recognized as the target central key point according to the central tendency vector field, where the target central key point is any one of the central key points; Determining the pose of the target object to be recognized according to the target non-central key point and the target central key point; The central tendency vector field includes multiple central tendency vector sub-fields, different central tendency vector sub-fields correspond to different non-central key point sets, and different non-central key point sets correspond to different key point categories; The step of, for a target central key point, determining a target non-central key point that belongs to the same target object to be recognized as the target central key point according to the central tendency vector field includes: For the target central key point, determining a non-central key point with the largest matching degree value with the target central key point from each non-central key point set to obtain target non-central key points of multiple categories; Determining the matching degree between the non-central key point and the target central key point through the following formula: Wherein, M is the value of the matching degree, is the central tendency vector from any non-central key point in the non-central key point set ( , ) to the target central key point ( , ), K is a preset value, ( , ) are the coordinates of a preset key point on the connection between the target central key point and the non-central key point, and the preset key point is obtained by dividing the connection into K segments.
2. The method according to claim 1, wherein After inputting the image to be recognized into the trained key point recognition model, the key point recognition model is used to: Perform key point recognition on the image to be recognized to obtain a central key point heat map and a non-central key point heat map; Determining the central key points from the central key point heat map and determining the non-central key points from the non-central key point heat map.
3. The method according to claim 1 or 2, characterized in that, The training process of the key point recognition model includes: Obtaining a sample image to be recognized, where the sample image to be recognized corresponds to a central key point label, a non-central key point label, and a central tendency vector field label. Among them, the sample image to be recognized includes multiple object samples to be recognized, and each object sample to be recognized corresponds to a sample central key point labeled with a central key point label and multiple sample non-central key points labeled with non-central key point labels; Training a trained key point recognition model according to the sample image to be recognized.
4. The method according to claim 3, characterized in that, The sample central key point corresponding to the object sample to be recognized is the mean value of the multiple sample non-central key points corresponding to the object sample to be recognized.
5. The method according to claim 4, wherein The central tendency vector field label represents the correlation degree between each sample non-central key point and each sample central key point. When the sample non-central key point and the sample central key point are on the same object sample to be recognized, the correlation degree is calculated through the following formula: Among them, is the value of the correlation degree, ( , ) are the coordinates of the central key point of the sample, ( , ) are the coordinates of the non-central key point of the sample; In the case where the sample non-central key point and the sample central key point correspond to different object samples to be recognized, the value of the correlation degree is set to 0.
6. An attitude recognition device, characterized in that, The device includes: An acquisition module for acquiring an image to be recognized, where the image to be recognized includes a plurality of objects to be recognized; An input module for inputting the image to be recognized into a trained key-point recognition model to obtain the central key points and non-central key points of each object to be recognized, and a central tendency vector field, where the central tendency vector field represents the correlation degree between each non-central key point and each central key point; A first determination module for determining, for a target central key point, a target non-central key point that belongs to the same target object to be recognized as the target central key point according to the central tendency vector field, where the target central key point is any one of the central key points; A second determination module for determining the pose of the target object to be recognized according to the target non-central key point and the target central key point; The central tendency vector field includes a plurality of central tendency vector sub-fields, different central tendency vector sub-fields correspond to different non-central key point sets, and different non-central key point sets correspond to different key point categories; The first determination module is further configured to: For the target central key point, determine, from each non-central key point set, a non-central key point with the largest matching degree value with the target central key point, to obtain the target non-central key points of multiple categories; A third determination module for determining the matching degree between the non-central key point and the target central key point through the following formula: where M is the value of the matching degree, is the central tendency vector from any non - central key point in the non - central key point set , to the target central key point , ), K is a preset value, and ( , ) are the coordinates of a preset key point on the connection between the target central key point and the non - central key point, and the preset key point is obtained by dividing the connection into K segments.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-5.
8. An electronic device, characterized in that, Including: A memory storing a computer program thereon; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Openpose-based multi-person posture detection method and system
CN111310625A
Key point detection method and device, electronic equipment and storage medium
CN111783882A
Posture recognition method and device, storage medium and electronic equipment
CN114139630A