Posture recognition method, device and electronic equipment
Patent Information
- Application Number
- CN202210791514.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2042-07-05
AI Technical Summary
[0009] The posture recognition method, apparatus, and electronic device provided in this disclosure, after receiving candidate posture information and key point information, further determine whether to respond to the posture indicated by the candidate posture information based on the key point information. That is, it determines whether to respond to the user's first posture based on the user's second part of key point information. Therefore, although the posture indicated by the candidate posture information is a predefined posture, the decision to respond to the first part of the posture is still made based on the second part of key point information, thereby avoiding erroneous responses to certain postures. In other words, this allows for a more accurate response to the user's first part of the posture.
Smart Images

Figure CN115188071B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of Internet technology, and in particular to a posture recognition method, apparatus and electronic device. Background Technology
[0002] With the development of science and technology, the functions of terminal applications are becoming increasingly sophisticated. For example, some terminal devices support gesture control. When a user performs a gesture, the terminal device may execute a corresponding command. For instance, a terminal device might be a smart music player; after a user performs a gesture, they can switch the currently playing song. Summary of the Invention
[0003] This disclosure is provided to briefly introduce the concepts, which will be described in detail in the subsequent Detailed Description section. This disclosure is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] This disclosure provides a posture recognition method, apparatus, and electronic device that can more accurately respond to the user's first part of the posture, that is, can identify the user's erroneous posture, thereby improving the accuracy of posture response.
[0005] In a first aspect, embodiments of this disclosure provide a posture recognition method, comprising: receiving candidate posture information and key point information, wherein the candidate posture information and the key point information correspond to the same user, the candidate posture information indicates that the posture of a first part is a predefined posture, and the key point information indicates key points of a second part of the user; and determining whether to respond to the posture indicated by the candidate posture information based on the key point information.
[0006] Secondly, embodiments of this disclosure provide a posture recognition device, comprising: a receiving unit, configured to receive candidate posture information and key point information, wherein the candidate posture information and the key point information correspond to the same user, the candidate posture information indicating a predefined posture of a first part, and the key point information indicating a key point of a second part of the user; and a determining unit, configured to determine, based on the key point information, whether to respond to the posture indicated by the candidate posture information.
[0007] Thirdly, embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the attitude recognition method as described in the first aspect.
[0008] Fourthly, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon that, when executed by a processor, implements the steps of the pose recognition method as described in the first aspect.
[0009] The posture recognition method, apparatus, and electronic device provided in this disclosure, after receiving candidate posture information and key point information, further determine whether to respond to the posture indicated by the candidate posture information based on the key point information. That is, it determines whether to respond to the user's first posture based on the user's second part of key point information. Therefore, although the posture indicated by the candidate posture information is a predefined posture, the decision to respond to the first part of the posture is still made based on the second part of key point information, thereby avoiding erroneous responses to certain postures. In other words, this allows for a more accurate response to the user's first part of the posture. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0011] Figure 1 This is a flowchart of an embodiment of the pose recognition method according to the present disclosure;
[0012] Figure 2 This is a partial user differentiation diagram according to an embodiment of the gesture recognition method of this disclosure;
[0013] Figure 3A and 3B These are schematic diagrams illustrating key points of another embodiment of the posture recognition method according to this disclosure;
[0014] Figure 4 This is a schematic diagram of a structure of an embodiment of the posture recognition device according to the present disclosure;
[0015] Figure 5 This is an example of a posture recognition method according to an embodiment of the present disclosure, which can be applied to an exemplary system architecture.
[0016] Figure 6 This is a schematic diagram of the basic structure of an electronic device provided according to an embodiment of the present disclosure. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0019] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0021] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0022] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0023] Please refer to Figure 1 This illustrates the flow of an embodiment of the pose recognition method according to the present disclosure. This pose recognition method can be applied to terminal devices (e.g., smart home appliances). Figure 1 The pose recognition method shown includes the following steps:
[0024] Step 101: Receive candidate pose information and key point information.
[0025] Here, candidate pose information and key point information both correspond to the same user. Candidate pose information indicates that the pose of the first part is a predefined pose, and key point information indicates the key points of the user's second part.
[0026] Here, a candidate pose information can correspond to a human body key point information.
[0027] As an example, the first part and the second part can be different. For instance, the first part could be the user's first hand, while the second part could be any other part of the user's body excluding the first hand. Of course, the specific parts that the first and second parts refer to can be determined based on the actual situation; here, we do not limit the specific parts of the first and second parts.
[0028] To better understand the first and second parts, you can combine... Figure 2 To understand, Figure 2 This can be understood as a diagram of the user's body parts, such as... Figure 2 As shown, the first part 201 and the second part 202 can be different parts of the user.
[0029] As an example, the executing entity can determine the form of the user's second body part by using the number and location of key points indicated by the key point information. For instance, if the second body part includes human limbs, the key point information can be used to determine whether the user is facing the executing entity, whether the user is sitting, or whether the user is standing, etc.
[0030] As an example, predefined poses can be defined according to the actual application scenario. For example, when the first part is the hand, when the user makes a palm gesture, a "V" sign gesture, or a fist gesture, it can be determined whether to define some or all of them as predefined poses based on the actual situation.
[0031] Step 102: Based on the key point information, determine whether to respond to the pose indicated by the candidate pose information.
[0032] As an example, the executing entity can determine the current form of the user's second part through key point information. Thus, based on the user's key point information, it can determine whether the user's first part made this gesture intentionally or unintentionally.
[0033] In some implementations, determining whether to respond to the posture indicated by the candidate posture information can also be understood as: determining whether the posture indicated by the candidate posture information is a posture that was accidentally triggered by the user.
[0034] For example, the executing entity can respond to a user's 'fist' gesture. If the user is facing the executing entity and makes the 'fist' gesture, it indicates that the user intends to control the executing entity. However, if the user is not facing the executing entity and makes the 'fist' gesture, it may be because the user is chatting with other users and makes this gesture, in which case it can be considered that the user does not intend to control the executing entity. In this case, it is obviously not necessary to respond to the gesture indicated by the candidate gesture information.
[0035] In related technologies, when receiving candidate posture information whose indicated posture is a predefined posture, one can directly respond to the posture indicated by that candidate posture information. However, this can lead to a large number of false responses.
[0036] As can be seen, in this disclosure, after receiving candidate pose information and key point information, it further determines whether to respond to the pose indicated by the candidate pose information based on the key point information. That is, it determines whether to respond to the user's first pose based on the user's second part of key point information. Therefore, although the pose indicated by the candidate pose information is a predefined pose, it still judges whether to respond to the first pose based on the second part of key point information, thus avoiding erroneous responses to certain poses. In other words, this allows for a more accurate response to the user's first part of pose.
[0037] In some embodiments, step 102 (determining whether to respond to the posture indicated by the candidate posture information based on key point information) may specifically include: determining the posture of the user's second part based on key point information; and determining whether to respond to the posture indicated by the candidate posture information based on the determined posture of the second part.
[0038] Here, the posture of the first part should correspond to the posture of the second part. For example, the first part is the user's first hand, and the second part is the user's other parts excluding the first hand. If the user's first hand is making a gesture towards the executing entity, the user's second part should also be facing the executing entity. If the verification shows that the posture of the user's second part is not facing the executing entity, obviously, the possible situation is that the user has mistakenly made a gesture with the first part. In this case, it can be determined that the posture indicated by the candidate posture information will not be responded to.
[0039] As can be seen, by using the posture of the second part, we can determine the user's current form. And by using the user's current form, we can determine the user's state when making the posture of the first part, such as whether the user is facing the executing entity. Thus, by determining whether the user is making a posture towards the executing entity, we can more accurately determine whether the user's current posture of the first part is directed towards the executing entity.
[0040] In some embodiments, determining the pose of the user's second part based on key point information may specifically include: determining the pose of the second part based on the number of key points indicated by the key point information and the positional relationship between the key points indicated by the key point information.
[0041] As an example, a keypoint pose mapping table can be pre-prepared, which can record the number of keypoints and the correspondence between the keypoint positions and poses.
[0042] Of course, in specific implementations, the specific method for determining the posture of the second part based on the number of key points and the positional relationship between the key points indicated by the key point information can be limited according to the actual situation.
[0043] As an example, this application scenario is a gesture recognition scenario. In this scenario, candidate pose information and key point information can both be obtained from video frame images. Therefore, the top left corner of the video frame can be used as the origin of the coordinate system, and the coordinates of each key point can be recorded in pixels. Then, the positional relationship of each key point can be determined based on its coordinates. Of course, in specific implementations, how to determine the positional relationship of key points can be limited according to the actual situation.
[0044] Here, by simultaneously utilizing the number of key points and their positional relationships, the posture of the second part can be determined more accurately. For example, the number of key points can determine whether the user's second part is facing the executing subject, while the positional relationships of the key points can further determine the current posture of the user's second part. This makes the posture judgment of the second part more accurate.
[0045] In some embodiments, the key point information includes at least one sub-key point information, which is used to indicate the key points of a sub-part. In this case, step 102 (determining whether to respond to the posture indicated by the candidate posture information based on the key point information) may specifically include: determining the verification method corresponding to each sub-key point information; verifying the sub-key point information according to the determined verification method; and determining whether to respond to the posture indicated by the candidate posture information based on the verification result of the sub-key point information.
[0046] As an example, key point information may include at least one sub-key point information. In this case, validating each sub-key point information can reduce the amount of data required for each validation.
[0047] Furthermore, since the verification results of sub-keypoint information can determine whether to respond to the posture indicated by the candidate posture information, the determination results can be made more accurate. For example, if a user's sub-part does not meet the verification conditions, it can be determined that the user will not respond to the posture indicated by the candidate posture information (that is, if a user's sub-part does not meet the verification conditions, it can be determined that the posture of the current user's first part is falsely triggered. The sub-part can be understood as the part indicated by the sub-keypoint information).
[0048] To better understand the relationship between the second part indicated by the key point information and the sub-part indicated by the sub-key point information, an example is given. For instance, the second part is any part other than the user's hand, while the sub-part can be understood as: eyes, head, shoulders, hips, etc.
[0049] In some embodiments, the above-mentioned verification of sub-key point information according to the determined verification method may specifically include: determining the verification order of each sub-key point information; verifying the sub-key point information according to the determined verification order; and determining whether to verify the next sub-part information based on the verification result of the previous sub-key point information.
[0050] In some implementation scenarios, if the verification result of the previous key information fails, then there is no need to verify the next key information. This improves the efficiency of determining whether to respond to the posture indicated by the candidate posture information, while also saving computational resources.
[0051] In some embodiments, the verification order for each sub-key point information can be determined based on the verification duration corresponding to each sub-key point information.
[0052] As an example, the second part may include multiple sub-parts, and the sub-parts indicated by each sub-keypoint information are different. This means that the verification time for each sub-keypoint information may be different. For example, factors such as the different number of keypoints in each sub-keypoint information or the different verification steps required for each sub-keypoint information will cause the verification time for each sub-keypoint information to be different.
[0053] As an example, the verification order could indicate that the sub-keypoint information with shorter verification times is verified first. This allows for a more efficient determination of whether to respond to the pose indicated by the candidate pose information. In other words, the verification order could indicate that each sub-keypoint information is verified sequentially according to its verification time, from shortest to longest.
[0054] In some embodiments, in response to determining that any sub-key point information verification fails, the pose indicated by the candidate pose information is determined not to respond.
[0055] In other words, if the verification of any sub-key point information fails, it can be determined that the posture indicated by the candidate posture information will not be responded to. In this way, not only can the amount of data to be processed be reduced, but processing efficiency can also be improved and computing resources can be saved.
[0056] For example, key point information may include three sub-key point information. After sorting the three sub-key point information, the sub-key point information ranked first can be verified first. If the sub-key point information ranked first fails to be verified, there is no need to verify the sub-key point information ranked second and third. Instead, it is directly determined that the posture indicated by the candidate posture information will not be responded to.
[0057] To illustrate this, consider the following example: Keypoint information includes sub-keypoint information A, sub-keypoint information B, and sub-keypoint information C. If verifying sub-keypoint information A takes 5 milliseconds, verifying sub-keypoint information B takes 15 milliseconds, and verifying sub-keypoint information C takes 20 milliseconds, then verifying sub-keypoint information A first will only require 5 milliseconds to determine whether to respond to the posture indicated by the candidate posture information. However, if sub-keypoint information B is verified first, followed by sub-keypoint information A, it may take 20 milliseconds to determine whether to respond to the posture indicated by the candidate posture information. Therefore, sorting and verifying the sub-keypoint information sequentially according to their verification time from shortest to longest allows for a more efficient determination of whether to respond to the posture indicated by the candidate posture information.
[0058] In some embodiments, key point information may include at least one of the following: limb key point information, eye key point information, hip joint key point information, and shoulder joint key point information.
[0059] As an example, the specific information included in key point information can be limited according to the actual situation, and there is no limit to the specific information included in human body key point information here.
[0060] In some embodiments, when the key point information includes limb key point information, eye key point information, hip joint key point information, and shoulder joint key point information, the determined verification order may be as follows: verify limb key point information, verify eye key point information, verify hip joint key point information, and verify shoulder joint key point information.
[0061] As an example, this verification method allows for a faster determination of whether to respond to the posture indicated by the candidate posture information. For instance, verifying limb keypoint information can determine whether the user is facing the executing subject, verifying eye keypoint information can determine whether the user is looking at the executing subject, and verifying hip joint keypoint information and shoulder joint keypoint information can further determine the user's posture, thereby more accurately determining whether the posture indicated by the candidate posture information is a falsely triggered posture.
[0062] In some embodiments, sub-key point information can be verified in the following ways:
[0063] The sub-keypoint information is validated using the pose of the sub-parts indicated by the sub-keypoint information and / or the confidence level of the sub-keypoint information.
[0064] As an example, the posture of a sub-part can reflect to some extent whether the posture performed by the user's first part is intentional or unintentional.
[0065] As an example, the confidence level of sub-keypoint information can characterize the degree of confidence of that sub-keypoint information, and the confidence level of sub-keypoint information is related to the background, lighting, etc. at the time of shooting. Therefore, based on the confidence level of sub-keypoint information, it can be determined whether the sub-keypoint information is reliable.
[0066] In other words, whether the pose of a sub-part does not meet the verification conditions or the information of a sub-key point does not meet the verification conditions, it can be determined that the information of the sub-key point does not meet the verification conditions (verification failed).
[0067] It should be noted that in actual application scenarios, whether to use the pose of the sub-part indicated by the sub-keypoint information to verify the sub-keypoint information, or to use the confidence level of the sub-keypoint information, or both, can be reasonably selected according to the actual situation. There is no limitation on the specific verification method here.
[0068] In some implementation scenarios, when simultaneously using the pose of the sub-part indicated by the sub-keypoint information to verify both the sub-keypoint information and its confidence level, the pose of the sub-part indicated by the sub-keypoint information can be verified first, followed by the confidence level of the sub-keypoint information. This way, if the pose of the sub-part indicated by the sub-keypoint information fails, there's no need to verify the confidence level of the sub-keypoint information, saving verification time. Furthermore, verifying the confidence level of sub-keypoint information is more complex and time-consuming than verifying the pose of the sub-part indicated by the sub-keypoint information; this verification method also saves computational resources and improves verification efficiency.
[0069] In some embodiments, the pose of the sub-part indicated by the sub-keypoint information can be determined in the following manner.
[0070] Based on the number of sub-keypoints indicated by the sub-keypoint information, the pose of the sub-part indicated by the sub-keypoint information is determined.
[0071] As an example, when a user performs a certain pose, the number of keypoints in the sub-parts indicated by the sub-keypoint information should be constant. For instance, when the user is facing the main body, the number of sub-keypoints on the user's torso should be 15. If, in this case, the number indicated by the sub-keypoint information is 10, it indicates that the user is not directly facing the main body. For easier understanding, this can be combined with... Figures 3A-3B To explain, Figures 3A-3B It is a key point distribution map of a certain part of the user's body. Figure 3A This can be understood as a key point distribution diagram of the part when the user faces the execution entity. Figure 3B It can characterize the distribution map of the currently detected key points, from Figure 3B visible, Figure 3B The key points in compared to Figure 3A It is clearly less than half, so at this time the user's front is not facing the executing entity, but the user's side is opposite the executing entity.
[0072] Here, by determining the pose of the sub-part indicated by the sub-keypoint information based on the number of sub-keypoints indicated by the sub-keypoint information, the pose of the sub-part can be determined more efficiently, and thus the verification of the sub-keypoint information can be achieved more efficiently.
[0073] To better understand the verification methods for various sub-keypoint information disclosed herein, the verification methods for limb keypoint information, eye keypoint information, hip joint keypoint information, and shoulder joint keypoint information are described below in sequence.
[0074] In some embodiments, the key point information may include limb key point information. In this case, the limb key point information can be verified in the following ways:
[0075] The limb keypoint information is verified based on the number of limb keypoints and / or the location of the limb keypoints indicated by the limb keypoint information.
[0076] As an example, if the number of limb key points does not match the preset number of limb key points, it indicates that the user's current posture may not be facing the executing entity. In this case, it can be determined that the limb key point information verification has failed.
[0077] As an example, if the position of the limb key point does not match the predefined position, it can also indicate that the user's current posture may not be facing the execution subject. In this case, it can be determined that the limb key point information verification has failed.
[0078] As an example, a mismatch between the positions of limb keypoints and their predefined positions can be understood as a mismatch between the distribution relationships between limb keypoints and the predefined distribution relationships between keypoints. For instance, three limb keypoints may be distributed in an acute triangle, while the predefined three keypoints may be distributed in an obtuse triangle.
[0079] As an example, a two-dimensional coordinate system can be established with the top-left corner of the first video frame image as the origin. Then, the pixels in the first video frame image can be used to represent the positions of limb keypoints. Of course, there are many other ways to represent the positions of limb keypoints in specific implementations, and this article does not limit the methods used to represent the positions of limb keypoints. It should be noted that the first video frame image here can be understood as the image used to obtain candidate pose information and keypoint information.
[0080] As can be seen, when the key point information includes limb key point information, the limb key point information can be verified by checking the number of limb key points and / or the position of limb key points.
[0081] Here, when the number of limb key points is used to determine whether the limb key point information is qualified, the amount of data that needs to be processed to verify the number of limb key points is small, so the verification result can be obtained quickly, that is, the verification result can be obtained efficiently.
[0082] Here, by using the location of limb key points to determine limb key point information for verification, the user's posture can be determined more accurately, thus making the verification results more accurate.
[0083] Here, when verifying limb keypoint information by simultaneously checking both the position and the number of limb keypoints, the number of limb keypoints can be checked first. Only if the number of keypoints meets the required condition can the position be checked. This not only improves verification efficiency but also makes the verification results more accurate.
[0084] In some embodiments, the limb key point information can also be verified by combining the confidence level of the limb key point information.
[0085] As an example, the confidence level of limb key point information can be determined based on the shadow ratio of the first video frame image and the background light intensity corresponding to the first video frame image.
[0086] In some embodiments, the human body key point information may further include eye key point information. In this case, the eye key point information can be verified in the following way:
[0087] The eye key point information can be verified based on the number of eye key points indicated by the eye key point information and / or the confidence level corresponding to the eye key point information.
[0088] As an example, if the number of eye key points indicated by the eye key point information does not match the preset number of eye key points, it can be determined that the eye key point information verification has failed. For example, if the number of eye key points indicated by the eye key point information is 5, with an error of 1, that is, when the number of eye key points indicated by the eye key point information is 4-6, the number of eye key points indicated by the eye key point information matches the preset number of eye key points. However, if the number of eye key points indicated by the eye key point information is 7, then the number of eye key points indicated by the eye key point information does not match the preset number of eye key points.
[0089] As an example, the confidence level corresponding to eye key point information can characterize the authenticity of eye key point information.
[0090] As an example, the confidence level of the eye keypoint information can be determined based on information such as the shadow intensity of the first video frame image and the amount of content included in the first video frame image. Of course, in specific implementations, the method for determining the confidence level of the eye keypoint information can be limited according to the actual situation.
[0091] As an example, when the confidence level of the eye keypoint information is lower than the preset confidence threshold for eye keypoints, it indicates that the eye keypoint information verification has failed. Of course, the preset confidence threshold for eye keypoints can be limited according to the actual situation.
[0092] In some implementation scenarios, when using the number of eye key points indicated by the eye key point information and the corresponding confidence level of the eye key point information to verify whether the eye key point information is qualified, the number of eye key points can be verified first and then the confidence level of the key point information can be verified; this can improve the efficiency of verification.
[0093] In some embodiments, the key point information may include hip joint key point information, and the hip joint key point information may be verified in the following manner:
[0094] The hip joint key point information is validated based on the number of hip joint key points indicated by the information and / or the confidence level corresponding to the hip joint key point information.
[0095] Similarly, the method for verifying hip joint key point information using the number of hip joint key points indicated by the hip joint key point information and / or the corresponding confidence level of the hip joint key point information is similar to the method for verifying eye key point information described above. For the sake of brevity, it will not be repeated here. This verification method can also reliably ensure the verification of hip joint key point information.
[0096] In some embodiments, the key point information may include hip joint key point information, and the hip joint key point information may be verified in the following manner:
[0097] Based on the number of hip joint key points indicated by the hip joint key point information, the hip joint key point information is first verified; in response to determining that the hip joint key point information meets the first verification condition, the hip joint key point information is second verified based on the confidence level of the hip joint key point information; in response to determining that the hip joint key point information meets the second verification condition, the difference coefficient between the first part and the hip joint is determined using the hip joint key point information and candidate posture information, and the hip joint key point information is third verified based on the difference coefficient between the first part and the hip joint.
[0098] As an example, performing three verifications on the hip joint key point information can make the verification results of the determined hip joint key point information more accurate, thereby making the judgment result of whether the posture indicated by the candidate posture information is a falsely triggered posture more accurate.
[0099] As an example, the difference coefficient between the first body part and the hip joint can also characterize the difference between the user's first body part and the hip joint. Normally, when a user performs a certain action, the difference coefficient between the first body part and the hip joint should be within a fluctuating range. If the difference coefficient between the first body part and the hip joint is not within this fluctuation range, it indicates that the user's current posture is incorrect.
[0100] Specifically, if the hip joint key point information meets the first verification condition, it indicates that the user's hip joint posture may meet the condition. If the hip joint key point information meets the second verification condition, it indicates that the authenticity of the hip joint key point information is relatively high. If the hip joint key point information meets the third verification condition, it indicates that the user's hip joint posture meets the condition at this time. Through three verifications, the accuracy of the verification can be guaranteed.
[0101] In some implementations, the number of hip joint key points indicated by the hip joint key point information can be verified first. This allows for the verification of simpler items first, followed by more complex items, thereby further improving verification efficiency.
[0102] In some implementations, the difference coefficient between the first site and the hip joint can be obtained in the following way:
[0103] A coordinate system is established with the top left corner of the first video frame as the origin. At this point, the vertical coordinate of the hip joint is set as follows: The vertical coordinate of the i-th key point in the first part is: Then there is
[0104] Coefficient of difference
[0105] The difference coefficient γ can be understood as the difference coefficient between the first part and the hip joint. When calculating the difference coefficient between the first part and the hip joint using the above formula, the calculated difference coefficient should be around 1.35, for example, it should be between 1.25 and 1.45. If the calculated difference coefficient γ exceeds this range, it can indicate that the third verification is unqualified.
[0106] In some embodiments, the key point information may include shoulder joint key point information, and the verification of the shoulder joint key point information is determined by the following method:
[0107] The shoulder joint key point information is verified based on the number of hip joint key points indicated by the shoulder joint key point information and / or the confidence level corresponding to the aforementioned shoulder joint key point information.
[0108] Similarly, the method for verifying hip joint key point information using the number of hip joint key points indicated by the hip joint key point information and / or the corresponding confidence level of the hip joint key point information is similar to the method for verifying eye key point information described above. For the sake of brevity, it will not be repeated here. This verification method can also effectively verify hip joint key point information.
[0109] In some embodiments, the key point information may include shoulder joint key point information, and the shoulder joint key point information may be verified in the following manner:
[0110] Based on the number of shoulder joint key points indicated by the shoulder joint key point information, the shoulder joint key point information is first verified; in response to determining that the shoulder joint key point information meets the first verification condition, the shoulder joint key point information is second verified based on the confidence level of the shoulder joint key point information; in response to determining that the shoulder joint key point information meets the second verification condition, the difference coefficient between the first part and the shoulder joint is determined using the shoulder joint key point information and candidate posture information, and the shoulder joint key point information is third verified based on the difference coefficient between the first part and the shoulder joint.
[0111] Here, performing three verifications on the shoulder joint key point information can make the verification results of the determined shoulder joint key point information more accurate, thereby making the judgment result of whether the posture indicated by the candidate posture information is a falsely triggered posture more accurate.
[0112] As an example, the difference coefficient between the first body part and the shoulder joint can also characterize the difference between the user's first body part and the shoulder joint. Normally, when a user performs a certain action, the difference coefficient between the first body part and the shoulder joint should be within a fluctuating range. If the difference coefficient between the first body part and the shoulder joint is not within this fluctuation range, it indicates that the user's current posture is incorrect.
[0113] Specifically, if the shoulder joint key point information meets the first verification condition, it indicates that the user's shoulder joint posture may meet the requirements. If the shoulder joint key point information meets the second verification condition, it indicates that the authenticity of the shoulder joint key point information is relatively high. If the shoulder joint key point information meets the third verification condition, it indicates that the user's shoulder joint posture is compliant with the requirements. Through three verifications, the accuracy of the verification can be guaranteed.
[0114] In some implementations, the number of shoulder joint key points indicated by the shoulder joint key point information is checked first. This allows for the verification of simpler verification items first, followed by more complex ones, thereby further improving verification efficiency.
[0115] In some implementations, the difference coefficient between the first part and the shoulder joint can be obtained in the following way:
[0116] A coordinate system is established with the top left corner of the first video frame as the origin. At this point, the vertical coordinate of the shoulder joint is set as follows: The horizontal coordinate of the i-th key point in the first part is: Then there is
[0117] Hand-shoulder difference coefficient
[0118]
[0119] Here, the difference coefficient β can be understood as the difference coefficient between the first part and the shoulder joint. When calculating the difference coefficient between the first part and the shoulder joint using the above formula, the calculated difference coefficient should be around -0.85, for example, it should be between -0.8 and -0.9. If the calculated difference coefficient β exceeds this range, it can indicate that the third verification is unqualified.
[0120] In some embodiments, when candidate pose information is used to indicate the hand pose of a user’s first hand, shoulder joint keypoint information includes information for indicating the shoulder keypoints of the user’s second hand.
[0121] In some implementations, the shoulder joint keypoint information includes information indicating the shoulder keypoints of the user's second hand. This allows for verification of the number of shoulder keypoints indicated by the shoulder joint keypoint information, where only the number of shoulder keypoints of the second hand needs to be verified, thus improving verification efficiency. Similarly, when calculating the difference coefficient between the target limb and the shoulder joint, only the shoulder keypoints of the second hand can be used, further improving verification efficiency.
[0122] Meanwhile, when the user is facing the execution subject, the second hand is usually unobstructed, which means that the shoulder image of the second hand may have less shadow. In other words, the key point information of the second hand may be more authentic. Thus, using the key point information of the second hand for inspection can make the verification results more accurate.
[0123] In some embodiments, the posture indicated by the response candidate posture information is determined, and whether the candidate posture information is the target posture information is determined based on the confidence level of the candidate posture information and the confidence level of the key point information.
[0124] Here, the attitude indicated by the target attitude information is the attitude to be executed.
[0125] It should be noted that there may be multiple user images in a single video frame, and correspondingly, multiple candidate pose information may be obtained; and multiple candidate pose information may be responded to (for example, the poses indicated by these multiple candidate pose information are not erroneously triggered poses).
[0126] Here, the poses indicated by multiple candidate pose information may be different. Therefore, it may be necessary to determine whether the target pose information exists among these multiple candidate pose information. For example, it can be determined based on the confidence level of the candidate pose information, and / or based on the degree of matching between the pose indicated by the candidate pose information and a predefined pose.
[0127] Among multiple candidate pose information, there may not be a target pose information. For example, multiple users may be performing a certain pose. If it is detected that these poses are not triggered erroneously, all of them will respond to these candidate pose information. However, since the actions performed by these users are not standard, these poses are not the poses to be executed, thus making these candidate pose information not the target pose information.
[0128] As an example, the pending posture can be understood as the posture of the executing entity waiting to be executed; that is, the executing entity may perform certain instructions based on this posture. However, typically, the executing entity will not immediately execute the target posture information after determining it, because at this point, the target posture information is only the target posture information corresponding to a single video frame. For accuracy, the executing entity may continue to verify whether the target posture information corresponding to subsequent video frames matches the target posture information corresponding to the current video frame. If so, the instructions corresponding to that target posture information can be executed. For example, the executing entity can then perform functions such as page turning, volume adjustment, turning on / off, etc.
[0129] Here, determining whether candidate pose information is the target pose information based on the confidence levels of both candidate pose information and keypoint information can be understood as determining the validity of a user-triggered pose that is not mistakenly triggered. Considering both the confidence levels of candidate pose information and keypoint information simultaneously can make the determination more accurate.
[0130] In some implementations, the similarity between the pose indicated by the candidate pose information and the standard pose can also be considered, which can further improve the accuracy of the determination result.
[0131] To better understand the execution process in this disclosure, one possible execution method of the executing entity is listed below.
[0132] In some embodiments, when the target pose information is included among the multiple candidate pose information, the first video frame image can be used as the reference video frame image, and a first preset number of video frame images can be verified; based on the number of matching video frame images in the first preset number, it is determined whether to execute the pose indicated by the target pose information.
[0133] Here, the target pose information determined by matching video frame images is matched with the target pose information determined in the reference video frame images.
[0134] As an example, by continuously verifying a first preset number of video frame images and determining whether to execute the pose indicated by the target pose information based on the number of matching video frame images in the first preset number, the accuracy of the execution by the executing entity can be further ensured.
[0135] For example, if the pose to be executed determined by the first video frame image is pose B, and if the first preset number of video frame images are checked, the determined poses to be executed do not match pose B. It can be seen that the pose to be executed in the first video frame image is determined incorrectly, so pose B can be omitted.
[0136] Of course, the first preset number can be set according to the actual situation, for example, it can be 7. At this time, if the number of matched video frame images is not less than 5, the posture indicated by the target posture information can be executed; otherwise, the posture indicated by the target posture information will not be executed.
[0137] In some embodiments, in response to determining that the number of matching video frame images in the first preset number is not less than a preset numerical threshold, the posture indicated by the target posture information determined by the reference video frame image is executed.
[0138] Here, if the number of matching video frames in the first preset number is not less than a preset numerical threshold, it indicates that the target pose information determined by most of the video frames in the first preset number matches the target pose information determined by the reference video frame. This allows it to be determined that the pose the user needs to execute is the pose indicated by the target information determined by the reference video frame. Therefore, the system can respond to the pose indicated by the target information determined by the reference video frame. This makes the execution process of the executing entity more precise.
[0139] It should be noted that the first preset number and the preset value threshold can be reasonably set according to the actual situation. For example, the first preset number is 7, and the preset value threshold can be 5.
[0140] In some embodiments, in response to determining that the number of matching video frame images in the first preset number is less than a preset numerical threshold, the determined reference video frame image is updated based on the target pose information determined by each video frame image in the first preset number of video frame images.
[0141] Here, if the number of matching video frames in the first preset number is less than a preset threshold, it indicates that the target pose information determined by most of the video frames in the first preset number does not match the target pose information determined by the reference video frame. In this case, the user may not need the execution subject to execute the pose indicated by the target pose information determined by the current reference video frame. Therefore, the determined reference video frame can be updated based on the target pose information determined by each video frame in the first preset number of video frames. This ensures that the execution subject can execute instructions more accurately.
[0142] In other words, determining that the number of matching video frame images in the first preset number is less than a preset numerical threshold can also be understood as the execution posture determined in most of the video frame images in the first preset number being different from the execution posture determined in the reference video frame image.
[0143] In some embodiments, the determined reference video frame image can be updated according to the temporal order of each video frame in the first preset number.
[0144] As an example, updating the established reference video frame images according to the existing time sequence can enable the user's gestures to be executed more quickly, that is, it can shorten the time required from the user making a gesture to the execution entity executing the gesture.
[0145] To facilitate understanding, let's take an example. The current video frame is video frame 1, and the first number of video frames can be video frames 2-8. If the target pose information determined by video frame 1 does not match the target pose information determined by most of the video frames 2-8, then video frame 2 can be used as the reference video frame. At this point, the target pose information determined by each video frame in video frames 3-9 can be compared with the target pose information determined by video frame 2. If video frame 8 is selected as the reference video frame, it is necessary to wait for video frame 15 to complete determining the target pose information before determining whether to execute the pose indicated by the target pose information. Therefore, by updating the reference video frame images sequentially, the time required from the user making a pose to the execution entity executing the pose can be shortened.
[0146] In some embodiments, both candidate pose information and key point information are obtained from the first video frame. Candidate pose information and key point information can be obtained from the first video frame by: inputting the first video frame image into a pre-built neural network model to obtain initial pose information and key point information; and filtering the obtained initial pose information based on a predefined pose to obtain candidate pose information.
[0147] Here, a pre-built neural network model can be obtained in the following way:
[0148] Obtain the set of training sample images and the initial neural network model.
[0149] The initial neural network model is trained using training sample images from the training sample image set to obtain the pre-built neural network model.
[0150] As an example, a pre-built neural network model is used to obtain initial pose information and human key point information from the first video frame image, making it faster and more accurate to obtain these information. This also speeds up the process from when the user makes a pose to when the subject executes that pose.
[0151] As an example, the initial posture information obtained can be filtered based on a predefined posture. This can save the amount of information that the executing entity needs to verify, and thus save verification time.
[0152] In some embodiments, before filtering the obtained initial pose information based on a predefined pose, the method may further include:
[0153] The validity of the initial attitude information is used to filter the acquired initial attitude information, and / or, information augmentation processing is performed on the initial attitude information.
[0154] Here, by using the validity of the initial posture information to filter the acquired initial posture information, the amount of initial posture information that needs to be processed later can be reduced, thereby speeding up the entire gesture recognition process.
[0155] Here, information augmentation processing of the initial posture information makes it easier to find the required features in the subsequent gesture recognition process, thereby speeding up the entire gesture recognition process.
[0156] In some application scenarios, when both the validity of the initial attitude information is used to filter the acquired initial attitude information and information augmentation processing is performed on the initial attitude information, the validity of the initial attitude information can be used to filter the acquired initial attitude information first, and then information augmentation processing can be performed on the initial attitude information. In this way, the amount of data to be augmented can be reduced, thereby speeding up the processing efficiency.
[0157] In some application scenarios, when both the legitimacy of the initial attitude information is used to filter the acquired initial attitude information and information augmentation processing is performed on the initial attitude information, the initial attitude information can be augmented first, and then the legitimacy of the initial attitude information can be used to filter the acquired initial attitude information. This can make the filtering effect better.
[0158] Of course, the specific order of processing can be reasonably set according to the actual usage.
[0159] Further reference Figure 4 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a posture recognition device, which is similar to... Figure 1 Corresponding to the illustrated posture recognition method embodiment, this device can be specifically applied to various electronic devices.
[0160] like Figure 4 As shown, the posture recognition device of this embodiment includes: a receiving unit 401, used to receive candidate posture information and key point information, wherein the candidate posture information and the key point information correspond to the same user, the candidate posture information indicates that the posture of the first part is a predefined posture, and the key point information indicates the key points of the second part of the user; and a determining unit 402, used to determine whether to respond to the posture indicated by the candidate posture information based on the key point information.
[0161] In some embodiments, the key point information includes at least one sub-key point information, which is used to indicate key points of sub-parts. The determining unit 402 is further configured to: determine the verification method corresponding to each sub-key point information; verify the sub-key point information according to the determined verification method; and determine whether to respond to the posture indicated by the candidate posture information based on the verification result of the sub-key point information.
[0162] In some embodiments, the determining unit 402 is further configured to: determine the verification order of each sub-key point information; verify the sub-key point information according to the determined verification order; and determine whether to verify the next sub-part information according to the verification result of the previous sub-key point information.
[0163] In some embodiments, the determining unit 402 is further configured to: in response to determining that the verification of any sub-key point information fails, determine the posture that does not respond to the candidate posture information.
[0164] In some embodiments, the determining unit 402 is further configured to: verify the sub-key point information in the following manner:
[0165] The sub-keypoint information is validated using the pose of the sub-parts indicated by the sub-keypoint information and / or the confidence level of the sub-keypoint information.
[0166] In some embodiments, the determining unit 402 is further configured to: determine the pose of the sub-part indicated by the sub-keypoint information in the following manner:
[0167] Based on the number of sub-keypoints indicated by the sub-keypoint information and the positional relationship between each sub-keypoint, the pose of the sub-part indicated by the sub-keypoint information is determined.
[0168] In some embodiments, the key point information includes hip joint key point information, and the determining unit 402 is further configured to: verify the hip joint key point information in the following manner:
[0169] Based on the number of hip joint key points indicated by the hip joint key point information, the hip joint key point information is first verified; in response to determining that the hip joint key point information meets the first verification condition, the hip joint key point information is second verified based on the confidence level of the hip joint key point information; in response to determining that the hip joint key point information meets the second verification condition, the difference coefficient between the first part and the hip joint is determined using the hip joint key point information and the candidate posture information, and the hip joint key point information is third verified based on the difference coefficient between the first part and the hip joint.
[0170] In some embodiments, the aforementioned human body key point information includes shoulder joint key point information, and the determining unit 402 is further configured to: verify the shoulder joint key point information in the following manner: perform a first verification of the shoulder joint key point information based on the number of shoulder joint key points indicated by the shoulder joint key point information; in response to determining that the shoulder joint key point information meets the first verification condition, perform a second verification of the shoulder joint key point information based on the confidence level of the shoulder joint key point information; in response to determining that the second verification of the shoulder joint key point information is qualified, determine the difference coefficient between the first part and the shoulder joint using the shoulder joint key point information and the candidate posture information, and perform a third verification of the shoulder joint key point information based on the difference coefficient between the first part and the shoulder joint.
[0171] In some embodiments, candidate pose information is used to indicate the hand pose of the user's first hand, and shoulder joint key point information includes information for indicating the shoulder key points of the user's second hand.
[0172] In some embodiments, the aforementioned key point information includes at least one of the following: limb key point information, eye key point information, hip joint key point information, and shoulder joint key point information.
[0173] In some embodiments, the apparatus is further configured to, in response to determining the posture indicated by the candidate posture information, determine whether the candidate posture information is target posture information based on the confidence level of the candidate posture information and the confidence level of the key point information, wherein the posture indicated by the target posture information is the posture to be executed.
[0174] In some embodiments, both candidate pose information and key point information are obtained from a first video frame. Furthermore, the apparatus is configured to obtain candidate pose information and key point information from the first video frame by: inputting the first video frame image into a pre-built neural network model to obtain initial pose information and key point information; filtering the obtained initial pose information based on a predefined pose to obtain candidate pose information; wherein the pre-built network model is obtained by:
[0175] Obtain the set of training sample images and the initial neural network model.
[0176] The initial neural network model is trained using the training sample images from the aforementioned training sample image set to obtain the aforementioned pre-built neural network model.
[0177] In some embodiments, the above-described apparatus is further configured to filter the acquired initial attitude information by utilizing the legitimacy of the initial attitude information, and / or to perform information enhancement processing on the initial attitude information.
[0178] Please refer to Figure 5 , Figure 5 An exemplary system architecture in which the gesture recognition method of one embodiment of this disclosure can be applied is illustrated.
[0179] like Figure 5 As shown, the system architecture may include terminal devices 501, 502, and 503, a network 504, and a server 505. Network 504 can be used as a medium to provide a communication link between terminal devices 501, 502, and 503 and server 505. Network 504 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0180] Terminal devices 501, 502, and 503 can interact with server 505 via network 504 to receive or send messages, etc. Various client applications, such as web browsers, search engines, and news apps, can be installed on terminal devices 501, 502, and 503. These client applications can receive user commands and perform corresponding functions, such as adding information to a message based on user instructions.
[0181] Terminal devices 501, 502, and 503 can be either hardware or software. When terminal devices 501, 502, and 503 are hardware, they can be various electronic devices with a display screen and support web browsing, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers, etc. When terminal devices 501, 502, and 503 are software, they can be installed in the aforementioned electronic devices. They can be implemented as multiple software programs or software modules (e.g., software programs or software modules used to provide distributed services) or as a single software program or software module. No specific limitations are imposed here.
[0182] Server 505 can be a server that provides various services, such as receiving information retrieval requests sent by terminal devices 501, 502, and 503, retrieving the corresponding display information according to the information retrieval request through various methods, and sending the relevant data for displaying the information to terminal devices 501, 502, and 503.
[0183] It should be noted that the information processing method provided in this embodiment can be executed by a terminal device, and correspondingly, the posture recognition device can be installed in terminal devices 501, 502, and 503. Furthermore, the posture recognition method provided in this embodiment can also be executed by a server 505, and correspondingly, the information processing device can be installed in server 505.
[0184] It should be understood that Figure 4 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0185] The following is for reference. Figure 6 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 5 The diagram shows the structure of the terminal device or server in this disclosure. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (e.g., vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0186] like Figure 6 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 508 into a random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of the electronic device 600. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0187] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0188] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0189] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0190] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0191] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0192] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: receive candidate posture information and key point information, wherein the candidate posture information and the key point information correspond to the same user, the candidate posture information indicates that the posture of a first part is a predefined posture, and the key point information indicates the key points of the user's second part; and, based on the key point information, determine whether to respond to the posture indicated by the candidate posture information.
[0193] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0194] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0195] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit is not necessarily a limitation on the unit itself; for example, a receiving unit can also be described as "a unit that receives candidate pose information and key point information".
[0196] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0197] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0198] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0199] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0200] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A pose recognition method, characterized in that, include: Receive candidate pose information and key point information, wherein the candidate pose information and the key point information correspond to the same user, the candidate pose information indicates that the pose of the first part is a predefined pose, and the key point information indicates the key points of the user's second part. Based on the key point information, determining whether to respond to the posture indicated by the candidate posture information, wherein the key point information includes at least one sub-key point information, the sub-key point information being used to indicate key points of a sub-part, and determining whether to respond to the posture indicated by the candidate posture information based on the key point information includes: Determine the verification method corresponding to each sub-key point information; The sub-key point information is verified according to the determined verification method, and based on the verification result of the sub-key point information, it is determined whether to respond to the posture indicated by the candidate posture information. The key point information includes shoulder joint key point information, and the shoulder joint key point information is verified in the following manner: Based on the number of shoulder joint key points indicated by the shoulder joint key point information, the shoulder joint key point information is verified for the first time. In response to determining that the shoulder joint key point information meets the first verification condition, a second verification is performed on the shoulder joint key point information based on the confidence level of the shoulder joint key point information; In response to the determination that the shoulder joint key point information has passed the second verification, the difference coefficient between the first part and the shoulder joint is determined using the shoulder joint key point information and the candidate posture information, and the shoulder joint key point information is verified for the third time based on the difference coefficient between the first part and the shoulder joint. The candidate posture information is used to indicate the hand posture of the user's first hand, and the shoulder joint key point information includes information to indicate the shoulder key points of the user's second hand.
2. The method according to claim 1, characterized in that, The verification of sub-key point information according to the determined verification method includes: Determine the verification order for each sub-key point information; The sub-key point information is verified according to the determined verification order, and the verification result of the previous sub-key point information is used to determine whether to verify the next sub-part information.
3. The method according to claim 1, characterized in that, The verification result based on the sub-keypoint information determines whether to respond to the pose indicated by the candidate pose information, including: In response to the determination that any sub-key point information verification fails, it is determined that the attitude indicated by the candidate attitude information will not be responded to.
4. The method according to claim 1, characterized in that, The sub-key point information is verified using the following method: The sub-keypoint information is validated using the pose of the sub-parts indicated by the sub-keypoint information and / or the confidence level of the sub-keypoint information.
5. The method according to claim 4, characterized in that, The pose of the sub-part indicated by the sub-keypoint information is determined in the following way: Based on the number of sub-keypoints indicated by the sub-keypoint information and the positional relationship between each sub-keypoint, the pose of the sub-part indicated by the sub-keypoint information is determined.
6. The method according to claim 1, characterized in that, The key point information includes hip joint key point information, and the hip joint key point information is verified in the following manner: The hip joint key point information is first verified based on the number of hip joint key points indicated by the hip joint key point information. In response to determining that the hip joint key point information meets the first verification condition, a second verification is performed on the hip joint key point information based on the confidence level of the hip joint key point information; In response to determining that the hip joint key point information meets the second verification condition, the difference coefficient between the first part and the hip joint is determined using the hip joint key point information and the candidate posture information, and the hip joint key point information is verified for the third time based on the difference coefficient between the first part and the hip joint.
7. The method according to claim 1, characterized in that, The key point information includes at least one of the following: limb key point information, eye key point information, hip joint key point information, and shoulder joint key point information.
8. The method according to claim 1, characterized in that, The method further includes: In response to determining the posture indicated by the candidate posture information, based on the confidence level of the candidate posture information and the confidence level of the key point information, it is determined whether the candidate posture information is the target posture information, wherein the posture indicated by the target posture information is the posture to be executed.
9. The method according to claim 1, characterized in that, Both candidate pose information and key point information are obtained from the first video frame, and are obtained from the first video frame in the following manner: The first video frame image is input into a pre-built neural network model to obtain initial pose information and key point information; Based on the predefined pose, the obtained initial pose information is filtered to obtain candidate pose information; The pre-built network model is obtained in the following way: Obtain the set of training sample images and the initial neural network model. The initial neural network model is trained using training sample images from the training sample image set to obtain the pre-built neural network model.
10. The method according to claim 9, characterized in that, Before filtering the obtained initial pose information based on a predefined pose, the method further includes: The validity of the initial attitude information is used to filter the acquired initial attitude information, and / or, information augmentation processing is performed on the initial attitude information.
11. A posture recognition device, characterized in that, include: A receiving unit is used to receive candidate posture information and key point information, wherein the candidate posture information and the key point information correspond to the same user, the candidate posture information indicates that the posture of the first part is a predefined posture, and the key point information indicates the key points of the user's second part. A determining unit is configured to determine, based on the key point information, whether to respond to the posture indicated by the candidate posture information, wherein the key point information includes at least one sub-key point information, the sub-key point information being used to indicate key points of a sub-part, and... The determining unit is further configured to determine whether to respond to the posture indicated by the candidate posture information based on the key point information in the following manner: determining the verification method corresponding to each sub-key point information; verifying the sub-key point information according to the determined verification method; and determining whether to respond to the posture indicated by the candidate posture information based on the verification result of the sub-key point information, wherein the key point information includes shoulder joint key point information, and... The determining unit is further configured to verify the shoulder joint key point information in the following manner: performing a first verification of the shoulder joint key point information based on the number of shoulder joint key points indicated by the shoulder joint key point information; in response to determining that the shoulder joint key point information meets the first verification condition, performing a second verification of the shoulder joint key point information based on the confidence level of the shoulder joint key point information; in response to determining that the second verification of the shoulder joint key point information is qualified, determining the difference coefficient between the first part and the shoulder joint using the shoulder joint key point information and the candidate posture information, and performing a third verification of the shoulder joint key point information based on the difference coefficient between the first part and the shoulder joint; wherein the candidate posture information is used to indicate the hand posture of the user's first hand, and the shoulder joint key point information includes information for indicating the shoulder key points of the user's second hand.
12. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-10.
13. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-10.
Citation Information
Patent Citations
Pedestrian posture recognition method and device, and unmanned vehicle
CN111907520A
Display equipment and control method of display equipment
CN113918010A
Gesture misrecognition prevention method, and electronic device
WO2022095983A1