Method, apparatus, terminal device and storage medium for identifying human body movements
By acquiring the standard action image corresponding to the target action image and replacing the local area, the problem of low recognition accuracy caused by occlusion information is solved, and a higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111529153.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-14
AI Technical Summary
In the prior art, due to the presence of occlusion information in the motion picture, the accuracy of the human body movement recognition results is low.
By acquiring the standard action image corresponding to the target action image, the standard action image is used to replace the target action image locally, reducing the influence of occlusion information and improving the recognition accuracy.
Effectively reduce the impact of occlusion information on human body movement recognition and improve the accuracy of human body movement recognition results.
Smart Images

Figure CN114267050B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human motion recognition, and particularly to a method, device, terminal device and storage medium for recognizing human motions. Background Art
[0002] Currently, a method for human motion recognition is disclosed, which acquires motion pictures generated when a user moves, compares the motion pictures with standard action pictures to obtain a comparison result, and obtains a human motion recognition result based on the final comparison result.
[0003] However, when using the existing method, when there is occlusion information in the motion pictures, it is difficult to accurately analyze the human motion in the action pictures, resulting in a low accuracy rate of the human motion recognition result. Summary of the Invention
[0004] The main object of the present invention is to provide a method, device, terminal device and storage medium for recognizing human motions, aiming to solve the technical problem that the accuracy rate of the human motion recognition result is low when using the existing method in the prior art.
[0005] To achieve the above object, the present invention proposes a method for recognizing human motions, the method comprising the following steps:
[0006] Acquire a target action picture including the human motion of a target user;
[0007] Acquire a standard action picture corresponding to the target action picture, wherein the human motion in the standard action picture has no occlusion information;
[0008] Use the standard action picture to perform local area replacement on the target action picture to obtain a replaced target action picture;
[0009] Obtain a human motion recognition result of the target user according to the replaced target action picture.
[0010] Optionally, before the step of acquiring the standard action picture corresponding to the target action picture, the method further comprises:
[0011] Divide the target action picture into multiple divided areas;
[0012] The step of acquiring the standard action picture corresponding to the target action picture comprises:
[0013] Acquire multiple standard action pictures corresponding to the multiple divided areas, one divided area corresponding to one standard action picture;
[0014] The step of using the standard action picture to perform local area replacement on the target action picture comprises:
[0015] Using the standard action images corresponding to each of the divided regions, perform local region replacement on each of the divided regions to obtain a plurality of replaced divided regions;
[0016] Based on the plurality of replaced divided regions, obtain a replaced target action image.
[0017] Optionally, the step of using the standard action images corresponding to each of the divided regions to perform local region replacement on each of the divided regions to obtain a plurality of replaced divided regions includes:
[0018] Divide each of the divided regions into a plurality of alternative blocks;
[0019] In the plurality of alternative blocks in each of the divided regions, determine the selected alternative block of each of the divided regions;
[0020] Rotate the preset image corresponding to each of the divided regions to obtain a rotated image corresponding to each of the divided regions;
[0021] According to the size of the selected alternative block of each of the divided regions, perform a scaling operation on the rotated image corresponding to each of the divided regions to obtain a scaled image corresponding to each of the divided regions;
[0022] Use the scaled image corresponding to each of the divided regions to replace the selected alternative block in each of the divided regions to obtain a plurality of replaced divided regions.
[0023] Optionally, the step of obtaining the human action recognition result of the target user according to the replaced target action image includes:
[0024] According to the replaced target action image, obtain a plurality of target coordinates corresponding to a plurality of key points;
[0025] According to the plurality of target coordinates, obtain a plurality of target joint angles;
[0026] Obtain a plurality of standard joint angles of the standard action image corresponding to the target action image;
[0027] According to the comparison result of the plurality of target joint angles and the plurality of standard joint angles, obtain the human action recognition result of the target user.
[0028] Optionally, the step of obtaining a plurality of target coordinates corresponding to a plurality of key points according to the replaced target action image includes:
[0029] Generate a first image with a size of 4a * 4a, a second image with a size of 2a * 2a, and a third image with a size of a * a based on the replaced target action image, where a is a natural number not equal to 0;
[0030] Perform convolutional downsampling on the first image to obtain a fourth image, and perform convolutional downsampling on the second image to obtain a fifth image;
[0031] Process the first image to obtain multiple first coordinates corresponding to multiple key points, and process the second image to obtain multiple second coordinates corresponding to multiple key points;
[0032] Overlay the fourth image and the second image to obtain a sixth image, and overlay the third image and the fifth image to obtain a seventh image;
[0033] Perform convolutional downsampling on the sixth image to obtain an eighth image;
[0034] Process the eighth image to obtain multiple third coordinates corresponding to multiple key points, and process the seventh image to obtain multiple fourth coordinates corresponding to multiple key points;
[0035] Obtain multiple target coordinates based on multiple first coordinates, multiple second coordinates, multiple third coordinates, and multiple fourth coordinates.
[0036] Optionally, before the step of obtaining multiple target coordinates based on multiple first coordinates, multiple second coordinates, multiple third coordinates, and multiple fourth coordinates, the method further includes:
[0037] Determine a first weight corresponding to multiple first coordinates, a second weight corresponding to multiple second coordinates, a third weight corresponding to multiple third coordinates, and a fourth weight corresponding to multiple fourth coordinates according to the sizes of the first image, the second image, the third image, and the fourth image, where the first weight, the second weight, the third weight, and the fourth weight decrease in sequence;
[0038] Obtain the target coordinate corresponding to each key point according to the first coordinate corresponding to each key point, the second coordinate corresponding to each key point, the third coordinate corresponding to each key point, the fourth coordinate corresponding to each key point, the first weight, the second weight, the third weight, and the fourth weight.
[0039] Optionally, the step of obtaining the human action recognition result corresponding to the action image according to multiple target joint angles and multiple standard joint angles includes:
[0040] Calculate the angle difference corresponding to each joint according to the target joint angle corresponding to each joint and the standard joint angle corresponding to each joint;
[0041] Obtain the joint weights corresponding to multiple joints according to the multiple angle differences, and the joint weights are positively correlated with the angle differences of the joints;
[0042] Calculate the human body movement similarity by using the multiple angle differences corresponding to multiple joints and the multiple joint weights corresponding to multiple joints;
[0043] Obtain the human body movement recognition result of the action image according to the comparison result between the human body movement similarity and the preset similarity threshold.
[0044] In addition, to achieve the above object, the present invention also proposes a human body movement recognition device, and the device includes:
[0045] A first acquisition module, configured to acquire a target action image including the human body movement of a target user;
[0046] A second acquisition module, configured to acquire a standard action image corresponding to the target action image, and there is no occlusion information in the human body movement in the standard action image;
[0047] A replacement module, configured to perform local area replacement on the target action image by using the standard action image to obtain a replaced target action image;
[0048] An acquisition module, configured to obtain the human body movement recognition result of the target user according to the replaced target action image.
[0049] In addition, to achieve the above object, the present invention also proposes a terminal device, and the terminal device includes: a memory, a processor, and a human body movement recognition program stored on the memory and running on the processor. When the human body movement recognition program is executed by the processor, the steps of the human body movement recognition method described in any one of the above are implemented.
[0050] In addition, to achieve the above object, the present invention also proposes a storage medium, and a human body movement recognition program is stored on the storage medium. When the human body movement recognition program is executed by a processor, the steps of the human body movement recognition method described in any one of the above are implemented.
[0051] The technical solution of the present invention proposes a method for recognizing human actions, which includes obtaining a target action image including the human actions of a target user; obtaining a standard action image corresponding to the target action image, where the human actions in the standard action image have no occlusion information; using the standard action image to perform local area replacement on the target action image to obtain a replaced target action image; and obtaining a recognition result of the human actions of the target user according to the replaced target action image.
[0052] In the existing methods, the image comparison result is obtained by comparing the target action image with the standard action image, and the recognition result of the human actions is obtained based on the image comparison result. When there is occlusion information in the target action image, it is difficult to recognize the human actions in the target action image, resulting in a low accuracy rate of the recognition result of the human actions. However, in the present invention, the local area of the target action image is replaced by the standard action image, which effectively reduces the influence of the occlusion information on the human actions in the target action image, so that the human actions in the replaced target action image can be accurately recognized, thereby improving the accuracy rate of the recognition result of the human actions. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0054] Figure 1 It is a schematic structural diagram of a terminal device for the hardware operating environment related to the solution of the embodiment of the present invention;
[0055] Figure 2 It is a schematic flowchart of the first embodiment of the method for recognizing human actions of the present invention;
[0056] Figure 3 It is a schematic diagram of an action image of the present invention;
[0057] Figure 4 For the present invention Figure 2 It is a refined flowchart of step S14 in the present invention;
[0058] Figure 5 It is a schematic structural diagram of multiple key points of the present invention;
[0059] Figure 6 It is a schematic block diagram of the first embodiment of the device for recognizing human actions of the present invention.
[0060] The implementation, functional features, and advantages of the present invention will be further described in conjunction with embodiments with reference to the accompanying drawings. Detailed implementation manners
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0062] Refer to Figure 1 , Figure 1 which is a schematic structural diagram of a terminal device for the hardware operating environment involved in the embodiment solution of the present invention.
[0063] Generally, the terminal device includes: at least one processor 301, a memory 302, and a recognition program for human body movements stored on the memory and operable on the processor. The recognition program for human body movements is configured to implement the steps of the recognition method for human body movements as described above.
[0064] The processor 301 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 301 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. The processor 301 may also include an AI (Artificial Intelligence) processor, which is used to process the operations related to the recognition method for human body movements, so that the recognition method model for human body movements can be autonomously trained and learned to improve efficiency and accuracy.
[0065] The memory 302 may include one or more storage media, which may be non-transitory. The memory 302 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory storage media in the memory 302 is used to store at least one instruction for being executed by the processor 301 to implement the human motion recognition method provided in the method embodiments of the present application.
[0066] In some embodiments, the terminal may further optionally include: a communication interface 303 and at least one peripheral device. The processor 301, the memory 302, and the communication interface 303 may be connected through a bus or signal lines. Each peripheral device may be connected to the communication interface 303 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 304, a display screen 305, and a power supply 306.
[0067] The communication interface 303 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 301 and the memory 302. In some embodiments, the processor 301, the memory 302, and the communication interface 303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 301, the memory 302, and the communication interface 303 may be implemented on a separate chip or circuit board, and the present embodiment does not limit this.
[0068] The radio frequency circuit 304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 304 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 304 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 304 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 304 may communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 304 may further include a circuit related to NFC (Near Field Communication), and the present application does not limit this.
[0069] The display screen 305 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 305 is a touch display screen, the display screen 305 also has the ability to collect touch signals on or above the surface of the display screen 305. The touch signals can be input to the processor 301 as control signals for processing. At this time, the display screen 305 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 305 can be one, the front panel of the electronic device; in other embodiments, the display screen 305 can be at least two, respectively arranged on different surfaces of the electronic device or in a folding design; in still other embodiments, the display screen 305 can be a flexible display screen, arranged on the curved surface or folding surface of the electronic device. Even, the display screen 305 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 305 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0070] The power supply 306 is used to supply power to each component in the electronic device. The power supply 306 can be alternating current, direct current, a primary battery, or a rechargeable battery. When the power supply 306 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology. Those skilled in the art can understand that Figure 1 the structure shown in does not constitute a limitation on the terminal device, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0071] In addition, an embodiment of the present invention also proposes a storage medium, on which a recognition program for human body movements is stored. When the recognition program for human body movements is executed by a processor, the steps of the recognition method for human body movements as described above are implemented. Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiment of the storage medium involved in the present application, please refer to the description of the method embodiment of the present application. By way of example, the program instructions can be deployed to be executed on a terminal device, or on multiple terminal devices corresponding to one location, or on multiple terminal devices distributed at multiple locations and interconnected by a communication network.
[0072] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the above storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0073] Based on the above hardware structure, an embodiment of the method for recognizing human body movements of the present invention is proposed.
[0074] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the method for recognizing human body movements of the present invention. The method is used for a terminal device, and the method includes the following steps:
[0075] Step S11: Obtain a target action image including the human body movement of the target user.
[0076] It should be noted that the execution subject of the present invention is a terminal device. The terminal device is installed with a program for recognizing human body movements. When the terminal device executes the program for recognizing human body movements, the steps of the method for recognizing human body movements of the present invention are implemented.
[0077] The target action image generally refers to an image of the target user doing a certain action collected by a camera or a webcam. For example, an image of a certain dance pose of user A is an action image. The target action image is generally selected as a high-definition image to improve the accuracy of the recognition result.
[0078] First, the target action image can be cropped (when the target action image is an image including multiple people, cropping processing is required; otherwise, cropping processing is not required), and the image corresponding to the main person is cropped out. Then, the middle prediction PL of the low-order hourglass HGL is used as a free detector to extract the effective human body image from the image corresponding to the main person. The effective human body image continues to the next step S12 for operation.
[0079] Refer to Figure 3 , Figure 3 which is a schematic diagram of the action image of the present invention. In Figure 3 , the part framed by the external large rectangular frame is the image corresponding to the main person. That is, in Figure 3 , the two people on the left are not the main people. Then, the image corresponding to the main person is detected to obtain the effective human body image framed by the small rectangular frame.
[0080] Step S12: Obtain a standard action image corresponding to the target action image, where there is no occlusion information in the human action in the standard action image.
[0081] Step S13: Use the standard action image to perform local region replacement on the target action image to obtain the replaced target action image.
[0082] The standard action image refers to the action image of the standard action corresponding to the human action in the target action image, and there should be no occlusion information in the standard action image. The occlusion information in the action image may refer to the information that the limbs in the action image are occluded. The target action image with occlusion information cannot be accurately recognized. Therefore, the present invention uses the standard action image to perform local replacement on the target action image to reduce the influence of the occlusion information in the target action image and further improve the recognition accuracy of the target action image. When there is no occlusion information in the target action image, the method of the present invention is still used for replacement, and the original recognition accuracy of the target action image will not be affected after replacement because there is no occlusion information in the standard action image.
[0083] Specifically, before the step of obtaining the standard action image corresponding to the target action image, the method further includes: dividing the target action image into multiple divided regions; the step of obtaining the standard action image corresponding to the target action image includes: obtaining multiple standard action images corresponding to the multiple divided regions, with one divided region corresponding to one standard action image; the step of using the standard action image to perform local region replacement on the target action image includes: using the standard action image corresponding to each divided region to perform local region replacement on each divided region to obtain multiple replaced divided regions; and obtaining the replaced target action image according to the multiple replaced divided regions.
[0084] Based on the above description, it is necessary to process the effective human body image. When dividing the target action image into multiple divided regions, first process the target action image into an effective human body image, and then divide the effective human body image into multiple divided regions.
[0085] First, the size of the effective human body image is in various forms and needs to be adjusted to an image with a size of 4a * 4a, where a is a non-zero natural number. In the present invention, a = 64 is a preferred choice.
[0086] For the convenience of different analyses of different human body parts, it is necessary to divide the 4a * 4a image into a head region, a torso region, and a leg region. Specifically, the 4a * 4a image is divided into 16 blocks with a size of a * a. The 4a * 4a image is divided into 4 rows. The first row is the head region, the second and third rows are the torso regions, and the fourth row is the leg region. At this time, each row includes 4 alternative blocks of the same size. That is, in the present invention, there are three types of multiple divided regions: the head region, the torso region, and the leg region. The user can set other division methods based on requirements, and the present invention does not make limitations. This is just a better choice here.
[0087] Usually, the standard action image is obtained from the standard image library. The standard image library includes the standard head image corresponding to the head region, the standard torso image corresponding to the torso region, and the standard leg image corresponding to the leg region. The images in the standard image library do not have occlusion information.
[0088] In specific applications, for the head region, randomly select 0 to 1 standard head images from the standard image library, randomly select 1 to 3 standard torso images from the standard image library, and randomly select 1 to 2 standard leg images from the standard image library. These randomly selected images are multiple standard action images, including the standard head image, the standard leg image, and the standard torso image.
[0089] For each divided region (one of the head region, the torso region, and the leg region), it is necessary to perform local replacement using the corresponding standard action image. For example, the head region needs to be locally replaced using the standard head image.
[0090] After local region replacement is completed for all multiple divided regions, recombine them into the original image, that is, obtain the target action image after replacement. For the other parts of the target action image after replacement that are not replaced, the position and information will not change.
[0091] Further, the step of using the standard action image corresponding to each divided region to perform local region replacement on each divided region to obtain multiple replaced divided regions includes: dividing each divided region into multiple alternative blocks; determining the selected alternative block of each divided region among the multiple alternative blocks in each divided region; rotating the preset image corresponding to each divided region to obtain the rotated image corresponding to each divided region; performing a scaling operation on the rotated image corresponding to each divided region according to the size of the selected alternative block of each divided region to obtain the scaled image corresponding to each divided region; using the scaled image corresponding to each divided region to replace the selected alternative block in each divided region to obtain multiple replaced divided regions.
[0092] As described above, a valid human body image is processed into an image of 4a * 4a, and then divided into 16 images of a * a, which are respectively related to three divided regions. Multiple small blocks of a * a corresponding to each divided region are an alternative block. In the present invention, preferably, the selected alternative block is usually randomly determined. For example, if the head region corresponds to a standard head image, then one of the 4 alternative blocks in the first row is randomly selected as the selected alternative block. The replaced alternative block is the selected alternative block.
[0093] Generally, multiple alternative blocks included in each of the divided regions are of the same size, and then the obtained scaled image has the same size as each block. In some embodiments, if the sizes of multiple alternative blocks are different, then the selected alternative block is first determined, and then the size of the final scaled image is set to the size of the selected alternative block. Generally, to improve the operation efficiency and accuracy, the sizes of all alternative blocks are the same.
[0094] For example, for a standard head image, it needs to be randomly rotated within [-60°, 60°] to obtain a rotated image. Then, for the rotated image, it is scaled to a scaled image of size a * a, and any alternative block in the head region is randomly replaced to obtain the replaced head region: In this embodiment, the scaled image is placed in any of the 4 blocks in the first row for replacement. It can be understood that the standard head image includes 0 or 1 image. For 0 standard head images, this replacement operation is not performed. If it includes 1 standard head image, then any block in the head region is replaced.
[0095] Another example, for a standard torso image, it needs to be randomly rotated within [-60°, 60°] to obtain a rotated image. Then, for the rotated image, it is scaled to a scaled image of size a * a, and any alternative block in the torso region is randomly replaced to obtain the replaced torso region: In this embodiment, the scaled image is placed in any alternative block among the 8 blocks in the second and third rows for replacement. It can be understood that the standard leg image includes 1, 2, or 3 images. For 1 standard torso image, any alternative block in the torso region is replaced. For 2 standard torso images, any two alternative blocks in the torso region are replaced; for 3 standard torso images, any three alternative blocks in the torso region are replaced.
[0096] Similarly, for the standard leg image, it is necessary to randomly rotate it to [-60°, 60°] to obtain a rotated image, and then scale the rotated image to become a scaled image of size a*a, and randomly replace any candidate block in the leg area to obtain the replaced leg area: in this embodiment, the scaled image is placed in any of the 4 candidate blocks in the fourth row for replacement. It can be understood that the standard leg image includes 1 or 2 images. For 1 standard leg image, any candidate block in the leg area is replaced. For 2 standard leg images, any two candidate blocks in the leg area are replaced; for 3 standard leg images, any three candidate blocks in the leg area are replaced. At this point, the corresponding replacement operations have been completed for the three divided areas.
[0097] Step S14: obtaining a human motion recognition result of the target user according to the replaced target motion image.
[0098] By analyzing the replaced target action image, the human body action recognition result of the target user is obtained. Since the replaced target action image has been partially replaced by the standard action image, the recognition accuracy of the replaced target action image is higher.
[0099] The technical solution of the present invention proposes a method for recognizing human motions, which includes obtaining a target motion image including the human motions of a target user; obtaining a standard motion image corresponding to the target motion image, wherein the human motions in the standard motion image do not have any occlusion information; using the standard motion image to replace a local area of the target motion image to obtain a replaced target motion image; and obtaining a human motion recognition result of the target user based on the replaced target motion image.
[0100] In the existing method, the target action image is compared with the standard action image to obtain the image comparison result, and the human action recognition result is obtained based on the image comparison result. When there is occluded information in the target action image, the human action in the target action image is difficult to be recognized, resulting in a low accuracy rate of the human action recognition result. In the present invention, the target action image is replaced with a local area using the standard action image, which effectively reduces the influence of the occlusion information on the human action in the target action image, so that the human action in the replaced target action image can be accurately recognized, thereby improving the accuracy rate of the human action recognition result.
[0101] Reference Figure 4 , Figure 4 For the present invention Figure 2 Flow chart of step S14 in detail; step S12 includes:
[0102] Step S21: Obtain multiple target coordinates corresponding to multiple key points according to the replaced target action image.
[0103] Step S22: Obtain multiple target joint angles according to the multiple target coordinates.
[0104] It should be noted that in the present invention, there are 16 multiple key points. The multiple key points refer to multiple key points corresponding to the human body. One key point corresponds to one target coordinate. The target coordinate can refer to the pixel coordinate or the coordinate in the world coordinate system. The present invention does not make a limitation.
[0105] Refer to Figure 5 , Figure 5 which is a schematic structural diagram of multiple key points of the present invention. In Figure 4 , the multiple key points include head 0, neck 1, chest 8, shoulders (2 and 5), elbows (3 and 6), wrists (4 and 7), pelvis 9, hips (10 and 13), knees (11 and 14), and ankles (12 and 15). That is, in the present invention, 16 target coordinates corresponding to 16 key points need to be obtained.
[0106] Further, the step of obtaining multiple target coordinates corresponding to multiple key points according to the replaced target action image includes: generating a first image with a size of 4a * 4a, a second image with a size of 2a * 2a, and a third image with a size of a * a according to the replaced target action image, where a is a natural number not equal to 0; performing convolutional downsampling on the first image to obtain a fourth image, and performing convolutional downsampling on the second image to obtain a fifth image; processing the first image to obtain multiple first coordinates corresponding to multiple key points, and processing the second image to obtain multiple second coordinates corresponding to multiple key points; superimposing the fourth image and the second image to obtain a sixth image, and superimposing the third image and the fifth image to obtain a seventh image; performing convolutional downsampling on the sixth image to obtain an eighth image; processing the eighth image to obtain multiple third coordinates corresponding to multiple key points, and processing the seventh image to obtain multiple fourth coordinates corresponding to multiple key points; obtaining multiple target coordinates according to the multiple first coordinates, multiple second coordinates, multiple third coordinates, and multiple fourth coordinates.
[0107] In the above steps, when obtaining the coordinates of multiple key points (the first coordinate, the second coordinate, the third coordinate, and the fourth coordinate), the heat map algorithm and the FC regression algorithm are used. In the present invention, one key point corresponds to one first coordinate, one second coordinate, one third coordinate, and one fourth coordinate. That is, the 16 key points of the present invention correspond to 16 first coordinates (denoted as wherein, i refers to the coordinates corresponding to the i-th key point), 16 second coordinates (denoted as ), 16 third coordinates (denoted as ), and 16 fourth coordinates (denoted as ), where i takes values from 1 to 16 in sequence.
[0108] Specifically, before the step of obtaining multiple target coordinates according to the multiple first coordinates, multiple second coordinates, multiple third coordinates, and multiple fourth coordinates, the method further includes: determining a first weight corresponding to the multiple first coordinates, a second weight corresponding to the multiple second coordinates, a third weight corresponding to the multiple third coordinates, and a fourth weight corresponding to the multiple fourth coordinates according to the sizes of the first image, the second image, the third image, and the fourth image, where the first weight, the second weight, the third weight, and the fourth weight decrease in sequence; obtaining the target coordinate corresponding to each key point according to the first coordinate corresponding to each key point, the second coordinate corresponding to each key point, the third coordinate corresponding to each key point, the fourth coordinate corresponding to each key point, the first weight, the second weight, the third weight, and the fourth weight.
[0109] Specifically, the first weight is denoted as w1, the second weight is denoted as w2, the third weight is denoted as w3, and the fourth weight is denoted as w4. The user can set the specific values of the corresponding weights based on requirements. In the present invention, w1 = 0.4, w2 = 0.3, w3 = 0.2, w4 = 0.1, that is, as described above, they decrease in sequence. For each key point, the corresponding first weight, second weight, third weight, and fourth weight are the same.
[0110] According to the first coordinate corresponding to each key point, the second coordinate corresponding to each key point, the third coordinate corresponding to each key point, the fourth coordinate corresponding to each key point, the first weight, the second weight, the third weight, and the fourth weight, use formula four to calculate the target coordinate corresponding to each key point; formula four is:
[0111]
[0112] wherein, (x i , y i ) represents the target coordinate of the i-th key point (i takes values from 0 to 15, a total of 16 key points), and w i respectively takes w1 = 0.4, w2 = 0.3, w3 = 0.2, w4 = 0.1.
[0113] In the present invention, the standard coordinates of each of the multiple key points corresponding to the standard action image, the standard coordinates of the key points can also be obtained by using the method of the present invention. Referring to the acquisition of the target coordinates above, however, when obtaining the standard coordinates, it is not necessary to execute steps S22 - S24, that is, directly adjust the size of the standard action image to obtain an image of 4a * 4a, and then use this image as the second image and execute the above method to obtain the corresponding standard coordinates.
[0114] 16 target coordinates, corresponding to multiple joints. In the present invention, each joint involves 3 key points. Using the 16 target coordinates, the target joint angles corresponding to each joint are obtained.
[0115] Specifically, using a preset comparison table, determine the multiple joints corresponding to the multiple key points. The preset comparison table includes the key points corresponding to different joints; based on the target coordinates corresponding to each joint, calculate the target joint angle corresponding to each joint.
[0116] Referring to Table 1, Table 1 is an example of the preset comparison table of the present invention, as follows:
[0117] Table 1
[0118]
[0119]
[0120] In Table 1, 16 key points correspond to 12 joints. The corresponding relationship between the numbers of each joint and the numbers of the key points is the corresponding relationship between the joint and the key point. Usually, for each joint, from left to right, they are the first key point (represented by s), the middle key point (represented by m), and the ending key point (represented by e). Referring to Figure 4 , for joint 1, it represents the angle between the user's head and shoulders.
[0121] According to the above preset comparison table, determine the key points involved in each joint, and then use the target coordinates corresponding to the key points to solve the target joint angles of each joint. One joint corresponds to one target joint angle.
[0122] Specifically, the step of calculating the target joint angle corresponding to each joint based on the target coordinates corresponding to each joint includes: based on the target coordinates corresponding to each joint, using Formula 1 to calculate the target joint angle corresponding to each joint; Formula 1 is:
[0123]
[0124] where, (x ps ,y ps ) is the target coordinate of the first key point of the pth joint, (xpm , y pm ) is the target coordinate of the intermediate key point of the p-th joint, (x pe , y pe ) is the target coordinate of the end key point of the p-th joint. What is solved is the cosine value, and the inverse trigonometric function is used to solve the specific angle value - the target joint angle.
[0125] Step S23: Obtain multiple standard joint angles of the standard action image corresponding to the target action image.
[0126] Step S24: Obtain the human action recognition result of the target user according to the comparison result of multiple target joint angles and multiple standard joint angles.
[0127] Generally, a user will make an imitation action to imitate a favorite standard action, so as to obtain the action image of the user making the corresponding imitation action, that is, the target action image in step S11. For a standard action, there will be a standard action image, and the standard action image also corresponds to the standard coordinates of multiple key points (in the present invention, the key points are the above 16 key points). Then, using a preset look-up table and multiple standard coordinates, according to the solution method of the target joint angle of the present invention, multiple standard joint angles corresponding to multiple joints are solved. One joint corresponds to one standard joint angle. In the embodiments of the present invention, the standard joint angles usually also include 12.
[0128] In some embodiments, the standard action image is fixed. Multiple standard joint angles can be obtained according to the method in the previous paragraph, and then multiple standard joint angles are stored for direct use in subsequent applications without having to repeatedly solve multiple standard joint angles, thereby saving computing resources and computing time.
[0129] The step of obtaining the human action recognition result corresponding to the action image according to multiple target joint angles and multiple standard joint angles includes: calculating the angle difference corresponding to each joint according to the target joint angle corresponding to each joint and the standard joint angle corresponding to each joint; obtaining the joint weights corresponding to multiple joints according to multiple angle differences, and the joint weights are positively correlated with the angle differences of the joints; calculating the human action similarity using multiple angle differences corresponding to multiple joints and multiple joint weights corresponding to multiple joints; obtaining the human action recognition result of the action image according to the comparison result of the human action similarity and a preset similarity threshold.
[0130] In the embodiments of the present invention, the target joint angle of a joint p is represented as θ p , and the corresponding standard joint angle is represented as θ' p , and their difference Δθ is solvedp , denoted as: Δθ p = θ p - θ′ p , and then continue to use the joint angle difference Δθ of joint p p , according to Formula 2, obtain the joint weight of joint p. Formula 2 is as follows:
[0131]
[0132] where, M p = Δθ p , n refers to the total number of joints, and W p is the joint weight of joint p. It can be seen that the joint weight is positively correlated with the joint angle difference. Then, based on the multiple angle differences corresponding to the multiple joints and the joint weights corresponding to the multiple joints, use Formula 3 to calculate the human action similarity; Formula 3 is as follows:
[0133]
[0134] where, s is the human action similarity, and C usually takes the value of 360.
[0135] In the present invention, the human action similarity is within the interval [0, 1]. When the human action similarity is 0, it means that the action in the action image is completely inconsistent with the standard action; when the human action similarity is 1, it means that the action in the action image is completely consistent with the standard action. The preset similarity threshold is 0.7. When s is greater than 0.7, it indicates that the action corresponding to the marked action image is accurate, otherwise, it is inaccurate.
[0136] In this embodiment, a specific method for obtaining the human action recognition result of the target user is disclosed. The joint angles corresponding to the key points can accurately reflect the user's action, so that the accuracy of human action recognition is relatively high.
[0137] Referring to Figure 6 , Figure 6 is the structural block diagram of the first embodiment of the human action recognition device of the present invention. The device is used for a terminal device. Based on the same inventive concept as the foregoing embodiment, the device includes:
[0138] The first acquisition module 10 is used to acquire a target action image including the human action of the target user;
[0139] The second acquisition module 20 is used to acquire a standard action image corresponding to the target action image, and there is no occlusion information in the human action in the standard action image;
[0140] The replacement module 30 is used to perform local area replacement on the target action image by using the standard action image to obtain the replaced target action image;
[0141] An obtaining module 40, configured to obtain a human motion recognition result of the target user according to the replaced target motion image.
[0142] It should be noted that since the steps executed by the device in this embodiment are the same as those in the foregoing method embodiment, the specific implementation manners and the achievable technical effects can refer to the foregoing embodiment, and will not be elaborated herein.
[0143] The above are only optional embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural transformation made under the inventive concept of the present invention by using the content of the specification and drawings of the present invention, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present invention.
Claims
1. A method for recognizing human body movements, characterized in that, The method includes the following steps: Obtain a target action image including the human action of a target user; Obtain a standard action image corresponding to the target action image, where the human action in the standard action image has no occlusion information; Use the standard action image to perform local area replacement on the target action image to obtain a replaced target action image; Obtain a human action recognition result of the target user according to the replaced target action image; Before the step of obtaining the standard action image corresponding to the target action image, the method further includes: Divide the target action image into multiple divided areas; The step of obtaining the standard action image corresponding to the target action image includes: Obtain multiple standard action images corresponding to the multiple divided areas, with one divided area corresponding to one standard action image; The step of using the standard action image to perform local area replacement on the target action image includes: Use the standard action image corresponding to each divided area to perform local area replacement on each divided area to obtain multiple replaced divided areas; Obtain the replaced target action image according to the multiple replaced divided areas; The step of using the standard action image corresponding to each divided area to perform local area replacement on each divided area to obtain multiple replaced divided areas includes: Divide each divided area into multiple alternative blocks; In the multiple alternative blocks in each divided area, determine the selected alternative block of each divided area; Rotate the preset image corresponding to each divided area to obtain a rotated image corresponding to each divided area; Perform a scaling operation on the rotated image corresponding to each divided area according to the size of the selected alternative block of each divided area to obtain a scaled image corresponding to each divided area; Use the scaled image corresponding to each divided area to replace the selected alternative block in each divided area to obtain multiple replaced divided areas.
2. The method according to claim 1, characterized in that, The step of obtaining the human action recognition result of the target user according to the replaced target action image includes: Obtain multiple target coordinates corresponding to multiple key points according to the replaced target action image; Obtain multiple target joint angles according to the multiple target coordinates; Obtain multiple standard joint angles of the standard action image corresponding to the target action image; Obtain the human action recognition result of the target user according to the comparison result of the multiple target joint angles and the multiple standard joint angles.
3. The method according to claim 2, wherein The step of obtaining multiple target coordinates corresponding to multiple key points according to the replaced target action image includes: Generate a first image with a size of 4a*4a, a second image with a size of 2a*2a, and a third image with a size of a*a from the replaced target action image, where a is a natural number not equal to 0; Perform convolutional downsampling on the first image to obtain a fourth image, and perform convolutional downsampling on the second image to obtain a fifth image; Process the first image to obtain multiple first coordinates corresponding to multiple key points, and process the second image to obtain multiple second coordinates corresponding to multiple key points; Overlay the fourth image and the second image to obtain a sixth image, and overlay the third image and the fifth image to obtain a seventh image; Perform convolutional downsampling on the sixth image to obtain an eighth image; Process the eighth image to obtain multiple third coordinates corresponding to multiple key points, and process the seventh image to obtain multiple fourth coordinates corresponding to multiple key points; Obtain multiple target coordinates based on the multiple first coordinates, the multiple second coordinates, the multiple third coordinates, and the multiple fourth coordinates.
4. The method according to claim 3, wherein Before the step of obtaining multiple target coordinates based on the multiple first coordinates, the multiple second coordinates, the multiple third coordinates, and the multiple fourth coordinates, the method further includes: Determine a first weight corresponding to the multiple first coordinates, a second weight corresponding to the multiple second coordinates, a third weight corresponding to the multiple third coordinates, and a fourth weight corresponding to the multiple fourth coordinates according to the sizes of the first image, the second image, the third image, and the fourth image, wherein the first weight, the second weight, the third weight, and the fourth weight decrease in sequence; Obtain the target coordinate corresponding to each key point according to the first coordinate corresponding to each key point, the second coordinate corresponding to each key point, the third coordinate corresponding to each key point, the fourth coordinate corresponding to each key point, the first weight, the second weight, the third weight, and the fourth weight.
5. The method according to claim 2, characterized in that, The step of obtaining the human action recognition result of the target user according to the comparison result between the multiple target joint angles and the multiple standard joint angles includes: Calculate the angle difference corresponding to each joint according to the target joint angle corresponding to each joint and the standard joint angle corresponding to each joint; Obtain joint weights corresponding to the multiple joints according to the multiple angle differences, and the joint weights are positively correlated with the angle differences of the joints; Calculate the human action similarity by using the multiple angle differences corresponding to the multiple joints and the multiple joint weights corresponding to the multiple joints; Obtain the human action recognition result according to the comparison result between the human action similarity and a preset similarity threshold.
6. An apparatus for recognizing human body movements, characterized in that, The device includes: A first acquisition module for acquiring a target action image including the human action of the target user; A second acquisition module for acquiring a standard action image corresponding to the target action image, and there is no occlusion information in the human action in the standard action image; A replacement module for locally replacing the target action image with the standard action image to obtain a replaced target action image; An acquisition module for obtaining the human action recognition result of the target user according to the replaced target action image; The device is further configured to divide the target action image into multiple divided regions; The second acquisition module is further configured to divide the target action image into multiple divided regions; The replacement module is further configured to perform local area replacement on each of the divided regions by using the standard action images corresponding to each of the divided regions, so as to obtain a plurality of replaced divided regions; and obtain a replaced target action image according to the plurality of replaced divided regions. The replacement module is further configured to divide each of the divided regions into a plurality of alternative blocks; determine a selected alternative block for each of the divided regions from the plurality of alternative blocks in each of the divided regions; rotate the preset image corresponding to each of the divided regions to obtain a rotated image corresponding to each of the divided regions; perform a scaling operation on the rotated image corresponding to each of the divided regions according to the size of the selected alternative block of each of the divided regions to obtain a scaled image corresponding to each of the divided regions; and replace the selected alternative blocks in each of the divided regions by using the scaled image corresponding to each of the divided regions to obtain a plurality of replaced divided regions.
7. A terminal device, characterized in that, The terminal device includes: a memory, a processor, and a human body motion recognition program stored on the memory and running on the processor. When the human body motion recognition program is executed by the processor, the steps of the human body motion recognition method according to any one of claims 1 to 5 are implemented.
8. A storage medium, characterized in that, A human body motion recognition program is stored on the storage medium. When the human body motion recognition program is executed by a processor, the steps of the human body motion recognition method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Human brain visual memory principle-based human body action identification method and system
CN105023000A