A method, device and equipment for human action recognition

By analyzing the relationship between basic movement elements and body state transitions in human body images, basic action combinations are constructed, and the problem of insufficient model expansion in the existing technology is solved, and flexible action recognition is achieved to adapt to the needs of different industries.

CN115019343BActive Publication Date: 2025-07-04ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210668939.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-07-04
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

The models trained in the prior art lack scalability when recognizing human movements and cannot adapt to the action recognition needs of different industries. The model needs to be retrained to adapt to new movements.

Method used

By receiving multi-frame images, analyzing the basic elements of the action, determining the relationship between human parts and between components and targets, and using the predefined transition relationship between limb states and basic action combinations to identify custom actions.

Benefits of technology

It realizes flexible action recognition capabilities, can adapt to the action recognition needs of different industries, and improves the scope of application and recognition efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019343B_ABST
    Figure CN115019343B_ABST
Patent Text Reader

Abstract

The present application discloses a human motion recognition method, device and equipment. By receiving multiple frames of images including a human body, parsing the motion basic elements in each frame of image according to the pre-defined motion basic elements constituting an action; determining the association relationships between various parts of the human body and the association relationships between various parts of the human body and other targets according to the parsed motion basic elements, so as to obtain the limb states corresponding to the human body in each frame of image; determining the basic action matched by the conversion relationship between the limb states corresponding to the multiple frames of images according to the sampling time according to the pre-defined conversion relationship between the limb states corresponding to different basic actions; and obtaining the custom action corresponding to the matched basic action according to the basic action combinations corresponding to different custom actions. Through this method, action recognition can be performed according to the requirements of different industries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technologies, and particularly to a human action recognition method, device, and equipment. Background Art

[0002] Action recognition is an important part in the process of the industry's digital transformation. However, due to different industry requirements, there are many differences in the requirements for actions, and the requirements are very fragmented. For example, for content review of video data, action recognition is one part of the content review, which is used to filter video data involving violence; in terms of action skill training, after calculating and analyzing the motion data transmitted by a data acquisition device, it is necessary to obtain the position and posture information of the user during movement, etc., so as to provide a basis for the user to share data, obtain action guidance, etc.; in the service industry, the actions of service personnel will be recognized to determine whether the behavior of the personnel meets the industry requirements.

[0003] The existing technology is to receive video data with multiple frames of raw image data; sample from the raw image data to obtain target image data; recognize the actions that appear in the video data based on the global features of the target image data to obtain global actions; recognize the actions that appear in the video data based on the local features of the target image data to obtain local actions; and fuse the global actions and the local actions into the target actions that appear in the video data. The design key point is to train a specific model to extract global features and local features, and then recognize specific actions. The model trained by this method has no scalability. If a new action appears, the model needs to be retrained, and it is impossible to perform action recognition according to the requirements of different industries. Summary of the Invention

[0004] In order to solve the problem that in the prior art, when performing human action recognition by training a model, if a new action appears, the model needs to be retrained, and it is impossible to perform action recognition according to the requirements of different industries, this application provides a human action recognition method, device, and equipment.

[0005] In a first aspect, this application provides a human action recognition method, and the method includes:

[0006] Receiving multiple frames of images including a human body, and parsing the action basic elements in each frame of image according to the pre-defined action basic elements that make up an action;

[0007] Determining the association relationships between various parts of the human body and the association relationships between various parts of the human body and other targets according to the parsed action basic elements, so as to obtain the limb state corresponding to the human body in each frame of image;

[0008] Determine the basic action that matches the conversion relationship between the limb states corresponding to the multi-frame images according to the sampling time, based on the pre-defined conversion relationship between the limb states corresponding to different basic actions;

[0009] Obtain the custom action corresponding to the matched basic action according to the combination of basic actions corresponding to different custom actions.

[0010] In a possible implementation manner, the basic action elements that make up the action defined in advance include at least one of the following:

[0011] Human body parts corresponding to different parts of the human body, key points corresponding to different parts during human movement, key points corresponding to different positions during the movement of human body parts, target areas related to human movement, and target objects related to human movement.

[0012] In a possible implementation manner, according to the parsed basic action elements, determine the association relationship between human body parts and the association relationship between human body parts and other targets, and obtain the limb state corresponding to the human body in each frame of image, including:

[0013] According to the parsed basic action elements, determine the association relationship between human body parts, the distance and angle between human body parts and target objects, and the static relationship or dynamic change relationship of the human body trajectory.

[0014] In a possible implementation manner, train a limb state classification model in advance with different sample images and corresponding limb states;

[0015] Among them, by inputting each frame of image into the limb state classification model, parsing the basic action elements in each frame of image through the limb state classification model, and according to the parsed basic action elements, determining the association relationship between human body parts and the association relationship between human body parts and other targets, the limb state corresponding to the human body in each frame of image is obtained.

[0016] In a possible implementation manner, according to the parsed basic action elements, determine the association relationship between human body parts and the association relationship between human body parts and other targets, and obtain the limb state corresponding to the human body in each frame of image, including:

[0017] Obtain the logical combination of the association relationship between human body parts and the association relationship between human body parts and other targets corresponding to different pre-defined limb states;

[0018] According to the parsed basic action elements, determine the logical combination corresponding to the association relationship between human body parts and the association relationship between human body parts and other targets, and obtain the limb state corresponding to the human body in each frame of image.

[0019] In a possible implementation manner, the logical combination corresponding to the association relationship between each component of the human body and the association relationship between each component of the human body and other targets includes at least one of the following:

[0020] The positional relationship between the key points corresponding to different components during human movement and the human body components corresponding to different parts of the human body;

[0021] The positional relationship between the key points corresponding to different components during human movement and the target area related to the human movement or the target object related to the human movement;

[0022] Compared with the previous frame of image data, the moving direction and moving speed of the key points corresponding to different components during human movement;

[0023] The distance relationship between the key points corresponding to different components during human movement;

[0024] The relationship between the line segments corresponding to different components during human movement, the relationship includes the ratio between the line segment lengths, whether the line segments intersect, and the angle of intersection of the line segments, and the human body component line segments are constructed by any two pre-defined key points of the human body components;

[0025] Finger pointing, the pointing includes at least one of pointing upward, pointing downward, pointing left, and pointing right.

[0026] In a possible implementation manner, after obtaining the limb state corresponding to the human body in each frame of image, it further includes:

[0027] Determine the human body direction of the human body relative to the acquisition device in each frame of image, and obtain the limb states corresponding to different human body directions;

[0028] Where the human body direction includes: the front of the human body facing the acquisition device, the back of the human body facing the acquisition device, the human body facing the acquisition device and tilted, the human body facing away from the acquisition device and tilted, the left side of the human body facing the acquisition device, and the right side of the human body facing the acquisition device.

[0029] In a possible implementation manner, the conversion relationship between the limb states corresponding to different predefined basic actions includes at least one of the following:

[0030] Pre-define at least one limb state corresponding to each basic action, the conversion relationship corresponding to the at least one limb state, and the corresponding at least one accompanying limb state and the conversion relationship of the at least one accompanying limb state;

[0031] Among them, the limb state is obtained by analyzing the association relationship between each component of the human body, and the accompanying conversion relationship is obtained by analyzing the distance and angle between each component of the human body and the object, and the static relationship or dynamic change relationship of the human body trajectory.

[0032] In a possible implementation, the conversion relationship corresponding to the at least one limb state includes a continuous mode in which the limb states corresponding to multiple frames of images are one, and a state conversion mode in which when the limb states corresponding to multiple frames of images are multiple, they are switched in a predetermined order;

[0033] The conversion relationship corresponding to the at least one accompanying limb state includes a continuous mode in which the accompanying limb states corresponding to multiple frames of images are one, and a state conversion mode in which when the accompanying limb states corresponding to multiple frames of images are multiple, they are switched in a predetermined order.

[0034] In a possible implementation, the following method is used to determine the basic action combinations corresponding to different custom actions:

[0035] Determine at least one basic action corresponding to the custom action, and combine the at least one basic action according to the time sequence, which is the basic action combination corresponding to the custom action.

[0036] In a second aspect, the present application provides a human action recognition device, and the device includes:

[0037] An analysis module, configured to receive multiple frames of images including a human body, and analyze the action basic elements in each frame of image according to the pre-defined action basic elements constituting the action;

[0038] A limb state determination module, configured to determine the association relationship between each component of the human body and the association relationship between each component of the human body and other targets according to the analyzed action basic elements, and obtain the limb state corresponding to the human body in each frame of image;

[0039] A basic action determination module, configured to determine the basic action matched by the conversion relationship between the limb states corresponding to the multiple frames of images according to the sampling time according to the pre-defined conversion relationship between the limb states corresponding to different basic actions;

[0040] An action determination module, configured to obtain the custom action corresponding to the matched basic action according to the basic action combination corresponding to different custom actions.

[0041] In a third aspect, the present application provides a human action recognition device, and the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the methods in the first aspect above.

[0042] Fourthly, the present application provides a computer storage medium storing a computer program for causing a computer to execute the method in the first aspect above.

[0043] The present application provides a human body motion recognition method, device and equipment. By analyzing pre-defined basic motion elements, the limb state corresponding to the human body in each frame of image is determined. Then, according to the pre-defined conversion relationship between limb states, the basic motion matching the conversion relationship between limb states in multiple frames of images is determined. Finally, according to the combination of basic motions corresponding to different custom motions, the custom motion corresponding to the matched basic motion is obtained. According to the above motion recognition method, the recognition of custom motions can be supported, with high flexibility and a wide range of applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 FIG. is a schematic diagram of an application scenario of a human body motion recognition method according to an exemplary embodiment of the present invention;

[0045] Figure 2 FIG. is a schematic flowchart of a human body motion recognition method according to an exemplary embodiment of the present invention;

[0046] Figure 3 FIG. is a schematic flowchart of custom motion construction according to an exemplary embodiment of the present invention;

[0047] Figure 4 FIG. is a schematic diagram of limb state construction according to an exemplary embodiment of the present invention;

[0048] Figure 5 FIG. is a schematic diagram of basic motion construction according to an exemplary embodiment of the present invention;

[0049] Figure 6 FIG. is a schematic diagram of custom motion construction according to an exemplary embodiment of the present invention;

[0050] Figure 7 FIG. is a schematic diagram of custom motion construction according to an exemplary embodiment of the present invention;

[0051] Figure 8 FIG. is a schematic diagram of a human body motion recognition device according to an exemplary embodiment of the present invention;

[0052] Figure 9 FIG. is a schematic diagram of a human body motion recognition device according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0053] The technical solutions in the embodiments of the present application will be clearly and elaborately described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] As Figure 1 shown in the schematic diagram of the application scenario of an action recognition method provided by an embodiment of the present application, the application scenario includes: a server 101, a database 102, and at least one collection device (the collection devices 103_1, 103_2, and 103_N shown in the figure). Among them, the collection device is used to collect multiple frames of images including a human body and send them to the server 101. The model of the collection device is not limited here, as long as multiple frames of images can be obtained. The server 101 is used to receive the multiple frames of images sent by the collection device and perform human action recognition. The database 102 is used to store the programs and data required for the human action recognition method.

[0055] Action recognition is an important part in the process of the digital transformation of the industry. However, due to different industry requirements, there are many differences in the requirements for actions, and the requirements are very fragmented. For example, for content review of video data, action recognition is a part of content review, which is used to filter video data involving violence; in terms of action skill training, after calculating and analyzing the motion data transmitted by the data collection device, the position and posture information of the user during movement can be obtained, so as to provide a basis for the user to share data, obtain action guidance, etc.; in the service industry, the actions of service personnel will be recognized to determine whether the behavior of the personnel meets the industry requirements.

[0056] However, the existing method of action recognition by training a classification model has no scalability and cannot perform action recognition according to the requirements of different industries. To address the problems existing in the prior art, an embodiment of the present application provides a human action recognition method. As Figure 2 shown, the method includes:

[0057] S201: Receive multiple frames of images including a human body, and parse the action basic elements in each frame of image according to the pre-defined action basic elements that make up the action.

[0058] Among them, the multiple frames of images are collected by a collection device that can obtain multiple frames of images. The action basic elements are pre-defined according to user requirements. In the embodiments provided by the present application, the pre-defined action basic elements that make up the action include at least one of the following:

[0059] Human body parts corresponding to different parts of the human body, key points corresponding to different parts during human body movements, key points corresponding to different positions during the movements of human body parts, target areas related to human body movements, and target objects related to human body movements.

[0060] The basic elements of the movement can include any one or any combination of the above, and can also include other types, such as other animals related to human body movements. Specifically, it can be expanded according to user needs or different industry needs, and no specific limitation is made here.

[0061] S202: According to the parsed basic elements of the movement, determine the association relationships between various human body parts and the association relationships between various human body parts and other targets, and obtain the limb states corresponding to the human body in each frame of the image.

[0062] After parsing multiple frames of images to obtain the basic elements of the movement included in the multiple frames of images, parse the association relationships between various human body parts and the association relationships between various human body parts and other targets. The association relationships can be, but are not limited to, pre-defined.

[0063] The limb state is defined as the association relationship between various human body parts, or the distance and angle between various human body parts and the target object, the static relationship or dynamic change relationship of the human body trajectory. Among them, the static relationship of the human body trajectory is that in the current frame of the image, the human body has a tendency to approach or move away from the target object, and the dynamic relationship of the human body trajectory is that compared with the previous frame of the image, the human body approaches or moves away from the target object in the current frame of the image.

[0064] S203: According to the conversion relationships between the limb states corresponding to different basic movements pre-defined, determine the basic movement matched by the conversion relationships between the limb states corresponding to the multiple frames of images according to the sampling time.

[0065] Among them, the basic movement is constructed by limb movements. In one possible implementation manner, the conversion relationships between the limb states corresponding to different pre-defined basic movements include at least one of the following:

[0066] Pre-define at least one limb state corresponding to each basic movement, and the conversion relationship corresponding to the at least one limb state, and the corresponding at least one accompanying limb state and the conversion relationship of the at least one accompanying limb state;

[0067] Among them, parse the association relationships between various human body parts to obtain the limb state, and parse the distance and angle between various human body parts and the object, the static relationship or dynamic change relationship of the human body trajectory to obtain the accompanying conversion relationship.

[0068] S204: According to the combination of basic movements corresponding to different custom movements, obtain the custom movement corresponding to the matched basic movement.

[0069] In a possible implementation, the following method is used to determine the basic action combinations corresponding to different custom actions:

[0070] Determine at least one basic action corresponding to the custom action, and combine the at least one basic action in chronological order to obtain the basic action combination corresponding to the custom action.

[0071] As Figure 3 shown, the human action recognition method provided in the embodiments of the present application first constructs the limb state by pre-defining the basic elements of the action, constructs the basic action through the pre-defined conversion relationship between the limb states, combines the basic actions in chronological order to obtain the custom action, and then recognizes the human action in multiple frames of images according to the process of constructing the custom action. If a new action appears, the action can be defined according to the above process of action customization, so as to supplement the action categories in the entire human action recognition system, and thus meet the human action recognition needs of different industries.

[0072] The basic elements of the action in S201 above include at least one of the human body parts corresponding to different parts of the human body, the key points corresponding to different parts during the human action, the key points corresponding to different positions when the human body parts move, the target area related to the human action, and the target object related to the human action.

[0073] Among them, the human body parts corresponding to different parts of the human body include: head, hands, feet, upper body, etc.; the key points corresponding to different parts during the human action include: wrists, elbows, knees, head, feet, etc.; the key points corresponding to different positions when the human body parts move can be when the hand moves, and its key points include: the tip of the thumb, the tip of the index finger, etc.; the target area related to the human action includes: the room where the human action is located, the street, the pre-defined space range, etc.; the target object related to the human action includes: teapot, book, cup, table, chair, etc. The above-listed basic elements of the action can be specifically adjusted according to user needs and actual situations.

[0074] The limb states in S202 above include two sources:

[0075] (1) The limb state classification model.

[0076] In a possible implementation, the limb state classification model is pre-trained through different sample images and the corresponding limb states;

[0077] Among them, by inputting each frame of image into the limb state classification model, parsing the basic elements of the action in each frame of image through the limb state classification model, and determining the association relationship between the human body parts and the association relationship between the human body parts and other targets according to the parsed basic elements of the action, the limb state corresponding to the human body in each frame of image is obtained.

[0078] The limb state classification model includes a human body state model and a hand state model. The human body state model includes: standing upright, lying prone, squatting on the ground, sitting on a seat, lying on a table, crossing legs, etc.

[0079] The hand state model includes: OK, giving a thumbs up, gesture 1, gesture 2, gesture 3, gesture 5, making a fist, etc.

[0080] The above-listed models can be specifically increased according to user needs and actual situations.

[0081] (2) The logical combination of the association relationships between various components of the human body and the association relationships between various components of the human body and other targets.

[0082] In a possible implementation manner, according to the parsed basic action elements, determine the association relationships between various components of the human body and the association relationships between various components of the human body and other targets, and obtain the corresponding limb state of the human body in each frame of image, including:

[0083] Obtain the logical combination of the association relationships between various components of the human body corresponding to different limb states and the association relationships between various components of the human body and other targets defined in advance;

[0084] According to the parsed basic action elements, determine the logical combination corresponding to the association relationships between various components of the human body and the association relationships between various components of the human body and other targets, and obtain the corresponding limb state of the human body in each frame of image.

[0085] In a possible implementation manner, the logical combination corresponding to the association relationships between various components of the human body and the association relationships between various components of the human body and other targets includes at least one of the following:

[0086] The positional relationship between the key points corresponding to different components during human movement and the human body components corresponding to different parts of the human body;

[0087] The positional relationship between the key points corresponding to different components during human movement and the target area related to the human movement, or between the key points corresponding to different components during human movement and the target object related to the human movement;

[0088] Compared with the previous frame of image data, the moving direction and moving speed of the key points corresponding to different components during human movement;

[0089] The distance relationship between the key points corresponding to different components during human movement;

[0090] The relationship between the line segments corresponding to different components during human movement, and the relationship includes the ratio between the line segment lengths, whether the line segments intersect, and the intersection angle of the line segments. The human body component line segments are constructed by any two predefined human body component key points.

[0091] Finger pointing, where the pointing includes at least one of upward pointing, downward pointing, leftward pointing, and rightward pointing.

[0092] The positional relationship between the key points corresponding to different components during the above human body movements and the human body components corresponding to different parts of the human body, as well as the positional relationship between the key points corresponding to different components during the human body movements and the target area related to the human body movements or the target object related to the human body movements can be simply described as the "relationship between points and frames"; compared with the previous frame of image data, the moving direction and moving speed of the key points corresponding to different components during the human body movements can be simply described as the "movement of points"; the distance relationship between the key points corresponding to different components during the human body movements can be simply described as the "relationship between points"; the relationship between the line segments corresponding to different components during the human body movements can be simply described as the "relationship between line segments".

[0093] Sort the logical combinations of the association relationships between the human body components corresponding to different limb states and the association relationships between the human body components and other targets predefined according to the preset table items, where the table items include association relationships, associated parties, relationship types, and corresponding parameters.

[0094] Specifically, as shown in Table 1:

[0095] Table 1

[0096]

[0097] Among them, the X-axis and Y-axis of the screen in Table 1 refer to the pixel coordinates generated by the acquisition device when acquiring images. Generally, the upper left corner of the image is the origin, and moving up / down / left / right refers to the movement direction relative to the human body, and the horizontal / vertical / screen distance refers to the movement in the pixel coordinates or the position in the pixel coordinates.

[0098] For the relationship between points, it is necessary to normalize their distances. Specifically, first obtain the actual distance between two key points, and then divide the actual distance by the predefined reference size to obtain the normalized distance between the two key points. The role of normalization is to make the absolute value of the physical system numerical value become a certain relative value relationship, which can simplify calculations and reduce the magnitude.

[0099] Considering the perspective relationship of the camera, the limb state is constructed by independently combining in different directions on the basis of the above association relationships.

[0100] In a possible implementation manner, after obtaining the limb state corresponding to the human body in each frame of image, it further includes:

[0101] Determine the human body direction of the human body relative to the acquisition device in each frame of image, and obtain the limb states corresponding to different human body directions;

[0102] Among them, the human body directions include: the front of the human body facing the acquisition device, the front and back of the human body facing the acquisition device, the human body facing the acquisition device and tilted, the human body facing away from the acquisition device and tilted, the left side of the human body facing the acquisition device, and the right side of the human body facing the acquisition device. During the process of human action recognition, the human body direction is obtained by analyzing the image. Under different human body directions, the analyzed limb states are different. For example, the limb state of the hand cannot be obtained when the front and back of the human body face the acquisition device. Therefore, when constructing the limb state in advance, the human body direction needs to be considered.

[0103] Such as Figure 4 shown, the limb state can be a combination of "AND" and "OR" of the above-mentioned association relationships. For example, when the association relationship A occurs and at the same time the association relationship B or the component relationship C occurs, the limb state is AB + AC.

[0104] Figure 4 It contains two columns, each column represents an "AND" association relationship, and the relationship between columns represents an "OR" relationship. Fill the two association relationships A and B in the first column, and fill the two association relationships A and C in the second column. Together they form a limb state, and different camera directions are also combined by an "OR" relationship. The specific relationship types, associated parties, and parameters of the association relationships A, B, and C need to be debugged and generated according to the actual business scenario.

[0105] The basic actions in the above S203 are specifically constructed through the following implementation methods:

[0106] The basic actions are composed of conversion relationships such as the continuation, switching, and accompaniment of limb states.

[0107] In a possible implementation manner, the conversion relationship corresponding to the at least one limb state includes a continuous mode in which the limb states corresponding to multiple frames of images are one, and a state conversion mode in which when the limb states corresponding to multiple frames of images are multiple, they are switched in a predetermined order;

[0108] The conversion relationship of the at least one accompanying limb state includes a continuous mode in which the accompanying limb states corresponding to multiple frames of images are one, and a state conversion mode in which when the accompanying limb states corresponding to multiple frames of images are multiple, they are switched in a predetermined order.

[0109] The specific combination method is such as Figure 5 shown, the limb state can only include one row, and can include one or more limb states. If only one limb state is included, it is a state continuous mode, such as Figure 5 in, only the limb state A is included during the time period Ta, such as crossing one's legs. The time period Ta includes multiple frames of images.

[0110] If it contains multiple limb states and switches in a predetermined order, it is a state transition mode. Figure 5 In the figure, the limb state is A during the Ta period and switches to B during the Tb period. Falling is caused by the transition between standing upright and falling down. Tab indicates the process of switching from A to B.

[0111] A line of accompanying pattern can be added to the continuous mode or transition mode of the limb state. Figure 5 In the example, during the Ta time period, the body state of sitting on a chair occurs while crossing the legs. The Ta time period includes multiple frames of images, each of which includes the body state A and its corresponding accompanying body state E. It can also be a series of multiple body states, adding an accompanying body state, such as Figure 5 The limb states B and C correspond to the same accompanying limb state F. There is also a conversion mode for the accompanying limb state. Figure 5 In the time period from Ta to Tc, the accompanying limb state switches from E to accompanying limb state F.

[0112] For complex business needs, you can define coherent and complex actions through the timing combination of basic actions. The specific construction process is as follows: Figure 6 shown.

[0113] Considering the intuitiveness and convenience of user configuration, the action definition includes the selection of basic actions and the related modification restrictions of basic actions. The modification content includes "posture", "body direction", "limbs", etc., and can also include other body parts such as "head", "shoulders", etc., which can be set according to actual conditions. "Posture" can be the posture of the human body, such as squatting, lying, lying prone, standing, sitting, bending over, etc. Human body direction, such as front, back, side, etc. "Limbs" are defined as human limbs, such as left hand, right hand, both hands, left foot, right foot, both feet, etc. In addition to the above modification conditions, it can be other modification conditions, such as top or oblique camera mounting. Dress modification, such as police uniform, nurse uniform, etc. For different industries, the same limb state may represent different basic actions. For example, in the service industry, when the limb state is bending over and reaching out, it may be handing something to the guest. For example, when the limb state of a nurse is bending over and reaching out, it may be giving an infusion to the patient.

[0114] The cross-legged action is a combination of two basic actions: vertical leg placement and cross-legged leg placement. Figure 7As shown in the figure, first construct the basic action of putting the foot upright: select the sitting posture, the front / side direction of the human body, and the limbs and feet to form the limb state, and construct the basic action of putting the foot upright through the limb state; then construct the basic action of crossing the legs: select the sitting posture and the front / side direction of the human body to form the limb state, and construct the basic action of crossing the legs through the limb state; finally, combine the basic action of putting the foot upright and the basic action of crossing the legs according to the time sequence to obtain the custom action of crossing the legs.

[0115] In order for each basic action to include these modification states, it is necessary to limit and configure the modification amount that the basic action can support, that is, to combine the basic actions. By customizing actions in the above way, the user's degree of freedom of use is improved. The user can design relevant actions by themselves through the provided elements, which has reusability, improves the reuse efficiency of the system, and reduces the maintenance cost.

[0116] Based on the same inventive concept, an embodiment of the present application provides a human action recognition device 800, as Figure 8 shown, the device includes:

[0117] An analysis module 801, configured to receive multiple frames of images including a human body, and analyze the action basic elements in each frame of image according to the pre-defined action basic elements for composing actions;

[0118] A determining limb state module 802, configured to determine the association relationship between each part of the human body and the association relationship between each part of the human body and other targets according to the analyzed action basic elements, and obtain the limb state corresponding to the human body in each frame of image;

[0119] A determining basic action module 803, configured to determine the basic action matched by the conversion relationship between the limb states corresponding to the multiple frames of images according to the pre-defined conversion relationship between the limb states corresponding to different basic actions;

[0120] A determining action module 804, configured to obtain the custom action corresponding to the matched basic action according to the basic action combination corresponding to different custom actions.

[0121] In a possible implementation manner, the analysis module 801 is configured to determine that the action basic elements for composing actions defined in advance include at least one of the following:

[0122] Human body parts corresponding to different parts of the human body, key points corresponding to different parts during human actions, key points corresponding to different positions during the actions of human body parts, target areas related to human actions, and target objects related to human actions.

[0123] In a possible implementation, the limb state determination module 802 is configured to determine the association relationships between various human body components and the association relationships between various human body components and other targets based on the parsed basic action elements, and obtain the limb state corresponding to the human body in each frame of image, including:

[0124] Based on the parsed basic action elements, determine the association relationships between various human body components, the distances and angles between various human body components and the target object, and the static relationships or dynamic change relationships of the human body trajectory.

[0125] In a possible implementation, the limb state determination module 802 is configured to pre-train a limb state classification model with different sample images and corresponding limb states;

[0126] Wherein, by inputting each frame of image into the limb state classification model, parsing the basic action elements in each frame of image through the limb state classification model, and based on the parsed basic action elements, determining the association relationships between various human body components and the association relationships between various human body components and other targets, the limb state corresponding to the human body in each frame of image is obtained.

[0127] In a possible implementation, the limb state determination module 802 is configured to determine the association relationships between various human body components and the association relationships between various human body components and other targets based on the parsed basic action elements, and obtain the limb state corresponding to the human body in each frame of image, including:

[0128] Obtain the logical combinations of the association relationships between various human body components corresponding to different pre-defined limb states and the association relationships between various human body components and other targets;

[0129] Based on the parsed basic action elements, determine the logical combinations corresponding to the association relationships between various human body components and the association relationships between various human body components and other targets, and obtain the limb state corresponding to the human body in each frame of image.

[0130] In a possible implementation, the limb state determination module 802 is configured to determine the logical combinations corresponding to the association relationships between various human body components and the association relationships between various human body components and other targets, including at least one of the following:

[0131] The positional relationships between the key points corresponding to different components during human body movement and the human body components corresponding to different parts of the human body;

[0132] The positional relationships between the key points corresponding to different components during human body movement and the target area related to the human body movement, or between the key points and the target object related to the human body movement;

[0133] The moving directions and moving speeds of key points corresponding to different parts during human body movements compared with the previous frame of image data;

[0134] The distance relationships between key points corresponding to different parts during human body movements;

[0135] The relationships between line segments corresponding to different parts during human body movements, where the relationships include the ratio between the lengths of line segments, whether the line segments intersect, and the angle of intersection of the line segments. The human body part line segments are constructed by any two predefined key points of human body parts;

[0136] Finger pointing, where the pointing includes at least one of pointing upward, pointing downward, pointing left, and pointing right.

[0137] In a possible implementation manner, after the determining limb state module 802 is used to obtain the limb state corresponding to the human body in each frame of image, it further includes:

[0138] Determine the human body direction of the human body relative to the acquisition device in each frame of image, and obtain the limb states corresponding to different human body directions;

[0139] Where the human body direction includes: the human body facing the acquisition device directly, the human body facing away from the acquisition device directly, the human body facing the acquisition device and tilting, the human body facing away from the acquisition device and tilting, the left side of the human body facing the acquisition device, and the right side of the human body facing the acquisition device.

[0140] In a possible implementation manner, the determining basic action module 803 is used for the conversion relationships between the limb states corresponding to different predefined basic actions, including at least one of the following:

[0141] Predefine at least one limb state corresponding to each basic action, and the conversion relationship corresponding to the at least one limb state, and the corresponding at least one accompanying limb state and the conversion relationship of the at least one accompanying limb state;

[0142] Among them, the limb state is obtained by analyzing the association relationships between various parts of the human body, and the accompanying conversion relationship is obtained by analyzing the distance and angle between various parts of the human body and objects, and the static relationship or dynamic change relationship of the human body trajectory.

[0143] In a possible implementation manner, the determining basic action module 803 is used to determine the conversion relationship corresponding to the at least one limb state, including a continuous mode where the limb states corresponding to multiple frames of images are one, and a state conversion mode where when the limb states corresponding to multiple frames of images are multiple, they are switched in a predetermined order;

[0144] The at least one conversion relationship of the accompanying limb states includes a continuous mode in which the accompanying limb state corresponding to multiple frames of images is one, and a state conversion mode in which the accompanying limb states corresponding to multiple frames of images are multiple and are switched in a predetermined order.

[0145] In a possible implementation manner, the determining action module 804 is configured to determine the basic action combinations corresponding to different custom actions in the following manner:

[0146] Determine at least one basic action corresponding to the custom action, and combine the at least one basic action in time sequence, which is the basic action combination corresponding to the custom action.

[0147] Based on the same inventive concept, the present application provides a human action recognition device, and the device includes:

[0148] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any one of the human action recognition methods in the embodiments of the present application.

[0149] Next, refer to Figure 9 to describe the electronic device 130 according to this embodiment of the present application. Figure 9 The displayed electronic device 130 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0150] As Figure 9 shown, the electronic device 130 is presented in the form of a general electronic device. The components of the electronic device 130 may include but are not limited to: the above-mentioned at least one processor 131, the above-mentioned at least one memory 132, and a bus 133 connecting different system components (including the memory 132 and the processor 131).

[0151] The processor 131 is configured to read and execute the instructions in the memory 132 so that the at least one processor can execute a human action recognition method provided in the above embodiment.

[0152] The bus 133 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a processor, or a local bus using any bus structure in a variety of bus structures.

[0153] The memory 132 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 1321 and / or a cache memory 1322, and may further include a read-only memory (ROM) 1323.

[0154] The memory 132 may also include a program / utilities 1325 having a set (at least one) of program modules 1324. Such program modules 1324 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0155] The electronic device 130 may also communicate with one or more external devices 134 (such as a keyboard, a pointing device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 130, and / or may communicate with any device that enables the electronic device 130 to communicate with one or more other electronic devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 135. Also, the electronic device 130 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 136. As shown in the figure, the network adapter 136 communicates with other modules for the electronic device 130 through a bus 133. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 130, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0156] In some possible implementation manners, various aspects of a human motion recognition method provided in this application may also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a human motion recognition according to various exemplary implementation manners of this application described above in this specification.

[0157] In addition, this application also provides a computer-readable storage medium, and the computer storage medium stores a computer program, and the computer program is used to cause a computer to execute the method described in any one of the above embodiments.

[0158] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.

[0159] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, causing a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or boxes Figure 1 one process or a plurality of processes and / or boxes Figure 1 and / or steps for implementing the functions specified in one box or a plurality of boxes.

[0160] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0161] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A human motion recognition method, characterized in that, The method includes: Receiving multiple frames of images including a human body, and parsing the action basic elements in each frame of image according to the pre-defined action basic elements of the composed actions; Determining the association relationships between the various parts of the human body and the association relationships between the various parts of the human body and other targets according to the parsed action basic elements, and obtaining the limb states corresponding to the human body in each frame of image; Determining the basic action matched by the conversion relationship between the limb states corresponding to the multiple frames of images according to the sampling time according to the pre-defined conversion relationship between the limb states corresponding to different basic actions. The pre-defined conversion relationship between the limb states corresponding to different basic actions includes at least one limb state corresponding to each pre-defined basic action, the conversion relationship corresponding to the at least one limb state, the corresponding at least one accompanying limb state, and part or all of the conversion relationship of the at least one accompanying limb state. Among them, parsing the association relationship between the various parts of the human body to obtain the limb state, and parsing the distance and angle between the various parts of the human body and the object, the static relationship or dynamic change relationship of the human body trajectory to obtain the accompanying conversion relationship; Obtaining the custom action corresponding to the matched basic action according to the combination of basic actions corresponding to different custom actions. The different custom actions are formed by selecting the corresponding basic actions and the relevant modification restrictions of the basic actions.

2. The method according to claim 1, wherein The pre-defined action basic elements of the composed actions include at least one of the following: Human body parts corresponding to different parts of the human body, key points corresponding to different parts during human actions, key points corresponding to different positions when human body parts move, target areas related to human actions, and target objects related to human actions.

3. The method according to claim 1, characterized in that, Determining the association relationships between the various parts of the human body and the association relationships between the various parts of the human body and other targets according to the parsed action basic elements, and obtaining the limb states corresponding to the human body in each frame of image, including: Determining the association relationships between the various parts of the human body, the distance and angle between the various parts of the human body and the target object, and the static relationship or dynamic change relationship of the human body trajectory according to the parsed action basic elements.

4. The method according to any one of claims 1 to 3, characterized in that, It also includes: Pre-training a limb state classification model with different sample images and the corresponding limb states; Among them, by inputting each frame of image into the limb state classification model, parsing the action basic elements in each frame of image through the limb state classification model, and determining the association relationships between the various parts of the human body and the association relationships between the various parts of the human body and other targets according to the parsed action basic elements, and obtaining the limb states corresponding to the human body in each frame of image.

5. The method according to any one of claims 1 to 3, characterized in that, Determining the association relationships between the various parts of the human body and the association relationships between the various parts of the human body and other targets according to the parsed action basic elements, and obtaining the limb states corresponding to the human body in each frame of image, including: Obtaining the logical combination of the association relationships between the various parts of the human body and the association relationships between the various parts of the human body and other targets corresponding to different pre-defined limb states; Determining the logical combination corresponding to the association relationships between the various parts of the human body and the association relationships between the various parts of the human body and other targets according to the parsed action basic elements, and obtaining the limb states corresponding to the human body in each frame of image.

6. The method according to claim 5, characterized in that The logical combinations corresponding to the association relationships between various components of the human body and the association relationships between various components of the human body and other targets include at least one of the following: The positional relationship between the key points corresponding to different components during human movement and the human body components corresponding to different parts of the human body; The positional relationship between the key points corresponding to different components during human movement and the target area related to the human movement, or between the key points corresponding to different components during human movement and the target object related to the human movement; The moving direction and moving speed of the key points corresponding to different components during human movement compared with the previous frame of image data; The distance relationship between the key points corresponding to different components during human movement; The relationship between the line segments corresponding to different components during human movement, the relationship including the ratio between the line segment lengths, whether the line segments intersect, and the angle of intersection of the line segments, and the human body component line segments are constructed by any two pre-defined key points of the human body components; Finger pointing, which includes at least one of upward pointing, downward pointing, leftward pointing, and rightward pointing.

7. The method according to claim 2, wherein After obtaining the limb state corresponding to the human body in each frame of image, it further includes: Determining the human body direction of the human body relative to the acquisition device in each frame of image, and obtaining the limb state corresponding to different human body directions; Wherein the human body direction includes: the front of the human body facing the acquisition device, the back of the human body facing the acquisition device, the human body facing the acquisition device and tilted, the human body facing away from the acquisition device and tilted, the left side of the human body facing the acquisition device, and the right side of the human body facing the acquisition device.

8. The method according to claim 1, wherein The conversion relationship corresponding to the at least one limb state includes a continuous mode in which the limb states corresponding to multiple frames of images are one, and a state conversion mode in which when the limb states corresponding to multiple frames of images are multiple, they are switched in a predetermined order; The conversion relationship corresponding to the at least one accompanying limb state includes a continuous mode in which the accompanying limb states corresponding to multiple frames of images are one, and a state conversion mode in which when the accompanying limb states corresponding to multiple frames of images are multiple, they are switched in a predetermined order.

9. The method according to claim 1, characterized in that The basic action combination corresponding to different custom actions is determined in the following manner: Determining at least one basic action corresponding to the custom action, and combining the at least one basic action in chronological order to obtain the basic action combination corresponding to the custom action.

10. A human motion recognition device, characterized in that, The device includes: An analysis module, configured to receive multiple frames of images including a human body, and analyze the action basic elements in each frame of image according to the pre-defined action basic elements constituting the action; A limb state determination module, configured to determine the association relationship between various components of the human body and the association relationship between various components of the human body and other targets according to the analyzed action basic elements, and obtain the limb state corresponding to the human body in each frame of image; Determine the basic action module, which is used to determine the basic action matched by the conversion relationship between the limb states corresponding to the multi-frame images according to the sampling time, according to the conversion relationship between the limb states corresponding to different predefined basic actions. The conversion relationship between the limb states corresponding to the different predefined basic actions includes at least one limb state corresponding to each predefined basic action, the conversion relationship corresponding to the at least one limb state, at least one accompanying limb state corresponding thereto, and part or all of the conversion relationship of the at least one accompanying limb state. Among them, the limb state is obtained by analyzing the association relationship between the components of the human body, and the accompanying conversion relationship is obtained by analyzing the distance and angle between the components of the human body and the object, the static relationship or the dynamic change relationship of the human body trajectory; Determine the action module, which is used to obtain the custom action corresponding to the matched basic action according to the combination of basic actions corresponding to different custom actions. The different custom actions are formed by selecting the corresponding basic actions and the related modification restrictions of the basic actions.

11. A human motion recognition device, characterized in that, The device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-9.

12. A computer storage medium, characterized in that, The computer storage medium stores a computer program, and the computer program is used to cause a computer to execute the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Method and device for comparing action in video, storage medium and electronic equipment

    CN113569753A

  • Behavior recognition system and behavior recognition method

    WO2018159542A1