A behavior recognition method, system, device and storage medium

Through the preset human posture estimation algorithm and feature extraction model, the joint nodes and dynamic and static combination features in the target video are extracted, and the problem of low behavior recognition accuracy in the prior art is solved, and high-precision recognition of complex and subtle behavioral actions is achieved.

CN114898248BActive Publication Date: 2025-06-20AGRICULTURAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210390053.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-14
Publication Date
2025-06-20
Estimated Expiration
2042-04-14

AI Technical Summary

Technical Problem

The prior art has low recognition accuracy for complex behavioral actions in behavior recognition, and methods based on bone sequences cannot accurately identify subtle behavioral actions or similar behavioral actions.

Method used

The preset human posture estimation calculation method is used to extract the node data of the target video, combined with the preset dynamic and static combination feature extraction model and the skeleton sequence feature extraction model, and extract dynamic and static combination feature vectors and skeleton sequence feature vectors to determine the behavioral status of the target character.

Benefits of technology

The accuracy of recognition of behavioral actions is improved, especially in identifying subtle, complex or similar actions, and the reduction of extraction accuracy caused by external factors in the prior art is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898248B_ABST
    Figure CN114898248B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a behavior recognition method, system, device, and storage medium. Among them, the method includes: using a preset human body pose estimation algorithm to extract joint point data from each frame of an image of a target video, obtaining a joint point data group of a target person in the target video; using the preset human body pose estimation algorithm to obtain a skeleton data group of the target person and a joint data group of the target person according to the joint point data group; based on a preset dynamic and static combined feature extraction model, extracting a dynamic and static combined feature vector group from the joint data group; using a preset skeleton sequence feature extraction model to extract feature vectors from the skeleton data group, obtaining a skeleton sequence feature vector of the target person; and determining the behavior state of the target person according to the skeleton sequence feature vector of the target person and the dynamic and static combined feature vector group. The present invention can improve the recognition accuracy of behavior actions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of behavior recognition, and particularly to a behavior recognition method, system, device, and storage medium. Background Art

[0002] Behavior recognition technology is a technology for recognizing and extracting the behavior characteristics of targets in videos. With the development of deep learning technology in recent years, deep learning has also begun to be applied to the field of behavior recognition technology to improve the accuracy of behavior recognition. Currently, behavior recognition methods based on deep learning are mainly divided into two categories, namely, spatio-temporal feature extraction methods based on multi-frame image sequences, and behavior recognition methods based on skeletal sequences.

[0003] However, for the spatio-temporal feature extraction method based on multi-frame image sequences, since its training data is easily affected by factors such as the complexity of the real scene, perspective changes, occlusion, etc., the generalization ability of the trained model in this method is insufficient, and complex behavior actions cannot be accurately recognized. For the behavior recognition method based on skeletal sequences, its recognition accuracy of behavior actions highly depends on the extraction accuracy of skeletal information. Moreover, existing skeletal information extraction methods cannot clearly express the relationship between joint points, making the existing behavior recognition methods based on skeletal sequences unable to accurately recognize subtle behavior actions or similar behavior actions. It can be seen that the recognition accuracy of behavior actions in the prior art is low. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a behavior recognition method, system, device, and storage medium to improve the recognition accuracy of behavior actions. The specific technical solutions are as follows:

[0005] A behavior recognition method, the method comprising:

[0006] Using a preset human pose estimation algorithm, extracting joint point data from each frame of the target video to obtain a joint point data group of the target person in the target video.

[0007] Using the preset human pose estimation algorithm, according to the joint point data group, obtaining a skeleton data group of the target person and a joint data group of the target person, wherein the skeleton data group includes skeleton image data in each frame of the target video, and the joint data group includes image data groups of multiple target joints, and each image data group of the target joint includes joint image data of the target joint in each frame of the target video.

[0008] Based on a preset dynamic and static combined feature extraction model, extract a dynamic and static combined feature vector group from the joint data group, where the dynamic and static combined feature vector group includes multiple dynamic and static combined feature vectors, and there is a one-to-one correspondence between the target joint and the dynamic and static combined feature vectors.

[0009] Use a preset skeleton sequence feature extraction model to extract feature vectors from the skeleton data group to obtain the skeleton sequence feature vector of the target person.

[0010] Determine the behavior state of the target person according to the skeleton sequence feature vector of the target person and the dynamic and static combined feature vector group.

[0011] Optionally, the preset dynamic and static combined feature extraction model is composed of a preset flow density algorithm, a preset target video sequence feature extraction model, and a preset dynamic and static feature aggregation algorithm. The process of extracting the dynamic and static combined feature vector group from the joint data group based on the preset dynamic and static combined feature extraction model includes:

[0012] For the image data group of each target joint in the joint data group: Use the preset flow density algorithm to perform optical flow extraction on the joint image data of each frame of the image data group of the target joint to obtain a target joint optical flow data group matching the image data group of the target joint. Among the target joint optical flow data group, it includes the target joint optical flow data corresponding to the joint image data in each frame of the target video. Use the preset target video sequence feature extraction model to respectively extract a first static feature vector group and a first dynamic feature vector group from the image data group of the target joint. Use the preset target video sequence feature extraction model to respectively extract a second static feature vector group and a second dynamic feature vector group from the target joint optical flow data group matching the image data group of the target joint. Use the preset dynamic and static feature aggregation algorithm to obtain a dynamic and static combined feature vector matching the image data group of the target joint according to the first static feature vector group, the first dynamic feature vector group, the second static feature vector group, and the second dynamic feature vector group.

[0013] Obtain the dynamic and static combined feature vector group matching the joint data group.

[0014] Optionally, the preset target video sequence feature extraction model is a first preset video sequence feature extraction model or a second preset video sequence feature extraction model. The training process of the preset target video sequence feature extraction model includes:

[0015] Use a first preset action data training set to pre-train an initial feature extraction model to obtain a first initial video sequence feature extraction model.

[0016] Perform sample extraction operations on the preset action recognition data set according to the preset number of frames to obtain training sample data. Use the training sample data and perform training and parameter tuning operations on the first initial video sequence feature extraction model according to the preset cross-view evaluation method to obtain the first preset video sequence feature extraction model.

[0017] Alternatively, pre-train the initial feature extraction model using the second preset action data training set to obtain a second initial video sequence feature extraction model.

[0018] Perform sample extraction operations on the preset action recognition data set according to the preset number of frames to obtain the training sample data. Use the training sample data and perform training and parameter tuning operations on the second initial video sequence feature extraction model according to the preset cross-view evaluation method to obtain the second preset video sequence feature extraction model.

[0019] Optionally, the use of the preset dynamic and static feature aggregation algorithm to obtain a dynamic and static combined feature vector that matches the target joint image data group according to the first static feature vector group, the first dynamic feature vector group, the second static feature vector group, and the second dynamic feature vector group includes:

[0020] According to the maximum value M of the target vector dimension i in the target static feature vector group i and the minimum value m i , through the formula:

[0021]

[0022] Obtain the target static action feature where the target static feature vector group is the first static feature vector group or the second static feature vector group, the target static action feature is the first static action feature or the second static action feature, the first static action feature matches the first static feature vector group, the second static action feature matches the second static feature vector group, p is the identifier of the target joint, stat is the static feature identifier, M i and the m i are obtained through the formula:

[0023]

[0024] where t is the image frame number identifier of the target video, T is the total number of image frames of the target video, is the value of the target static feature vector of the t-th frame image in the target joint image data group.

[0025] According to the The said M i and the said m i , through the formula:

[0026]

[0027] obtain the target dynamic action feature wherein, the target dynamic action feature is the first dynamic action feature or the second dynamic action feature, the first dynamic action feature matches the first dynamic feature vector group, the second dynamic action feature matches the second dynamic feature vector group, the dym is the dynamic feature identifier, wherein the Δm i and the ΔM i , through the formula:

[0028]

[0029] are obtained, where the Δt is the preset frame number interval.

[0030] Utilize the preset dynamic and static feature aggregation algorithm to splice the first static action feature and the first dynamic action feature on the target vector dimension i to obtain the depth image feature vector.

[0031] Utilize the preset dynamic and static feature aggregation algorithm to splice the second static action feature and the second dynamic action feature on the target vector dimension i to obtain the optical flow image feature vector.

[0032] Aggregate and regularize the depth image feature vector and the optical flow image feature vector to obtain the dynamic and static combined feature vector.

[0033] Optionally, determining the behavior state of the target person according to the skeleton sequence feature vector of the target person and the dynamic and static combined feature vector group includes:

[0034] Splice the skeleton sequence feature vector and the dynamic and static combined feature vector group to obtain the behavior feature vector of the target person.

[0035] Send the behavior feature vector to a preset behavior classifier so that the preset classifier determines the behavior label matching the feature vector and determines the content of the behavior label as the behavior state of the target person.

[0036] A behavior recognition system, the system includes:

[0037] A first data extraction module, which uses a preset human body pose estimation algorithm to extract joint point data from each frame of the target video to obtain the joint point data group of the target person in the target video.

[0038] The second data extraction module uses the preset human body pose estimation algorithm to obtain the skeleton data set of the target person and the joint data set of the target person according to the joint point data set. Among them, the skeleton data set includes the skeleton image data in each frame of the target video, and the joint data set includes the image data sets of multiple target joints. Each image data set of the target joint includes the joint image data of the target joint in each frame of the target video.

[0039] The first feature vector extraction module extracts a dynamic and static combined feature vector set from the joint data set based on a preset dynamic and static combined feature extraction model. Among them, the dynamic and static combined feature vector set includes multiple dynamic and static combined feature vectors, and there is a one-to-one correspondence between the target joints and the dynamic and static combined feature vectors.

[0040] The second feature vector extraction module uses a preset skeleton sequence feature extraction model to extract feature vectors from the skeleton data set to obtain the skeleton sequence feature vector of the target person.

[0041] The behavior recognition module is used to determine the behavior state of the target person according to the skeleton sequence feature vector of the target person and the dynamic and static combined feature vector set.

[0042] Optionally, the first feature vector extraction module is configured as follows:

[0043] For the image data sets of each target joint in the joint data set: use the preset flow density algorithm to perform optical flow extraction on the joint image data of each frame of the image data set of the target joint to obtain a target joint optical flow data set matching the image data set of the target joint. Among them, the target joint optical flow data set includes the target joint optical flow data corresponding to the joint image data in each frame of the target video. Use the preset target video sequence feature extraction model to respectively extract a first static feature vector set and a first dynamic feature vector set from the image data set of the target joint. Use the preset target video sequence feature extraction model to respectively extract a second static feature vector set and a second dynamic feature vector set from the target joint optical flow data set matching the image data set of the target joint. Use the preset dynamic and static feature aggregation algorithm to obtain a dynamic and static combined feature vector matching the image data set of the target joint according to the first static feature vector set, the first dynamic feature vector set, the second static feature vector set, and the second dynamic feature vector set.

[0044] Obtain the dynamic and static combined feature vector set matching the joint data set.

[0045] Optionally, the system further includes:

[0046] A model training module, which pre-trains the initial feature extraction model using the first preset action data training set to obtain the first initial video sequence feature extraction model.

[0047] Perform a sample extraction operation on the preset action recognition data set according to the preset number of frames to obtain training sample data, and use the training sample data to train and tune the parameters of the first initial video sequence feature extraction model according to the preset cross-view evaluation method to obtain the first preset video sequence feature extraction model.

[0048] Alternatively, pre-train the initial feature extraction model using the second preset action data training set to obtain the second initial video sequence feature extraction model.

[0049] Perform the sample extraction operation on the preset action recognition data set according to the preset number of frames to obtain the training sample data, and use the training sample data to train and tune the parameters of the second initial video sequence feature extraction model according to the preset cross-view evaluation method to obtain the second preset video sequence feature extraction model.

[0050] Optionally, the first feature vector extraction module is further configured to:

[0051] According to the maximum value M of the target vector dimension i in the target static feature vector group i and the minimum value m i , through the formula:

[0052]

[0053] Obtain the target static action feature wherein, the target static feature vector group is the first static feature vector group or the second static feature vector group, the target static action feature is the first static action feature or the second static action feature, the first static action feature matches the first static feature vector group, the second static action feature matches the second static feature vector group, p is the identifier of the target joint, stat is the static feature identifier, M i and the m i are obtained through the formula:

[0054]

[0055] wherein, t is the image frame number identifier of the target video, T is the total number of image frames of the target video, is the value of the target static feature vector of the t-th frame image in the image data group of the target joint.

[0056] According to the above The above M i and the above m i , through the formula:

[0057]

[0058] Obtain the target dynamic action feature wherein, the target dynamic action feature is the first dynamic action feature or the second dynamic action feature, the first dynamic action feature matches the first dynamic feature vector group, the second dynamic action feature matches the second dynamic feature vector group, the dym is a dynamic feature identifier, wherein, the Δm i and the ΔM i , through the formula:

[0059]

[0060] Obtain, where the Δt is a preset frame number interval.

[0061] Using the preset dynamic and static feature aggregation algorithm, splice the first static action feature and the first dynamic action feature on the target vector dimension i to obtain a depth image feature vector.

[0062] Using the preset dynamic and static feature aggregation algorithm, splice the second static action feature and the second dynamic action feature on the target vector dimension i to obtain an optical flow image feature vector.

[0063] Aggregate and regularize the depth image feature vector and the optical flow image feature vector to obtain the dynamic and static combined feature vector.

[0064] Optionally, the behavior recognition module is set to:

[0065] Splice the skeleton sequence feature vector and the dynamic and static combined feature vector group to obtain the behavior feature vector of the target person.

[0066] Send the behavior feature vector to a preset behavior classifier, so that the preset classifier determines the behavior label matching the feature vector, and determines the content of the behavior label as the behavior state of the target person.

[0067] A behavior recognition device, the device includes:

[0068] A processor;

[0069] A memory for storing the executable instructions of the processor.

[0070] Wherein the processor is configured to execute the instructions to implement the behavior recognition method as described in any one of the above.

[0071] A computer storage medium, when the instructions in the computer-readable storage medium are executed by the processor of the behavior recognition device, enables the device to execute the behavior recognition method as described in any one of the above.

[0072] A behavior recognition method, system, device and storage medium provided by an embodiment of the present invention, through a preset pose estimation algorithm, identifies a target person in a target video, and extracts a skeleton data group and a joint data group of the target person, so that the present solution improves the acquisition accuracy of data for behavior recognition compared with the prior art. At the same time, the present solution uses a preset dynamic and static combined feature extraction model to realize the recognition of the action relationship between joint points, thereby improving the recognition accuracy of subtle actions, complex actions and similar actions. Finally, the present invention extracts a skeleton sequence feature vector to represent the behavior of the target person, avoiding the risk of reduced extraction accuracy caused by external influences in the prior art and improving the accuracy of subsequent behavior recognition. It can be seen that the present invention can improve the recognition accuracy of behavior actions.

[0073] Of course, it is not necessary for any product or method implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0075] Figure 1 It is a flowchart of a behavior recognition method provided by an embodiment of the present invention;

[0076] Figure 2 It is a block diagram of a behavior recognition system provided by an optional embodiment of the present invention;

[0077] Figure 3 It is a block diagram of a behavior recognition device provided by another optional embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0078] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0079] An embodiment of the present invention provides a behavior recognition method, as Figure 1 shown, the behavior recognition method includes:

[0080] S101. Using a preset human body pose estimation algorithm, extract joint point data from each frame of the target video to obtain a joint point data group of the target person in the target video.

[0081] Among them, the above target video can be the sequential image data of the target scene collected by a depth camera that can output depth images (RGB color mode and DepthMap, RGB+D). Since the depth image can intuitively reflect the geometric shape and the coordinates of each pixel point in the actual scene, therefore, extracting joint point data based on the above depth image can improve the recognition efficiency and recognition accuracy of behavior recognition.

[0082] Optionally, in an alternative embodiment of the present invention, the above preset human body pose estimation algorithm can be an algorithm that recognizes the target person in the target video and uses joint point information to extract the joint point data of the target person from the target video. Among them, the above preset pose estimation algorithm can be an algorithm constructed based on open source human pose recognition (OpenPose). Since the OpenPose pose recognition algorithm uses the affinity field of human key points for pose estimation, the recognition accuracy of the target in the target video is improved in this solution compared with the existing pose estimation strategies.

[0083] S102. Using a preset human body pose estimation algorithm, according to the joint point data group, obtain a skeleton data group of the target person and a joint data group of the target person, where the skeleton data group includes skeleton image data in each frame of the target video, and the joint data group includes an image data group of multiple target joints, and each image data group of the target joint includes the joint image data of the target joint in each frame of the target video.

[0084] Optionally, in an alternative embodiment of the present invention, the data type in the above joint point data group can be a joint name label and a joint point pixel value. Among them, the above joint name label can be used to distinguish different joints. The joint point pixel value can be used for subsequent joint data extraction.

[0085] Optionally, in another alternative embodiment of the present invention, the above-described implementation manner of obtaining the skeleton data group may be:

[0086] It is set that the joints in the above joint point data group include: skull, shoulder joint, elbow joint, wrist joint, waist joint, hip joint, knee joint, ankle joint and foot joint.

[0087] Using a preset human pose estimation algorithm, select the coordinates of the joint points in the above joint point data group that can enclose the largest range. For the convenience of description, the above-selected joint points are respectively set as: skull, elbow joint, hip joint and foot joint. Then, with the coordinates of each of the above-selected joint points as the center and a preset extension distance as the radius, delimit the extraction range. And extract the joint point data identified within this extraction range. Since the data outside the above extraction range will not be extracted, and the above extension distance can be set according to the actual application scenario. Therefore, in the joint point data set obtained through the above extraction operation, the background image data in the target video is removed, making the accuracy of the data used for subsequent behavior recognition in this solution improved.

[0088] Optionally, in another alternative embodiment of the present invention, since the joint data group of the above target task includes the joint image data of each target joint in each frame image of the target video. Therefore, compared with the prior art that only realizes behavior recognition based on skeletal feature data, the present invention can realize the recognition of the action relationship between joint points. Thus, the recognition accuracy of subtle actions, complex actions and similar actions is improved. Furthermore, the accuracy of the final behavior recognition of the target person is improved.

[0089] Among them, the implementation manner of obtaining the joint image data of the target joint may be:

[0090] Based on the skeleton image data in each frame image of the obtained target video, determine the positions of each joint point in the skeleton image data according to the joint label content of each joint point data. And with the coordinates of the positions of each joint point as the center, set the joint extraction range. For example, if the target joint is the elbow joint, then according to the joint label content "elbow", search for the coordinates of the joint point position that matches this joint label content. With the coordinates of the joint point position of the elbow joint as the center and a ten-pixel extraction radius for joint extraction, the extracted joint image data can represent the position data of the elbow joint in this frame image.

[0091] S103. Based on a preset dynamic and static combined feature extraction model, extract a dynamic and static combined feature vector group from the joint data group, where the dynamic and static combined feature vector group includes multiple dynamic and static combined feature vectors, and there is a one-to-one correspondence between the target joint and the dynamic and static combined feature vector.

[0092] Optionally, in an alternative embodiment of the present invention, the above-mentioned dynamic and static combined feature vector may be a feature vector obtained by splicing a static feature vector and a dynamic feature vector. Among them, the above-mentioned static feature vector may characterize the action change of the target in the time domain, and the above-mentioned dynamic feature vector may characterize the change of the structural features of the target.

[0093] Since the prior art mainly realizes behavior recognition through static feature vectors, and static features cannot characterize the correlation between behavior actions and local joints, the prior art has low recognition accuracy for subtle, complex or similar actions. Therefore, by extracting the above-mentioned dynamic and static combined feature vector, the present invention can improve the recognition accuracy for subtle, complex or similar actions compared with the prior art.

[0094] S104. Use a preset skeleton sequence feature extraction model to extract feature vectors from the skeleton data group to obtain the skeleton sequence feature vector of the target person.

[0095] Optionally, in an alternative embodiment of the present invention, since the prior art extracts the behavior of the target person based on depth image data, its extraction accuracy is easily affected by factors such as camera perspective and scene interference objects. Therefore, the present invention extracts the skeleton sequence feature vector to characterize the behavior of the target person, avoiding the risk of reduced extraction accuracy caused by external influences in the prior art and improving the accuracy of subsequent behavior recognition.

[0096] It should be noted that in practical applications, the above-mentioned steps S103 and S104 as Figure 1 shown can be executed simultaneously or successively. The present invention does not overly limit the execution order of the above-mentioned steps S103 and S104.

[0097] S105. Determine the behavior state of the target person according to the skeleton sequence feature vector of the target person and the dynamic and static combined feature vector group.

[0098] It should be noted that in practical applications, there are various implementation manners of the above-mentioned step S105 as Figure 1 shown. Here, an example is provided:

[0099] Splice the above-mentioned skeleton sequence feature vector of the target person and the dynamic and static combined feature vector group to obtain a spliced feature vector.

[0100] Use a preset classification algorithm to match the above-mentioned spliced feature vector with the feature vector identifiers stored in the local database, and determine the behavior name of the successfully matched local feature vector identifier as the behavior state of the above-mentioned target person.

[0101] Through a preset pose estimation algorithm, the present invention identifies a target person in a target video and extracts a skeleton data set and a joint data set of the target person, so that the present solution improves the acquisition accuracy of data for behavior recognition compared with the prior art. At the same time, the present solution uses a preset dynamic and static combined feature extraction model to identify the action relationship between joint points, thereby improving the recognition accuracy of subtle actions, complex actions and similar actions. Finally, the present invention represents the behavior of the target person by extracting the skeleton sequence feature vector, avoiding the risk of reduced extraction accuracy caused by external influences in the prior art and improving the accuracy of subsequent behavior recognition. It can be seen that the present invention can improve the recognition accuracy of behavior actions.

[0102] Optionally, the preset dynamic and static combined feature extraction model is composed of a preset flow density algorithm, a preset target video sequence feature extraction model, and a preset dynamic and static feature aggregation algorithm. Based on the preset dynamic and static combined feature extraction model, a dynamic and static combined feature vector set is extracted from the joint data set, including:

[0103] For the image data set of each target joint in the joint data set: using the preset flow density algorithm, optical flow extraction is performed on the joint image data of each frame of the image data set of the target joint to obtain a target joint optical flow data set matching the image data set of the target joint. Among them, in the target joint optical flow data set, it includes the target joint optical flow data corresponding to the joint image data in each frame of the target video. Using the preset target video sequence feature extraction model, a first static feature vector set and a first dynamic feature vector set are respectively extracted from the image data set of the target joint. Using the preset target video sequence feature extraction model, a second static feature vector set and a second dynamic feature vector set are respectively extracted from the target joint optical flow data set matching the image data set of the target joint. Using the preset dynamic and static feature aggregation algorithm, according to the first static feature vector set, the first dynamic feature vector set, the second static feature vector set, and the second dynamic feature vector set, a dynamic and static combined feature vector matching the image data set of the target joint is obtained.

[0104] Obtain a dynamic and static combined feature vector set matching the joint data set.

[0105] Those skilled in the art can understand that for the construction and selection of the above preset flow density algorithm, the specific type of the above preset flow density algorithm can be determined through a cross-platform computer vision software library (Calc Optical Flow Pyr LK, Opencv), and the present invention will not elaborate too much here.

[0106] Optionally, in an alternative embodiment of the present invention, since the optical flow image data can represent the motion information of the target in the image. Therefore, the target joint optical flow data in each frame image obtained through the above optical flow extraction operation can reflect the motion information of the target joint in the target video.

[0107] Optionally, in another alternative embodiment of the present invention, the above-mentioned preset target video sequence feature extraction model may be a model constructed based on a two-stream architecture. Among them, the above two-stream architecture has two independent recognition streams in space and time respectively, and determines the final classification by fusing the output vectors of the two recognition streams. Since the model with a two-stream structure combines the control features and time features in the same image data, compared with the prior art method of solely recognizing time features, the above-mentioned preset target video sequence feature extraction model constructed based on the two-stream architecture can improve the accuracy of behavior recognition in each frame image.

[0108] It should be noted that the present invention does not overly limit the extraction order of the above-mentioned preset target video sequence feature extraction model for extracting the above-mentioned static feature vector groups and dynamic feature vector groups.

[0109] Optionally, the preset target video sequence feature extraction model is the first preset video sequence feature extraction model or the second preset video sequence feature extraction model. The training process of the preset target video sequence feature extraction model includes:

[0110] Using the first preset action data training set to pre-train the initial feature extraction model to obtain the first initial video sequence feature extraction model.

[0111] Performing a sample extraction operation on the preset action recognition data set according to a preset number of frames to obtain training sample data. Using the training sample data, training and parameter adjustment operations are performed on the first initial video sequence feature extraction model according to a preset cross-view evaluation method to obtain the first preset video sequence feature extraction model.

[0112] Or, using the second preset action data training set to pre-train the initial feature extraction model to obtain the second initial video sequence feature extraction model.

[0113] Performing a sample extraction operation on the preset action recognition data set according to a preset number of frames to obtain training sample data. Using the training sample data, training and parameter adjustment operations are performed on the second initial video sequence feature extraction model according to a preset cross-view evaluation method to obtain the second preset video sequence feature extraction model.

[0114] Optionally, in an alternative embodiment of the present invention, since the above-mentioned initial feature extraction model has generalization ability, when the generalization degree is high, it is likely to cause a decrease in the feature extraction accuracy of the trained preset target video sequence feature extraction model. Therefore, through the above-mentioned pre-training, the generalization degree of the preset target video sequence feature extraction model can be reduced, thereby improving the accuracy of feature extraction.

[0115] Those skilled in the art can understand that in practical applications, there are various selections for the above-mentioned first preset action data training set and second preset action data training set. For example: Computer Vision System Recognition Project (ImageNet) and Real Action Video Action Recognition Data Set (UCF-101).

[0116] Optionally, in another alternative embodiment of the present invention, the above-mentioned preset action recognition data set may be part or all of the data in the human skeleton data set (NTU RGB+D), and this data set includes 40 types of normal actions and 9 types of health-related actions.

[0117] Optionally, in another alternative embodiment of the present invention, since the action data in the above-mentioned preset action recognition data set is stored in the form of a video, therefore, the corresponding number of frames of data can be extracted from the above-mentioned preset action recognition data set through the above-mentioned preset number of frames to improve the accuracy of the training sample data. Among them, the above-mentioned preset number of frames can be determined according to the number of frames of the target video. For example, if the length of the target video is sixty frames, then the above-mentioned preset number of frames is set to sixty.

[0118] Optionally, in another alternative embodiment of the present invention, the specific implementation manner of the above-mentioned preset cross-view evaluation method may be:

[0119] Since the data in the above-mentioned preset action recognition data set is the data collected by three depth cameras at different angles, therefore, select the data collected by two of the depth cameras as training samples, and select the data collected by the remaining one depth camera as test samples for training and parameter adjustment, which can improve the extraction accuracy of the above-mentioned target video sequence feature extraction model. Optionally, using the preset dynamic and static feature aggregation algorithm, according to the first static feature vector group, the first dynamic feature vector group, the second static feature vector group and the second dynamic feature vector group, obtain the dynamic and static combined feature vector matching the target joint image data group, including: according to the maximum value M i and the minimum value m i , through the formula:

[0120]

[0121] obtain the target static action feature Among them, the target static feature vector group is the first static feature vector group or the second static feature vector group, the target static action feature is the first static action feature or the second static action feature, the first static action feature matches the first static feature vector group, the second static action feature matches the second static feature vector group, p is the identifier of the target joint, stat is the static feature identifier, M i and m i are obtained through the formula:

[0122]

[0123] where t is the image frame number identifier of the target video, T is the total number of image frames of the target video, is the value of the target static feature vector of the t-th frame image in the image data group of the target joint.

[0124] According to M i and m i , through the formula:

[0125]

[0126] the target dynamic action feature is obtained Among them, the target dynamic action feature is the first dynamic action feature or the second dynamic action feature, the first dynamic action feature matches the first dynamic feature vector group, the second dynamic action feature matches the second dynamic feature vector group, dym is the dynamic feature identifier, where Δm i and ΔM i , through the formula:

[0127]

[0128] are obtained, where Δt is the preset frame number interval.

[0129] Using the preset dynamic and static feature aggregation algorithm, the first static action feature and the first dynamic action feature are concatenated on the target vector dimension i to obtain the depth image feature vector.

[0130] Using the preset dynamic and static feature aggregation algorithm, the second static action feature and the second dynamic action feature are concatenated on the target vector dimension i to obtain the optical flow image feature vector.

[0131] The depth image feature vector and the optical flow image feature vector are aggregated and regularized to obtain the dynamic and static combined feature vector.

[0132] Optionally, in an alternative embodiment of the present invention, since the accuracy of the optical flow data in expressing the motion information of the target in the image is higher than that of the motion information of the target in the depth image data. Therefore, when splicing the depth image feature vector and the optical flow image feature vector, the weight value of the optical flow image feature vector can be increased to improve the data accuracy of the above-mentioned combined dynamic and static feature vector.

[0133] Optionally, in another alternative embodiment of the present invention, the process of aggregating and regularizing the depth image feature vector and the optical flow image feature vector to obtain the combined dynamic and static feature vector can splice the depth image feature vector and the optical flow image feature vector of the same frame image through a preset feature aggregation algorithm. Then, using a preset regularization algorithm to regularize the spliced feature vector to reduce the dimension and parameter value of the obtained combined dynamic and static feature vector. Thereby improving the recognition accuracy and efficiency of determining the behavior state of the target task subsequently.

[0134] Optionally, determining the behavior state of the target person according to the skeleton sequence feature vector and the combined dynamic and static feature vector group of the target person includes:

[0135] Splicing the skeleton sequence feature vector and the combined dynamic and static feature vector group to obtain the behavior feature vector of the target person.

[0136] Sending the behavior feature vector to a preset behavior classifier so that the preset classifier determines the behavior label matching the feature vector and determines the content of the behavior label as the behavior state of the target person.

[0137] Those skilled in the art can understand that in practical applications, there are various types of the above-mentioned preset behavior classifiers, such as Support Vector Machine (SVM) and Logistic Regression (LR). The present invention does not make excessive limitations on this.

[0138] Corresponding to the above method embodiment, the present invention also provides a behavior recognition system, as Figure 2 shown, the behavior recognition system includes:

[0139] The first data extraction module 201 extracts joint point data from each frame image of the target video by using a preset human body pose estimation algorithm to obtain a joint point data group of the target person in the target video.

[0140] The second data extraction module 202 uses a preset human body pose estimation algorithm to obtain a skeleton data group of the target person and a joint data group of the target person according to the joint point data group. The skeleton data group includes skeleton image data in each frame of the target video, and the joint data group includes an image data group of multiple target joints. Each image data group of the target joint includes joint image data of the target joint in each frame of the target video.

[0141] The first feature vector extraction module 203 extracts a dynamic and static combined feature vector group from the joint data group based on a preset dynamic and static combined feature extraction model. The dynamic and static combined feature vector group includes multiple dynamic and static combined feature vectors, and there is a one-to-one correspondence between the target joints and the dynamic and static combined feature vectors.

[0142] The second feature vector extraction module 204 uses a preset skeleton sequence feature extraction model to extract feature vectors from the skeleton data group to obtain the skeleton sequence feature vector of the target person.

[0143] The behavior recognition module 205 is used to determine the behavior state of the target person according to the skeleton sequence feature vector of the target person and the dynamic and static combined feature vector group.

[0144] Optionally, the above first feature vector extraction module 203 is set to:

[0145] For the image data group of each target joint in the joint data group: Use a preset optical flow density algorithm to perform optical flow extraction on the joint image data of each frame of the image data group of the target joint to obtain a target joint optical flow data group matching the image data group of the target joint. The target joint optical flow data group includes the target joint optical flow data corresponding to the joint image data in each frame of the target video. Use a preset target video sequence feature extraction model to extract a first static feature vector group and a first dynamic feature vector group from the image data group of the target joint respectively. Use a preset target video sequence feature extraction model to extract a second static feature vector group and a second dynamic feature vector group from the target joint optical flow data group matching the image data group of the target joint respectively. Use a preset dynamic and static feature aggregation algorithm to obtain a dynamic and static combined feature vector matching the image data group of the target joint according to the first static feature vector group, the first dynamic feature vector group, the second static feature vector group, and the second dynamic feature vector group.

[0146] Obtain a dynamic and static combined feature vector group matching the joint data group.

[0147] Optionally, the above behavior recognition system further includes:

[0148] The model training module uses the first preset action data training set to pre-train the initial feature extraction model and obtains the first initial video sequence feature extraction model.

[0149] Perform a sample extraction operation on the preset action recognition data set according to the preset number of frames to obtain training sample data. Use the training sample data and perform a training and parameter adjustment operation on the first initial video sequence feature extraction model according to the preset cross-view evaluation method to obtain the first preset video sequence feature extraction model.

[0150] Alternatively, use the second preset action data training set to pre-train the initial feature extraction model to obtain the second initial video sequence feature extraction model.

[0151] Perform a sample extraction operation on the preset action recognition data set according to the preset number of frames to obtain training sample data. Use the training sample data and perform a training and parameter adjustment operation on the second initial video sequence feature extraction model according to the preset cross-view evaluation method to obtain the second preset video sequence feature extraction model.

[0152] Optionally, the above first feature vector extraction module 203 is further configured as:

[0153] According to the maximum value M i and the minimum value m i of the target vector dimension i in the target static feature vector group, through the formula:

[0154]

[0155] Obtain the target static action feature where the target static feature vector group is the first static feature vector group or the second static feature vector group, the target static action feature is the first static action feature or the second static action feature, the first static action feature matches the first static feature vector group, the second static action feature matches the second static feature vector group, p is the identifier of the target joint, stat is the static feature identifier, M i and m i are obtained through the formula:

[0156]

[0157] where t is the image frame number identifier of the target video, T is the total number of image frames of the target video, is the value of the target static feature vector of the t-th frame image in the image data group of the target joint;

[0158] According to M i and m i , through the formula:

[0159]

[0160] Obtain the target dynamic action feature Among them, the target dynamic action feature is the first dynamic action feature or the second dynamic action feature. The first dynamic action feature matches the first dynamic feature vector group, and the second dynamic action feature matches the second dynamic feature vector group. dym is the dynamic feature identifier, where Δm i and ΔM i , through the formula:

[0161]

[0162] Obtained, where Δt is the preset frame number interval;

[0163] Using the preset dynamic and static feature aggregation algorithm, splice the first static action feature and the first dynamic action feature on the target vector dimension i to obtain the depth image feature vector;

[0164] Using the preset dynamic and static feature aggregation algorithm, splice the second static action feature and the second dynamic action feature on the target vector dimension i to obtain the optical flow image feature vector;

[0165] Aggregate and regularize the depth image feature vector and the optical flow image feature vector to obtain the dynamic and static combined feature vector.

[0166] Optionally, the above behavior recognition module 205 is set to:

[0167] Splice the skeleton sequence feature vector with the dynamic and static combined feature vector group to obtain the behavior feature vector of the target person;

[0168] Send the behavior feature vector to the preset behavior classifier, so that the preset classifier determines the behavior label matching the feature vector, and determines the content of the behavior label as the behavior state of the target person.

[0169] An embodiment of the present invention provides a behavior recognition device, as Figure 3 shown, the device includes:

[0170] Processor 301;

[0171] A memory 302 for storing executable instructions of the processor 301.

[0172] Among them, the processor 301 is configured to execute instructions to implement the behavior recognition method as described in any one of the above.

[0173] An embodiment of the present invention provides a computer storage medium. When the instructions in the computer-readable storage medium are executed by a processor of a behavior recognition device, the device can execute the behavior recognition method as described in any one of the above.

[0174] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one storage chip. The memory is an example of a computer-readable medium.

[0175] Computer-readable media include both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0176] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system, or computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0177] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0178] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content.

[0179] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A behavior recognition method, characterized in that, The method includes: Using a preset human pose estimation algorithm to extract joint point data from each frame image of the target video, and obtaining a joint point data group of the target person in the target video; Using the preset human pose estimation algorithm, according to the joint point data group, obtaining a skeleton data group of the target person and a joint data group of the target person, wherein the skeleton data group includes skeleton image data in each frame image of the target video, and the joint data group includes image data groups of multiple target joints, and each image data group of the target joint includes joint image data of the target joint in each frame image of the target video; Based on a preset dynamic and static combined feature extraction model, extracting a dynamic and static combined feature vector group from the joint data group, wherein the preset dynamic and static combined feature extraction model consists of a preset flow density algorithm, a preset target video sequence feature extraction model, and a preset dynamic and static feature aggregation algorithm group. The process of extracting the dynamic and static combined feature vector group includes: for each image data group of the target joint in the joint data group: using the preset flow density algorithm to perform optical flow extraction on the joint image data of each frame image in the image data group of the target joint, and obtaining a target joint optical flow data group matching the image data group of the target joint, wherein the target joint optical flow data group includes target joint optical flow data corresponding to the joint image data in each frame image of the target video; using the preset target video sequence feature extraction model to respectively extract a first static feature vector group and a first dynamic feature vector group from the image data group of the target joint; using the preset target video sequence feature extraction model to respectively extract a second static feature vector group and a second dynamic feature vector group from the target joint optical flow data group matching the image data group of the target joint; using the preset dynamic and static feature aggregation algorithm group to obtain a dynamic and static combined feature vector matching the image data group of the target joint according to the first static feature vector group, the first dynamic feature vector group, the second static feature vector group, and the second dynamic feature vector group, obtaining the dynamic and static combined feature vector group matching the joint data group, the dynamic and static combined feature vector group includes multiple dynamic and static combined feature vectors, and there is a one-to-one correspondence between the target joint and the dynamic and static combined feature vector; Using a preset skeleton sequence feature extraction model to perform feature vector extraction on the skeleton data group, and obtaining a skeleton sequence feature vector of the target person; Concatenating the skeleton sequence feature vector and the dynamic and static combined feature vector group to obtain a behavior feature vector of the target person; Sending the behavior feature vector to a preset behavior classifier, so that the preset classifier determines a behavior label matching the feature vector, and determining the content of the behavior label as the behavior state of the target person.

2. The method according to claim 1, characterized in that, The preset target video sequence feature extraction model is the first preset video sequence feature extraction model or the second preset video sequence feature extraction model. The training process of the preset target video sequence feature extraction model includes: Using a first preset action data training set to pre-train an initial feature extraction model to obtain a first initial video sequence feature extraction model; Performing a sample extraction operation on a preset action recognition data set according to a preset number of frames to obtain training sample data, and using the training sample data to train and tune the parameters of the first initial video sequence feature extraction model according to a preset cross-view evaluation method to obtain the first preset video sequence feature extraction model; Alternatively, using a second preset action data training set to pre-train the initial feature extraction model to obtain a second initial video sequence feature extraction model; Performing a sample extraction operation on the preset action recognition data set according to the preset number of frames to obtain the training sample data, and using the training sample data to train and tune the parameters of the second initial video sequence feature extraction model according to the preset cross-view evaluation method to obtain the second preset video sequence feature extraction model.

3. The method according to claim 1, characterized in that, The obtaining of the static and dynamic combined feature vector matching the target joint image data group according to the first static feature vector group, the first dynamic feature vector group, the second static feature vector group and the second dynamic feature vector group by using the preset static and dynamic feature aggregation algorithm group includes: According to the maximum value M of the target vector dimension i in the target static feature vector group i and the minimum value m i , through the formula: Obtain the target static action feature Wherein, the target static feature vector group is the first static feature vector group or the second static feature vector group, the target static action feature is the first static action feature or the second static action feature, the first static action feature matches the first static feature vector group, the second static action feature matches the second static feature vector group, p is the identifier of the target joint, stat is the static feature identifier, M i and the m i are obtained through the formula: obtained, where t is the image frame number identifier of the target video, and T is the total number of image frames of the target video, is the value of the target static feature vector of the t-th frame image in the image data group of the target joint; According to the said the said M i and the said m i , through the formula: Obtain the target dynamic action feature Among them, the target dynamic action feature is the first dynamic action feature or the second dynamic action feature. The first dynamic action feature matches the first dynamic feature vector group, and the second dynamic action feature matches the second dynamic feature vector group. The dym is a dynamic feature identifier, where the Δm i and the ΔM i , through the formula: Obtaining, where Δt is a preset frame number interval; Using the preset static and dynamic feature aggregation algorithm to splice the first static action feature and the first dynamic action feature on the target vector dimension i to obtain a depth image feature vector; Using the preset static and dynamic feature aggregation algorithm to splice the second static action feature and the second dynamic action feature on the target vector dimension i to obtain an optical flow image feature vector; Aggregating and regularizing the depth image feature vector and the optical flow image feature vector to obtain the static and dynamic combined feature vector.

4. A behavior recognition system, characterized in that, The system includes: A first data extraction module that uses a preset human body pose estimation algorithm to extract joint point data from each frame image of a target video to obtain a joint point data group of a target person in the target video; A second data extraction module that uses the preset human body pose estimation algorithm to obtain a skeleton data group of the target person and a joint data group of the target person according to the joint point data group, where the skeleton data group includes skeleton image data in each frame image of the target video, and the joint data group includes a group of image data of multiple target joints, and each group of image data of the target joint includes joint image data of the target joint in each frame image of the target video; The first feature vector extraction module extracts a combined dynamic and static feature vector group from the joint data group based on a preset combined dynamic and static feature extraction model. The preset combined dynamic and static feature extraction model consists of a preset flow density algorithm, a preset target video sequence feature extraction model, and a preset combined dynamic and static feature aggregation algorithm group. The process of extracting the combined dynamic and static feature vector group includes: for the image data group of each target joint in the joint data group: using the preset flow density algorithm to perform optical flow extraction on the joint image data of each frame of the image data group of the target joint, obtaining a target joint optical flow data group that matches the image data group of the target joint. In the target joint optical flow data group, it includes the target joint optical flow data corresponding to the joint image data in each frame of the target video; using the preset target video sequence feature extraction model to respectively extract a first static feature vector group and a first dynamic feature vector group from the image data group of the target joint; using the preset target video sequence feature extraction model to respectively extract a second static feature vector group and a second dynamic feature vector group from the target joint optical flow data group that matches the image data group of the target joint; using the preset combined dynamic and static feature aggregation algorithm group to obtain a combined dynamic and static feature vector that matches the image data group of the target joint according to the first static feature vector group, the first dynamic feature vector group, the second static feature vector group, and the second dynamic feature vector group, obtaining the combined dynamic and static feature vector group that matches the joint data group. The combined dynamic and static feature vector group includes multiple combined dynamic and static feature vectors, and there is a one-to-one correspondence between the target joint and the combined dynamic and static feature vector; The second feature vector extraction module uses a preset skeleton sequence feature extraction model to perform feature vector extraction on the skeleton data group to obtain the skeleton sequence feature vector of the target person; The behavior recognition module is used to splice the skeleton sequence feature vector and the combined dynamic and static feature vector group to obtain the behavior feature vector of the target person; sending the behavior feature vector to a preset behavior classifier, so that the preset classifier determines the behavior label that matches the feature vector, and determines the content of the behavior label as the behavior state of the target person.

5. The system according to claim 4, characterized in that, The system further includes: The model training module uses a first preset action data training set to pre-train the initial feature extraction model to obtain a first initial video sequence feature extraction model; Performing a sample extraction operation on a preset action recognition data set according to a preset number of frames to obtain training sample data, and using the training sample data to perform training and parameter adjustment operations on the first initial video sequence feature extraction model according to a preset cross-view evaluation method to obtain the first preset video sequence feature extraction model; Or, using a second preset action data training set to pre-train the initial feature extraction model to obtain a second initial video sequence feature extraction model; Perform sample extraction operations on the preset action recognition data set according to the preset number of frames to obtain the training sample data. Use the training sample data to train and tune the second initial video sequence feature extraction model according to the preset cross-view evaluation method to obtain the second preset video sequence feature extraction model.

6. A behavior recognition device, characterized in that, The device includes: A processor; A memory for storing instructions executable by the processor; Wherein the processor is configured to execute the instructions to implement the behavior recognition method according to any one of claims 1 to 3 above.

7. A computer storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the behavior recognition device, the device is enabled to execute the behavior recognition method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • A method for extract posture features and related device are disclosed

    CN109344803A

  • Action recognition method and device based on multi-feature fusion of key points, medium and equipment

    CN114299615A