Behavior recognition method based on time-space matching degree of key points of people and objects
By constructing the space-time matrix of key points and interactive animals and calculating the similarity, the problem of low accuracy in behavior recognition in the prior art is solved, and high accuracy behavior recognition in complex scenarios is achieved.
Patent Information
- Application Number
- CN202510041321.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-10
AI Technical Summary
The accuracy of existing behavior recognition technologies has decreased in high noise environments and lacks sufficient modeling of the relative position and movement of objects and human bodies, resulting in poor recognition of complex interactive behaviors.
A behavior recognition method based on the spatial and temporal matching degree of key points between people and objects is proposed. By preprocessing the input video, the timing information of key points of human body and interactive animals is extracted, the space-time matrix is constructed, and the similarity between the space-time matrix and the timing template is calculated to identify behavioral actions.
It improves the accuracy of behavior recognition in complex scenarios, enhances the stability of action recognition, reduces the situation of misjudgment and fuzzy recognition, and can effectively perform action recognition in diverse scenarios.
Smart Images

Figure CN120032422A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision technology, and in particular to a behavior recognition method based on the spatiotemporal matching of key points between people and objects. Background Art
[0002] Existing behavior recognition technologies mainly rely on video analysis to identify behaviors by extracting human motion or posture features. Common methods include extracting spatial features based on convolutional neural networks (CNN), processing temporal information with recurrent neural networks (RNN) and long short-term memory networks (LSTM), and posture estimation methods based on human key point detection. These methods are widely used in monitoring, smart home, sports analysis and other fields, and can accurately identify basic behavior patterns. Most existing methods focus on the analysis of human posture or motion, while ignoring the relative relationship between the human body and surrounding objects, lack effective spatiotemporal information modeling, and cannot fully capture the impact of the interaction between objects and the human body on behavior.
[0003] Although the existing technology has achieved certain success in some scenarios, it still has certain shortcomings. First, the existing methods are easily interfered in high-noise environments, resulting in a decrease in the accuracy of behavior recognition. Secondly, the lack of sufficient modeling of the relative position and movement of objects and human bodies leads to poor results in the recognition of complex interactive behaviors. The existing technology generally lacks support for standardized action libraries, making it difficult to achieve high-precision action classification in some specific application scenarios. Therefore, there is an urgent need for a new technical solution that can effectively solve these problems and improve the ability to recognize behavior in complex scenarios. Summary of the invention
[0004] The present invention aims to at least solve the technical problem of low accuracy of behavior recognition caused by ignoring the association between interactive objects and human behavior in the prior art, and particularly innovatively proposes a behavior recognition method based on the spatiotemporal matching of key points between people and objects.
[0005] In order to achieve the above-mentioned object of the present invention, the present invention provides a behavior recognition method based on the spatiotemporal matching degree of key points of people and objects, comprising the following steps:
[0006] S1, preprocessing the input video;
[0007] S2, converts the spatiotemporal information of human body movements into the temporal information of key points of the human body and the interactive objects;
[0008] S3, constructing a space-time matrix from the time series information of the fixed time window;
[0009] S4, calculating the similarity between the spatiotemporal matrix and the temporal template;
[0010] S5, selecting the time sequence template with the highest similarity, the action corresponding to the time sequence template is the behavior action of the space-time matrix.
[0011] Preferably, the specific steps of step S1 are:
[0012] S11, processing the input video into a single-frame image through Opencv;
[0013] S12, performing denoising and filtering on the image processed by S11;
[0014] S13, performing center cropping on the preprocessed image and adjusting the size to a required size to standardize the image size.
[0015] Preferably, the specific steps of step S2 are:
[0016] S21, obtaining coordinates of key points of the human body and coordinates of interactive objects from the preprocessed video; wherein the coordinates of the key points of the human body are obtained through the yolo-pose model, and the coordinates of the interactive objects are obtained through the yolo object detection model;
[0017] S22, calculating the relative distance between each key point and the interactive object, wherein the relative distance is obtained by Euclidean distance calculation, and the calculated distance is stored for subsequent use.
[0018] Preferably, the step S2 further includes:
[0019] If multiple human bodies are detected, the person with the highest confidence is selected as the target;
[0020] If the target is detected to have multiple interactive objects, the interactive object with the greatest confidence is selected.
[0021] Preferably, the maximum confidence level includes:
[0022] For the confidence of a person, the confidence is first calculated using formula (1). If the confidence value exceeds the set threshold, it is determined to be a person. Then, the confidence of the key points is calculated using formula (2), and the maximum confidence is selected.
[0023] For interactive objects, the confidence is calculated directly using formula (1), and the maximum confidence is selected;
[0024] Formula (1) is as follows:
[0025] Conf=P(object)*IoU(pred,gt)
[0026] Among them, Conf represents the final confidence of each target box;
[0027] P(object) is the probability that an object exists in the prediction box;
[0028] IoU(pred,gt) is the maximum IoU value between the predicted box and all real boxes;
[0029] Formula (2) is as follows:
[0030]
[0031] Among them, σ is the Sigmoid function, which is used to normalize the probability of whether the key point exists;
[0032] It represents the predicted value of the key point output by the network, and can be obtained as a value between 0 and 1 through the Sigmoid function;
[0033] Preferably, the specific steps of step S3 are:
[0034] S31, the relative distances between different key points and the interactive objects are in different sequences, as the columns of the matrix;
[0035] S32, for information whose sequence length is greater than the fixed time window, equal-interval sampling is adopted, and for information whose sequence length is less than the fixed time window, linear interpolation is adopted to meet the set time window length requirement;
[0036] S33, combining the data processed in step S32 to obtain a space-time matrix, wherein the rows of the space-time matrix represent time frames, and the columns of the space-time matrix represent the distances between different key points and the interactive object in different time frames.
[0037] Preferably, the specific steps of step S4 are:
[0038] S41, based on the standard action, generating a timing template corresponding to each action, wherein the timing template is in a matrix form, wherein rows represent time frames, and columns represent distances between different key points and the interactive object in different time frames; these timing templates contain the time-space trajectory information of the key points of the human body and the interactive object;
[0039] S42, performing a convolution operation on the time series template and the space-time matrix through the same filter to obtain a characteristic signal of the time series template and a characteristic signal of the space-time matrix;
[0040] S43, calculating the norms of the two characteristic signals;
[0041] S44, normalizing the two characteristic signals after the norm calculation, and then calculating the cross-correlation of the two signals; preferably, the calculation formula of the cross-correlation is:
[0042]
[0043] Where r(i) represents the cross-correlation of the signal;
[0044] n is the signal length;
[0045] x(i) represents the value of the normalized timing template signal when the signal position is i;
[0046] y(t+i) represents the value of the normalized space-time matrix signal when the signal position is t+i;
[0047] t represents the time delay between signals x(i) and y(i);
[0048] i is the index, traversing the samples of x(i) and y(i). For each t value, i starts from the starting index of the signal to the end position of the signal; the maximum value of t cannot exceed n-1. For actions with long duration, the t value should be selected to be larger, otherwise smaller.
[0049] Preferably, the key points of the human body are joint points of the human body. The key points of the human body generally refer to joint points, and may also be other types of points defined by user.
[0050] In summary, due to the adoption of the above technical solution, the present invention adopts a pure mathematical calculation method for human behavior recognition. Compared with the use of a neural network model for processing, the method of the present invention has the advantages of optimizing the calculation process, reducing processing delays, improving calculation speed, and saving calculation resources. In addition, the key points of the human body are not limited to specific nodes, but cover a wider range of features. Moreover, the present invention can make full use of the correlation between interactive objects and behavioral actions, and convert the interactive relationship between people and objects into the features required by the model, thereby significantly improving the accuracy of behavioral action recognition. Specifically: by comparing with standard action data, the present invention can accurately classify behavioral actions according to the degree of match between the current action and the standard action. This approach not only enhances the stability of action recognition, but also effectively avoids misjudgment and fuzzy recognition. Compared with traditional methods, the present invention no longer relies solely on the training data of the model for judgment, but introduces a standard action library, making the recognition of unknown actions more reliable. The method can deeply analyze the relative movement of the human body and objects, providing a more accurate basis for action recognition. In complex scenes, the present invention can accurately capture tiny motion differences, thereby further improving the accuracy of action recognition. At the same time, due to the construction method of the spatiotemporal matrix, the model can flexibly adapt to different input data and scene changes, and can effectively perform action recognition regardless of complex backgrounds, different lighting conditions, or interference between dynamic objects. This adaptability and robustness make up for the shortcomings of existing models in diverse scenarios, making the present invention have significant advantages in the field of behavioral action recognition.
[0051] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0053] Figure 1 It is a flow chart of the present invention.
[0054] Figure 2 This is the data processing result in this case. DETAILED DESCRIPTION
[0055] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0056] The present invention first pre-processes the input video so that the subsequent model recognition has a higher accuracy, then extracts the corresponding coordinates of the key points recognized in each frame, calculates and records the key point coordinates, records the coordinates of the object, calculates and counts the distance between the interactive object and the center point of the human body, uses the counted distance as a space-time matrix, and calculates it with the data of the obtained standard action to obtain the matching degree between the current action and the standard action for action classification, and obtains the result of current behavior action recognition.
[0057] The present invention proposes a behavior recognition method based on the spatiotemporal matching of key points of people and objects. Figure 1 As shown, the following steps are included:
[0058] S1, preprocessing the input video;
[0059] S11, processing the input video into a single-frame image through Opencv;
[0060] S12, denoising and filtering the image processed by S11
[0061] S13. Perform center cropping on the preprocessed image and adjust the size to a required size to standardize the image size.
[0062] S2, converts the spatiotemporal information of human body movements into the temporal information of key points of the human body and interactive objects,
[0063] S21, obtaining coordinates of key points of the human body and coordinates of interactive objects from the preprocessed video; wherein the coordinates of the key points of the human body are obtained through the yolo-pose model, and the coordinates of the interactive objects are obtained through the yolo object detection model;
[0064] Specifically: the key point detection model outputs the skeleton key point sequence of the person detected in the video in the order of N, T, V, C according to the set key point sequence. In addition, it also outputs the recognition result of human behavior; N represents the maximum number of detected objects in the video, T represents the video frame length, V represents the number of key points, and C represents the number of channels, i.e., the horizontal coordinate, the vertical coordinate, and the confidence level. In the present invention, considering that the person will not only interact with the interactive objects in different scenes, but also with other people, N is set to 2. When there is only one person in the video or only one person can be detected, the other object cannot be detected, and 0 is output at the corresponding position of the key point;
[0065] This involves the issue of primary and secondary targets. When two or more people are detected at the same time, how to choose? The solution adopted in the present invention is to sort by human confidence, with the person with the highest confidence as the primary target and the person with the second highest confidence as the secondary target. Only the primary and secondary targets are retained, and the remaining targets are regarded as irrelevant persons and are not output. The target detection model outputs the center point sequence of the detected interactive objects. Here, there will also be a situation where multiple targets are detected in one frame of image. For this situation, the present invention selects the interactive object with the highest confidence to output;
[0066] The calculation of confidence involves three aspects: one is the score of the object box, that is, the evaluation of whether it contains the target object; the second is the probability that the target box is judged to be a "person" category; the third is the detection of key points in the box, among which the confidence of the key points will directly affect the final confidence assessment of the person.
[0067] Specifically, the calculation of confidence relies on the following formula:
[0068] Conf=P(object)*IoU(pred,gt)
[0069] Among them, Conf represents the final confidence of each target box;
[0070] P(object) is the probability that an object exists in the prediction box;
[0071] IoU(pred,gt) is the maximum IoU value between the predicted box and all real boxes;
[0072] The confidence of interactive objects can be calculated using the above formula. For the confidence of people, the confidence is first calculated using the above formula. If the confidence value exceeds the set threshold, it is determined to be a person. Then the confidence of the key points is calculated using the following formula. The confidence of the key points is:
[0073]
[0074] Among them, σ is the Sigmoid function, which is used to normalize the probability of whether the key point exists;
[0075] It represents the predicted value of the key point output by the network, and can be obtained as a value between 0 and 1 through the Sigmoid function;
[0076] S22, calculating the acquired coordinates and simultaneously calculating the relative distances between the multiple key points and the interactive object;
[0077] Specifically, the distance is calculated using the Euclidean distance, and the specific formula is: All distances are relative distances to avoid errors caused by viewing angle jitter. In addition to using Euclidean distance to calculate the relative distance between key points and interactive objects, other distance calculation methods can also be used, such as Manhattan distance, Chebyshev distance, cosine similarity, etc.
[0078] Among them, (x 1 ,y 1 ) is the coordinate of the interactive object in the video, (x 0 ,y 0 ) is the coordinate information of the key points;
[0079] S23, storing the calculated distance for subsequent use;
[0080] S3, constructing a matrix of the time series information of the fixed time window;
[0081] S31, the relative distances between different key points and the interactive objects are in different sequences, as the columns of the matrix;
[0082] S32, for information whose sequence length is greater than the fixed time window, equal interval sampling is adopted, and for information whose sequence length is less than the fixed time window, linear interpolation is adopted to meet the set requirements;
[0083] Specifically, equal-interval sampling is to select samples at fixed intervals while also taking into account the representativeness of the data; linear interpolation means that when data points are sparse, linear interpolation can be used to fill in missing data;
[0084] The linear interpolation formula is as follows:
[0085]
[0086] where y is the estimated value at x;
[0087] is the slope between two known points;
[0088] S33, combining the data processed in step S32 to obtain a space-time matrix, wherein the rows of the space-time matrix represent time frames, and the columns of the space-time matrix represent the distances between different key points and the interactive object in different time frames; the data format of the space-time matrix is a matrix of m rows × n columns, and the values of the matrix are shown in the following table.
[0089]
[0090] Among them, the rows represent the time frames, and the columns represent the distances of different key points to the interactive objects in different time frames;
[0091] S34, finally processing the data into an arrangement according to the relative distance data between the key points and the interactive objects;
[0092] S4, classify the temporal information using template matching algorithm;
[0093] S41, based on the standard actions, generating a timing template corresponding to each action, wherein the timing template is in a matrix form, wherein rows represent time frames, and columns represent distances between different key points and the interactive object in different time frames; these templates contain the time-space trajectory information of the key points of the human body and the interactive object;
[0094] S42, performing a convolution operation on the time series template and the space-time matrix through the same filter to obtain a characteristic signal of the time series template and a characteristic signal of the space-time matrix;
[0095] S43, calculating the norms of the two characteristic signals;
[0096] The Euclidean norm (L2 norm) for a vector v = [v 1 ,v 1 ,....,v n-1 ,v n ] is defined as:
[0097]
[0098] In terms of the choice of norm, since the Euclidean norm has better smoothness, the L2 norm is more suitable for the current signal processing task than the L1 norm;
[0099] S44, normalizing the two characteristic signals after the norm calculation, and then calculating the cross-correlation of the two signals;
[0100] Specifically, the normalized feature signals of x(t) and y(t) are the template signal and the target signal, respectively;
[0101] Specifically, in this case, the normalization method uses the Min-Max normalization method to normalize the data range to 0-1;
[0102] The following formula is used to calculate the cross-correlation:
[0103]
[0104] Among them, r(i) represents the cross-correlation of the signal. The larger r(i) is, the greater the matching degree of the signal is.
[0105] t represents the time delay between signals x(i) and y(i), i is the time point index of the signal, traverses the samples of x(i) and y(i), and for each t value, i starts from the starting index of the signal until the end position of the signal;
[0106] The process of calculating cross-correlation is as follows Figure 2 As shown in the figure: x(i) represents the template signal (blue curve), y(i) represents the target signal (green curve), and the signal is shifted by sliding comparison with the template signal and changing the t value. In the cross-correlation calculation, the calculation process corresponds to the iteration of t in the formula. When t=0, the starting points of the template signal and the target signal are aligned. Then, as t increases, the target signal gradually shifts to the right to match the other parts of the template signal.
[0107] The red curve is the matching degree change curve. The horizontal axis represents the offset t, and the vertical axis represents the cumulative matching degree (i.e., the cross-correlation value). When the curve reaches a peak, it means that the template signal and the target signal have reached the best alignment state at a certain offset t, and the matching degree is the highest.
[0108] Assuming the signal length is n, the maximum value of t cannot exceed n-1. For actions with a long duration, the t value should be larger, otherwise smaller.
[0109] S5, classify the actions according to similarity;
[0110] The present invention provides a behavior recognition method based on the spatiotemporal matching of key points between people and objects, so as to solve the problem of low behavior recognition accuracy caused by ignoring the association between interactive objects and human behaviors in the prior art.
[0111] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A behavior recognition method based on the spatiotemporal matching of key points between people and objects, characterized in that: The following steps are involved: S1, preprocessing the input video; S2, converts the spatiotemporal information of human body movements into the temporal information of key points of the human body and the interactive objects; S3, constructing a space-time matrix from the time series information of the fixed time window; S4, calculating the similarity between the spatiotemporal matrix and the temporal template; S5, selecting the time sequence template with the highest similarity, the action corresponding to the time sequence template is the behavior action of the space-time matrix.
2. According to claim 1, a behavior recognition method based on the spatiotemporal matching of key points between people and objects is characterized in that: The specific steps of step S1 are: S11, processing the input video into a single-frame image through Opencv; S12, performing denoising and filtering on the image processed by S11; S13, performing center cropping on the preprocessed image and adjusting the size to a required size to standardize the image size.
3. The behavior recognition method based on the spatiotemporal matching of key points between people and objects according to claim 1 is characterized in that: The specific steps of step S2 are: S21, obtaining coordinates of key points of the human body and coordinates of interactive objects from the preprocessed video; wherein the coordinates of the key points of the human body are obtained through the yolo-pose model, and the coordinates of the interactive objects are obtained through the yolo object detection model; S22, calculating the relative distance between each key point and the interactive object.
4. The behavior recognition method based on the spatiotemporal matching of key points between people and objects according to claim 1 is characterized in that: The step S2 further comprises: If multiple human bodies are detected, the person with the highest confidence is selected as the target; If the target is detected to have multiple interactive objects, the interactive object with the greatest confidence is selected.
5. The method for behavior recognition based on the spatiotemporal matching of key points between people and objects according to claim 4, characterized in that: The maximum confidence level includes: For the confidence of a person, the confidence is first calculated using formula (1). If the confidence value exceeds the set threshold, it is determined to be a person. Then, the confidence of the key points is calculated using formula (2), and the maximum confidence is selected. For interactive objects, the confidence is calculated directly using formula (1), and the maximum confidence is selected; Formula (1) is as follows: Conf=P(object)*IoU(pred,gt) Among them, Conf represents the final confidence of each target box; P(object) is the probability that an object exists in the prediction box; IoU(pred,gt) is the maximum IoU value between the predicted box and all real boxes; Formula (2) is as follows: Among them, σ is the Sigmoid function, which is used to normalize the probability of whether the key point exists; is the key point prediction value.
6. The behavior recognition method based on the spatiotemporal matching of key points between people and objects according to claim 1 is characterized in that: The specific steps of step S3 are: S31, the relative distances between different key points and the interactive objects are in different sequences, as the columns of the matrix; S32, for information whose sequence length is greater than the fixed time window, equal-interval sampling is adopted, and for information whose sequence length is less than the fixed time window, linear interpolation is adopted to meet the set time window length requirement; S33, combining the data processed in step S32 to obtain a space-time matrix, wherein the rows of the space-time matrix represent time frames, and the columns of the space-time matrix represent the distances between different key points and the interactive object in different time frames.
7. The method for behavior recognition based on the spatiotemporal matching of key points between people and objects according to claim 1, characterized in that: The specific steps of step S4 are: S41, based on the standard action, generating a timing template corresponding to each action, wherein the timing template is in a matrix form, wherein rows represent time frames, and columns represent distances between different key points and the interactive object in different time frames; S42, performing a convolution operation on the time series template and the space-time matrix through the same filter to obtain a characteristic signal of the time series template and a characteristic signal of the space-time matrix; S43, calculating the norms of the two characteristic signals; S44, normalizing the two characteristic signals after the norm calculation, and then calculating the cross-correlation of the two signals.
8. The method for behavior recognition based on the spatiotemporal matching of key points between people and objects according to claim 7, characterized in that: The formula for calculating the cross-correlation is: Where r(i) represents the cross-correlation of the signal; n is the signal length; x(i) represents the value of the normalized timing template signal when the signal position is i; y(t+i) represents the value of the normalized space-time matrix signal when the signal position is t+i; t represents the time delay between signals x(i) and y(i).
9. The method for behavior recognition based on the spatiotemporal matching of key points between people and objects according to claim 1, characterized in that: The key points of the human body are joint points of the human body.
Citation Information
Patent Citations
Method for measuring the space-time multi-variant hydrological time series similarity
CN108537247A
Method and device for detecting human-object interaction relationship in video
CN112464875A
Interactive video action comprehensive identification and evaluation system and method
CN114677765A
Human-object interaction action recognition method based on multi-feature fusion
CN116311506A
Action category identification method and device fused with visual knowledge graph
CN116978113A