Dual-modal abnormal skeleton data correction method based on spatio-temporal information and optical flow extraction
Through a dual-modal method based on spatiotemporal information and optical flow extraction, optical flow information is used to correct abnormal bone data, which solves the problem of insufficient detection accuracy of bone data in complex motion scenarios, and improves the accuracy of bone data.
Patent Information
- Application Number
- CN202211374668.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-04
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-11-04
AI Technical Summary
In the prior art, it is difficult to obtain accurate bone data in a single bone data mode in complex motion scenarios, especially in abnormal situations such as occlusion and motion blur, the detection accuracy of bone data is limited.
The dual-mode abnormal bone data correction method based on spatiotemporal information and optical flow extraction is adopted. The human posture estimator is trained to generate an initial posture, perform abnormal detection, and use optical flow information to perform forward and reverse optical flow correction. Combined with the bidirectional correction results, abnormal bone data is predicted and corrected.
Effectively detect and correct abnormal bone data in human body in the video to improve the accuracy of bone data.
Smart Images

Figure CN115619680B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image and video processing and computer vision, and particularly relates to a dual-modal abnormal skeleton data correction method based on spatio-temporal information and optical flow extraction. Background Art
[0002] In recent years, human action recognition and evaluation have been important research fields in computer vision and have been widely applied to real-world scenarios such as human-computer interaction, video monitoring, and video retrieval. Currently, there are multiple modalities for studying human actions, such as RGB videos, skeleton data, optical flow, human shapes, etc. The commonly used skeleton data is a lightweight structural data that is not easily affected by the background and has higher computational efficiency and robustness. However, in practical applications, the acquisition of skeleton data depends on the detection accuracy of key-point detection algorithms.
[0003] Although key-point detection algorithms have achieved high accuracy and effectiveness, they are still challenged and restricted by abnormal situations such as occlusion and motion blur in complex motion scenarios. Studying human actions relying solely on a single skeleton data modality has reached a bottleneck, and how to obtain more accurate skeleton data in complex scenarios has become one of the main challenges currently. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a dual-modal abnormal skeleton data correction method based on spatio-temporal information and optical flow extraction, which can effectively detect abnormal skeleton data of humans in videos, and finally predict and correct the abnormal skeleton data to improve the accuracy of the skeleton data.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A dual-modal abnormal skeleton data correction method based on spatio-temporal information and optical flow extraction, comprising the following steps:
[0007] Step S1: Obtain a human motion video dataset and preprocess it to obtain a training dataset, and train a computer vision-based human pose estimator based on the training dataset;
[0008] Step S2: Generate an initial pose for each video frame of the input video sequence according to the trained human pose estimator, and perform abnormal detection on the initial pose. If an abnormal frame is detected, then proceed to Step S3;
[0009] Step S3: Search for the nearest reliable pre-order frame and the nearest reliable post-order frame before and after the abnormal frame respectively, and extract the optical flow information of the continuous sequence between the nearest reliable pre-order and post-order frames
[0010] Step S4: According to the obtained optical flow information, perform forward optical flow correction and inverse forward optical flow correction respectively in chronological order, and combine the results of the two-way correction to predict and correct abnormal skeletal data.
[0011] Further, the specific steps of Step S1 are as follows:
[0012] Step S11: Obtain a publicly available human motion video dataset;
[0013] Step S12: Process the completely occluded, blurred, mirrored, and irrelevant elements of the human body in the human motion video dataset, screen and organize them to construct a dataset;
[0014] Step S13: Label the dataset, divide the dataset into a training set and a test set according to a certain proportion, and train a computer vision-based human pose estimator based on the training dataset.
[0015] Further, the specific steps of Step S2 are as follows: For an input video sequence I v , which contains N frames, the t-th frame is denoted as I t , t = 1, 2, …, N; First, for each frame I t , t = 1, 2, …, N of the input video sequence, generate the initial pose of the video sequence using the computer vision-based human pose estimator trained, obtaining E(I t ), I t , t = 1, 2, …, N. Then, perform anomaly detection on the generated initial poses E(I t ), I t , t = 1, 2, …, N to obtain the set of abnormal skeletal point data A t for the I t -th frame, which is expressed as follows:
[0016]
[0017]
[0018] where I v represents the input video sequence, I t represents the t-th frame of the video sequence, N represents that the video sequence is divided into N frames in total, E(·) represents the human pose estimation operation, D(·) represents the anomaly detection operation, A t represents the set of abnormal skeletal data detected in the frame I t , represents the data of the k-th abnormal skeletal point in the video frame I t , and H and W are the height and width of the video frame image respectively.
[0019] Further, the generation of the initial pose using the human pose estimator is specifically as follows: The input video sequence I v is divided into multiple frames and processed frame by frame. For a frame image I t , t = 1, 2, …, N, the trained computer vision-based human pose estimator is used to detect the initial pose E(I t ), I t , t = 1, 2, …, N, including the position information of the i-th (i = 1, 2, …, 17) bone point in the I t -th frame and its confidence
[0020] Further, the anomaly detection is specifically as follows: Anomaly detection is performed on the generated initial poses E(I t ), t = 1, 2, …, N. The bone points in frame I t with a confidence lower than λ, and the bone points in frame I t whose Euclidean distance from the previous frame I t-1 or the next frame I t+1 is greater than D are uniformly defined as abnormal bone points. Then, the position information of the abnormal bone point k is the data of the abnormal bone point k Let R(i) represent the situation of the bone point i in the t-th frame I t . R(i) = 0 indicates that the bone point i is an abnormal bone point, and R(i) = 1 indicates that the bone point i is a normal bone point, which is expressed as follows:
[0021]
[0022] where represents the confidence of the i-th (i = 1, 2, …, 17) bone point in the I t -th frame, and λ and D are constants represents the horizontal and vertical coordinates of the position of the i-th (i = 1, 2, …, 17) bone point in the I t -th frame
[0023] Further, step S3 is specifically as follows:
[0024] Step S31: The frame containing abnormal bone data is called an abnormal frame I a . Anomaly detection is performed frame by frame in the forward direction of the abnormal frame to find the nearest reliable previous frame I front for the specific abnormal bone data of this abnormal frame. Similarly, anomaly detection as described in step S23 is performed frame by frame in the backward direction of the abnormal frame to find the nearest reliable subsequent frame I back for the specific abnormal bone data of this abnormal frame;
[0025] Step S32: Use the FlowFormer-based optical flow recognition algorithm to extract the optical flow information of the most recent reliable continuous sequence of preceding and following frames, thereby obtaining the forward optical flow information of the most recent reliable continuous sequence of preceding and following frames. and reverse optical flow information
[0026] Furthermore, the forward optical flow correction is specifically as follows:
[0027] For the video frame I in the forward sequence tf Abnormal Frame I t The bone point data corresponding to the kth abnormal bone point in According to the forward optical flow information Calculate the next video frame I tf+1 Abnormal Frame I t The bone point data corresponding to the kth abnormal bone point in The abnormal skeleton data correction is performed in sequence until the abnormal frame is completed, as shown below:
[0028]
[0029]
[0030] in, Represents video frame I tf Abnormal Frame I t The bone point data corresponding to the k-th abnormal bone point in , Indicates abnormal frame I t The forward correction of the kth abnormal bone point, Indicates abnormal frame I t The corresponding forward optical flow information between the continuous sequences of the most recent credible preceding and following frames, tf represents the video frame number between the most recent credible preceding frame and the abnormal frame, and front represents the most recent credible preceding frame I front The video frame number.
[0031] Furthermore, the reverse optical flow correction is specifically as follows:
[0032] For the video frame I in the reverse sequence tb Abnormal Frame I t The bone point data corresponding to the kth abnormal bone point in According to the forward optical flow information obtained Calculate the next video frame I tb-1 Abnormal Frame I t The bone point data corresponding to the kth abnormal bone point in The abnormal skeleton data correction is performed in sequence until the abnormal frame is completed, as shown below:
[0033]
[0034]
[0035] Among them, represents the video frame I tb corresponding to the k-th abnormal bone point in the abnormal frame I t in the bone point data, represents the reverse correction of the k-th abnormal bone point of the abnormal frame I t ; represents the reverse optical flow information between consecutive sequences between the nearest credible pre- and post-order frames corresponding to the abnormal frame I t tb represents the video frame serial number between the nearest credible post-order frame and the abnormal frame, and back represents the video frame serial number of the nearest credible post-order frame I back ;
[0036] Furthermore, combining the bidirectional correction results, predicting and correcting abnormal bone data specifically includes: for the k-th abnormal bone point of the abnormal frame I t the corrected bone data is obtained by forward optical flow correction the corrected bone data is obtained by the reverse optical flow correction process After fusing the two in a certain proportion, the final corrected bone data is obtained It is expressed as follows:
[0037]
[0038] Among them, represents the final result of correcting the k-th abnormal bone point of the abnormal frame I t ; represents the corrected bone data obtained by the forward optical flow correction process represents the corrected bone data obtained by the reverse optical flow correction process, and λ1 and λ2 are the weights of the forward optical flow correction result and the reverse optical flow correction result.
[0039] The present invention has the following beneficial effects compared with the prior art:
[0040] The present invention can effectively detect the abnormal bone data of the human body in the video, finally predict and correct the abnormal bone data, and improve the accuracy of the bone data. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is the flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0042] The present invention will be further described below with reference to the drawings and embodiments.
[0043] Please refer to Figure 1 , the present invention provides a dual-modal abnormal skeleton data correction method based on skeleton spatio-temporal information and optical flow extraction, including the following steps:
[0044] Step S1: Obtain a dataset and preprocess it, and train a human pose estimator based on computer vision.
[0045] In this embodiment, step S1 specifically includes the following steps:
[0046] Step S11: Obtain a publicly available human motion video dataset;
[0047] Step S12: Preprocess the obtained dataset, process the seriously abnormal and uncorrectable influencing factors such as complete occlusion, blurriness, mirroring, and irrelevant elements of the human body in the dataset, screen and sort out suitable images and video segments, and complete the construction of the dataset;
[0048] Step S13: Label the dataset, divide the dataset into a training set and a test set according to a certain ratio, and use the training set to train a human pose estimator based on ViTPose.
[0049] Step S2: Use the human pose estimator trained in S1 to generate the initial pose of each video frame for the input video sequence, and perform anomaly detection on the initial pose.
[0050] In this embodiment, step S2 specifically includes the following steps:
[0051] Step S21: For an input video sequence I v , assuming it contains N frames, the t-th frame is denoted as I t , t = 1, 2, …, N. First, for each frame I t , t = 1, 2, …, N of the input video sequence, use the human pose estimator trained in step S1 to generate the initial pose of the video sequence, obtaining E(I t ), I t , t = 1, 2, …, N, and then perform anomaly detection on the generated initial pose E(I t ), I t , t = 1, 2, …, N to obtain the set of abnormal skeleton point data A t of the I t -th frame, which is expressed as follows:
[0052]
[0053]
[0054] Among them, I v represents the input video sequence, It denotes the t-th frame of the video sequence, N denotes that the video sequence is divided into N frames in total, E(·) denotes the human pose estimation operation, D(·) denotes the anomaly detection operation, and A t denotes the detected set of abnormal skeleton data in frame I t in denotes the video frame I t data of the k-th abnormal skeleton point in, where H and W are the height and width of the video frame image respectively;
[0055] Step S22: The process of generating the initial pose using the human pose estimator in Step S21 is described as follows. The input video sequence I v is divided into multiple frames and processed frame by frame. For a frame image I t , t = 1, 2, …, N, use the trained ViTPose-based human pose estimator in Step S1 to detect the initial pose E(I t ), I t , t = 1, 2, …, N, including the position information of the i-th (i = 1, 2, …, 17) skeleton point in the I t -th frame and its confidence (In the present invention, the commonly used skeleton points of the human body are used as the skeleton point positions, including the nose, left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, and left and right ankles of the human body, a total of 17 skeleton points);
[0056] Step S23: The process of anomaly detection in Step S21 is described as follows. Anomaly detection is performed on the initial pose E(I t ), t = 1, 2, …, N generated in Step S22. Taking the search for abnormal skeleton points in frame I t as an example, the skeleton points with a confidence t lower than λ in frame I , the skeleton points in frame I t whose Euclidean distance from the previous frame I t-1 or the next frame I t+1 is greater than D are uniformly defined as abnormal skeleton points. Then the position information of the abnormal skeleton point k is the data of the abnormal skeleton point k Let R(i) denote the situation of the skeleton point i in the t-th frame I t . R(i) = 0 indicates that the skeleton point i is an abnormal skeleton point, and R(i) = 1 indicates that the skeleton point i is a normal skeleton point, which is expressed as follows:
[0057]
[0058] Among them, denotes the I tThe confidence of the i-th (i = 1, 2, …, 17) skeleton point in the frame, where λ and D are constants. Denote the I-th t The abscissa and ordinate of the position of the i-th (i = 1, 2, …, 17) skeleton point in the frame.
[0059] Step S3: Search for the nearest reliable pre-order frame and the nearest reliable post-order frame before and after the abnormal frame respectively, and extract the optical flow information of the continuous sequence between the nearest reliable pre-order and post-order frames.
[0060] In this embodiment, step S3 specifically includes the following steps:
[0061] Step S31: The frame containing abnormal skeleton data is called abnormal frame I a , and perform the abnormal detection described in step S23 frame by frame in the forward direction of the abnormal frame to find the nearest reliable pre-order frame I of the specific abnormal skeleton data for this abnormal frame front , and perform the abnormal detection described in step S23 frame by frame in the backward direction of the abnormal frame to find the nearest reliable post-order frame I of the specific abnormal skeleton data for this abnormal frame back ;
[0062] Step S32: Use the optical flow recognition algorithm based on FlowFormer to extract the optical flow information of the continuous sequence between the nearest reliable pre-order and post-order frames, and thus the forward optical flow information between the continuous sequences of the nearest reliable pre-order and post-order frames can be obtained and the reverse optical flow information
[0063] Step S4: According to the optical flow information obtained in S3, perform optical flow correction forward and backward in chronological order respectively, and combine the two-way correction results to predict and correct the abnormal skeleton data.
[0064] In this embodiment, step S4 specifically includes the following steps:
[0065] Step S41: Take the k-th abnormal skeleton point of abnormal frame I t as an example to illustrate the forward optical flow correction process. Starting from the nearest reliable pre-order frame I obtained in step S31 front , extract the continuous optical flow sequence between the nearest reliable pre-order frame and the abnormal frame in chronological order forward. Specifically, for the video frame I in the forward sequence tf corresponding to the k-th abnormal skeleton point in abnormal frame I t in the skeleton point data Calculate the skeleton point data in the next video frame I corresponding to the k-th abnormal skeleton point in abnormal frame I tf+1 according to the forward optical flow information obtained in step S32 t in the skeleton point data The correction of the abnormal skeleton data up to the abnormal frame is completed, as shown below:
[0066]
[0067]
[0068] Among them, represents the video frame I tf in the abnormal frame I t corresponding to the k-th abnormal skeleton point in the data of the skeleton point, represents the forward correction of the k-th abnormal skeleton point of the abnormal frame I t , represents the forward optical flow information between consecutive sequences between the nearest credible previous and subsequent frames corresponding to the abnormal frame I t , tf represents the video frame number between the nearest credible previous frame and the abnormal frame, and front represents the video frame number of the nearest credible previous frame I front ;
[0069] Step S42: Take the k-th abnormal skeleton point of the abnormal frame I t as an example to illustrate the reverse optical flow correction process. Starting from the nearest credible subsequent frame I back obtained in step S31, extract the continuous optical flow sequence between the abnormal frames in reverse chronological order. Specifically, for the video frame I tb in the reverse sequence corresponding to the k-th abnormal skeleton point in the abnormal frame I t in the data of the skeleton point According to the forward optical flow information obtained in step S32 tb-1 calculate the data of the skeleton point corresponding to the k-th abnormal skeleton point in the next video frame I t in the abnormal frame I The correction of the abnormal skeleton data up to the abnormal frame is completed, as shown below:
[0070]
[0071]
[0072] Among them, represents the video frame I tb in the abnormal frame I t corresponding to the k-th abnormal skeleton point in the data of the skeleton point, represents the reverse correction of the k-th abnormal skeleton point of the abnormal frame I t , represents the abnormal frame I tThe reverse optical flow information between the corresponding consecutive sequences of the nearest reliable previous and subsequent frames. tb represents the video frame number between the nearest reliable subsequent frame and the abnormal frame, and back represents the video frame number of the nearest reliable subsequent frame I back of the video frame number;
[0073] Step S43: For the k-th abnormal bone point of the abnormal frame I t The corrected bone data is obtained by the forward optical flow correction process in step S41 The corrected bone data is obtained by the reverse optical flow correction process in step S42 After fusing the two in a certain proportion, the final corrected bone data is obtained It is expressed as follows:
[0074]
[0075] Among them, represents the final result of correcting the k-th abnormal bone point of the abnormal frame I t , represents the corrected bone data obtained by the forward optical flow correction process in step S41 represents the corrected bone data obtained by the reverse optical flow correction process in step S42. λ1 and λ2 are the weights of the forward optical flow correction result and the reverse optical flow correction result
[0076] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope of the present invention
Claims
1. A dual-modal abnormal skeletal data correction method based on spatio-temporal information and optical flow extraction, characterized in that It includes the following steps: Step S1: Obtain a human motion video dataset and preprocess it to obtain a training dataset, and train a computer vision-based human pose estimator based on the training dataset; Step S2: Generate the initial pose of each video frame for the input video sequence according to the trained human pose estimator, and perform anomaly detection on the initial pose. If an abnormal frame is detected, go to Step S3; Step S3: Search for the nearest credible previous frame and the nearest credible subsequent frame before and after the abnormal frame respectively, and extract the optical flow information of the continuous sequence between the nearest credible previous and subsequent frames; Step S4: According to the obtained optical flow information, perform forward optical flow correction and inverse forward optical flow correction in chronological order respectively, and combine the results of the two-way correction to predict and correct the abnormal skeletal data; The specific steps of step S2 are as follows: For an input video sequence I v , which contains N frames, and the t-th frame is denoted as I t , where t = 1, 2, …, N; First, for each frame I t of the input video sequence, a computer vision-based human pose estimator trained to generate the initial pose of the video sequence, obtaining E(I t ), I t , where t = 1, 2, …, N. Then, perform anomaly detection on the generated initial poses E(I t ), I t , where t = 1, 2, …, N, to obtain the set A t of abnormal skeleton point data for the I t -th frame, which is expressed as follows: Among them, I v represents the input video sequence, and I t represents the t-th frame of the video sequence. N represents that the video sequence is divided into N frames in total. E(·) represents the human pose estimation operation, and D(·) represents the anomaly detection operation. A t represents the set of abnormal skeletal data detected in the frame I t . represents the data of the k-th abnormal skeletal point in the video frame I t . H and W are the height and width of the video frame image respectively; Generate an initial pose using a human pose estimator as follows: Divide the input video sequence I v into multiple frames and process them frame by frame. For a frame image I t , t = 1, 2, …, N, use the trained computer vision-based human pose estimator to detect the initial pose E(I t ), I t , t = 1, 2, …, N, including the position information of the i-th (i = 1, 2, …, 17) bone point in the I t -th frame and its confidence The abnormal detection is specifically as follows: perform abnormal detection on the generated initial poses E(I t ), t = 1, 2, …, N. Define the skeleton points in frame I t with confidence lower than λ and the skeleton points in frame I t with Euclidean distance greater than D from the previous frame I t-1 or the next frame I t+1 as abnormal skeleton points. Then, the position information of the abnormal skeleton point k is the data of the abnormal skeleton point k Use R(i) to represent the situation of the skeleton point i in the t-th frame I t . R(i) = 0 indicates that the skeleton point i is an abnormal skeleton point, and R(i) = 1 indicates that the skeleton point i is a normal skeleton point, which is expressed as follows: Among them, represents the confidence of the i-th (i = 1, 2, …, 17) skeletal point in the I-th t frame, where λ and D are constants. represents the horizontal and vertical coordinates of the position of the i-th (i = 1, 2, …, 17) skeletal point in the I-th t frame; The specific content of Step S3 is as follows: Step S31: The frame containing abnormal bone data is called an abnormal frame I a , perform abnormal detection frame by frame in the forward direction of the abnormal frame to find the nearest credible preceding frame I of the specific abnormal bone data for this abnormal frame front , similarly perform the abnormal detection described in step S23 frame by frame in the backward direction of the abnormal frame to find the nearest credible succeeding frame I of the specific abnormal bone data for this abnormal frame back ; Step S32: Use the optical flow recognition algorithm based on FlowFormer to extract the optical flow information of the continuous sequence between the nearest reliable previous and subsequent frames, thereby obtaining the forward optical flow information between the continuous sequences of the nearest reliable previous and subsequent frames and the reverse optical flow information The forward optical flow correction is specifically as follows: For the video frame I in the forward sequence tf and the abnormal frame I t the skeleton point data corresponding to the k-th abnormal skeleton point According to the forward optical flow information calculate the skeleton point data corresponding to the k-th abnormal skeleton point in the next video frame I tf+1 and the abnormal frame I t Execute sequentially until the correction of the abnormal skeleton data at the abnormal frame is completed, which is expressed as follows: Execute sequentially until the correction of the abnormal skeleton data at the abnormal frame is completed, which is expressed as follows: Among them, represents the video frame I tf corresponding to the abnormal frame I t and the bone point data corresponding to the k-th abnormal bone point in it, represents the forward correction of the k-th abnormal bone point of the abnormal frame I t ; represents the forward optical flow information between consecutive sequences of the nearest reliable pre- and post-order frames corresponding to the abnormal frame I t , tf represents the video frame number between the nearest reliable pre-order frame and the abnormal frame, and front represents the video frame number of the nearest reliable pre-order frame I front ; The inverse optical flow correction is specifically as follows: For the video frame I in the reverse sequence tb and the abnormal frame I t the bone point data corresponding to the k-th abnormal bone point According to the forward optical flow information calculate the bone point data corresponding to the k-th abnormal bone point in the next video frame I tb-1 and the abnormal frame I t in sequence until the correction of the abnormal bone data at the abnormal frame is completed, which is expressed as follows: Among them, represents the video frame I tb corresponding to the abnormal frame I t and the bone point data corresponding to the k-th abnormal bone point in represents the reverse correction of the k-th abnormal bone point of the abnormal frame I t ; represents the reverse optical flow information between consecutive sequences of the nearest reliable front and back frames corresponding to the abnormal frame I t , tb represents the video frame number between the nearest reliable subsequent frame and the abnormal frame, and back represents the video frame number of the nearest reliable subsequent frame I back ; Combining the two-way correction results, predicting and correcting abnormal skeleton data, specifically: for the k-th abnormal skeleton point in the abnormal frame I t the corrected skeleton data is obtained by forward optical flow correction the corrected skeleton data is obtained by the reverse optical flow correction process After fusing the two according to a certain ratio, the final corrected skeleton data is obtained It is expressed as follows: Among them, represents the final result of correcting the k-th abnormal skeleton point of the abnormal frame I t , represents the corrected skeleton data obtained from the forward optical flow correction process, represents the corrected skeleton data obtained from the reverse optical flow correction process, and λ1 and λ2 are the weights of the forward optical flow correction result and the reverse optical flow correction result.
2. The dual-modal abnormal skeleton data correction method based on spatio-temporal information and optical flow extraction according to claim 1, wherein The specific content of Step S1 is as follows: Step S11: Obtain a publicly available human motion video dataset; Step S12: Process the human body that is completely blocked, blurred, mirrored, or contains irrelevant elements in the human motion video dataset, filter and sort it to construct a dataset; Step S13: Label the dataset, divide the dataset into a training set and a test set according to a certain ratio, and train a computer vision-based human pose estimator based on the training dataset.