An Action Prediction Method and System Based on Lattice Optical Flow
Through the action prediction method based on grid point optical flow, combined with the joint node depth regression network and the Seq2seq_attention network, real-time 3D human body movement prediction is realized, solving the balance problem of efficiency and accuracy in the existing technology, and significantly improving the accuracy and speed of action prediction.
Patent Information
- Application Number
- CN202210568887.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-24
AI Technical Summary
The prior art has problems with efficiency and accuracy in real-time 3D human posture prediction, especially in fields such as autonomous driving and sports, and it is difficult to achieve fast and accurate action prediction.
Using the action prediction method based on grid point optical flow, the human body area is detected through an ordinary RGB camera, combined with the joint node depth regression network and the Seq2seq_attention network, the optical flow vector of the joint node is calculated and a three-dimensional coordinate pair is generated to realize real-time 3D human body motion prediction.
This method not only improves the ability to extract features, but also maintains a balance of accuracy and speed, eliminates redundant information in the optical flow, significantly accelerates the calculation of optical flow vectors, and can be adapted to domestic AI acceleration cards, with good versatility, flexibility and portability.
Smart Images

Figure CN115100559B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a method and system for action prediction based on lattice optical flow. Background Art
[0002] Human pose estimation is a human pose detection and analysis technology based on computer vision. Literally, it can be understood as the position estimation of the human pose (including joint points such as the head, hands, feet, etc.). Its mainstream classifications include: single-person pose estimation, multi-person pose estimation, human pose tracking, and 3D human pose estimation. Its main component is the modeling of the human body. There are three most commonly used human body models: skeleton-based models, contour-based models, and volume-based models. Since CPM (Compressible Packing Mode), neural networks have been able to implicitly model the feature map representation and the spatial position relationship of joint points end-to-end. Input a region of an image, and output a vector with spatial information. The number of channels is generally the number of human joint points (or the number of joint points plus 1). By finding the maximum response position (x, y coordinates) for each channel on the output heat map, the position of the corresponding joint point can be found. This heat map representation method is widely used in the problem of the human skeleton.
[0003] Current human pose detection algorithms are widely used in fields such as human-computer interaction, autonomous driving, and intelligent security. For example, in human-computer interaction, intelligent devices such as drones can be operated through human gestures; in an autonomous driving system, potential dangers can be predicted by judging the pose of pedestrians, thereby providing decision-making support for the system; in intelligent security, dangerous actions of the target population in the monitoring can be judged by monitoring the human pose, and early warnings can be given in a timely manner. In addition, human pose estimation has an important application background in fields such as action recognition, film and television production, motion sensing games, and human positioning. For example, using pose estimation to analyze the movement and actions of athletes, etc. In the application of action recognition, by real-time detecting human actions, it can be used to detect whether a person falls, for teaching and training in sports, dance and other fields, and for understanding body language; in the application scenario of motion capture, by detecting the human pose, graphics, styles, special effects, etc. can be loaded onto the human body, promoting the progress of the augmented reality field.
[0004] Merely recognizing and analyzing human actions seems to be insufficient. For example, in the field of autonomous driving, potential risks can be avoided during autonomous driving through pedestrian road prediction. The normal reaction time range of a person is between 0.1s - 0.5s. If real-time 3D human pose prediction can be achieved and applied to autonomous driving, sports, or entertainment projects, the reaction time of people in these fields can be greatly improved.
[0005] Based on the above situation, the present invention proposes a method and system for action prediction based on lattice optical flow. Summary of the Invention
[0006] In order to make up for the deficiencies of the prior art, the present invention provides a simple and efficient method and system for action prediction based on lattice optical flow.
[0007] The present invention is implemented by the following technical solutions:
[0008] A method for action prediction based on lattice optical flow, characterized by comprising the following steps:
[0009] S1. Use an ordinary RGB camera as the input device of the target detector, and use the target detector to crop the target human body area from the RGB frames in the input video;
[0010] S2. Send the detected single-person area into the human body pose estimation network to extract the 2D human joint position information in the input picture, and obtain the depth information corresponding to the joint points through the joint point depth regression network for the obtained 2D human joint points;
[0011] The joint point depth regression network includes two branches: depth estimation based on the local area of the image and depth estimation based on global joint points. After integrating the joint point depth information of the two branches, it is used as the corresponding joint point depth information;
[0012] S3. Calculate the lattice points around the target human body area using the Lucas-Kanade sparse optical flow calculation method, and use the two-dimensional joint positions to average the optical flow around the joint points, calculate the average optical flow vector near the joint points, and obtain the changes of each joint point in the time domain and the correlation between adjacent frames;
[0013] S4. Transmit the joint point depth information and joint optical flow to the Seq2seq_attention network for prediction, and generate three-dimensional coordinate pairs to represent the three-dimensional pose.
[0014] In step S1, a custom ShuffleNet_CBAM_k5 network is used to perform regression on the joint point positions;
[0015] The custom process is as follows:
[0016] Expand the convolution kernel of the 3×3 depth convolution in the ShuffleNet v2 convolutional neural network, replace it with a 5×5 depth convolution, modify the padding attribute padding to 2, and apply the channel attention and spatial attention mechanisms in the attention model CBAM to the ShuffleNet v2 convolutional neural network.
[0017] In the step S2, the joint depth regression network is an LSTM network that can be trained end-to-end;
[0018] To solve the problem of the inherent ambiguity in a single view when lifting 2D poses to 3D poses, the joint depth regression network utilizes 2D joint points and local regions of image body parts;
[0019] The two branches of the joint depth regression network are the human skeleton joint point mapping network and the image Patch LSTM network;
[0020] The human skeleton joint point mapping network learns global depth information based on 2D joint points, uses the 2D coordinates of the image as input and encodes it. The encoding module contains two residual blocks, and each residual block is composed of a linear layer, a BN layer, a ReLU activation function, and Dropout;
[0021] The image Patch LSTM network learns using the local image information of each part of the human body, takes the two-dimensional original image as input, and extracts input features through convolution.
[0022] The outputs of the human skeleton joint point mapping network and the image Patch LSTM network are used as the input of the second-level Seq2seq_attention network through the max-pooling layer. The predicted depth of each joint point is the logarithmic space after the output of the hidden layer of the second-level Seq2seq_attention network passes through the fully connected layer.
[0023] In the step S3, first remove the irrelevant information in the optical flow map, including the optical flow redundant information brought by camera movement and the background redundant optical flow information;
[0024] Since the two-dimensional joint point estimation runs in parallel and is faster than the optical flow calculation in most cases, when the Lucas-Kanade sparse optical flow calculation method obtains the two-dimensional joint point position, it directly determines the rectangular grid point region with side length d max near the position;
[0025] The Lucas-Kanade sparse optical flow calculation method sets a candidate rectangular calculation region with side length d max for each joint point, then evenly selects grid points within this rectangular region and calculates the optical flow vector of each grid point, and finally takes the average value of the optical flow vectors in this region as the optical flow vector of this joint point.
[0026] To remove the optical flow redundant information brought by camera movement, first, camera movement estimation needs to be performed:
[0027] Traverse the global pixels to statistically calculate the optical flow histogram in the background, select the optical flow direction with the highest frequency as the camera movement direction, and obtain the average of the background optical flow speed magnitude as the camera movement speed magnitude; after obtaining the camera movement vector, correct the optical flow of each pixel point.
[0028] The Seq2seq_attention network is based on the traditional Seq2seq network structure and consists of two parts: an encoder and a decoder;
[0029] In the encoding stage, the encoder processes the input content and then outputs it as a vector as an intermediate representation; in the decoding stage, the decoder receives this vector and then generates the output prediction sequence.
[0030] The encoder uses a GRU (Gated Recurrent Neural Network) network, and the decoder uses a GRU network with an attention mechanism and residual connections added; when setting the model, set the size of the vector generated by the encoder.
[0031] The Seq2seq_attention network transfers the states of all hidden layers in the encoding stage; in the decoding stage, a hidden layer state h is generated through GRU, and the generated hidden layer state h is processed through the attention mechanism to generate a vector c. After concatenating the hidden layer state h and the vector c to obtain a new vector, it is output through the FC layer; then it continues to repeat at the next time node.
[0032] To solve the problem of discontinuous predicted values for the first frame, residual connections are introduced into the traditional Seq2seq network to learn the speed instead of learning the human pose itself. That is, the prediction for each frame is equivalent to predicting the change value of the speed, rather than predicting the human pose itself. Therefore, the prediction for the first frame is simplified to a prediction of 0 speed or close to 0 speed. Adding a link between the input and output of each GRU network can learn the change in speed.
[0033] A system for an action prediction method based on lattice optical flow, characterized by comprising:
[0034] Human pose estimation module: The load uses an ordinary RGB (red, green, and blue) camera as the input device of the target detector, and uses the target detector to crop the target human body area from the RGB frames in the input video;
[0035] Send the detected single-person area into the human pose estimation network to extract the 2D human joint position information in the input picture, and obtain the depth information corresponding to the joint points of the obtained 2D human joint points through the joint point depth regression network;
[0036] The joint point depth regression network includes two branches: depth estimation based on local image regions and depth estimation based on global joint points. After integrating the joint point depth information of the two branches, it is used as the corresponding joint point depth information.
[0037] Optical flow vector calculation module: responsible for calculating the lattice points around the target human body region using the Lucas-Kanade sparse optical flow calculation method, and averaging the optical flow around the joint points using the two-dimensional joint positions to calculate the average optical flow vector near the joint points, obtaining the changes of each joint point in the time domain and the correlation between adjacent frames.
[0038] 3D pose prediction module: responsible for transmitting the joint point depth information and joint optical flow to the Seq2seq_attention network for prediction, generating three-dimensional coordinate pairs to represent the three-dimensional pose.
[0039] The beneficial effects of the present invention are as follows: The action prediction method and system based on lattice optical flow not only enhance the ability of feature extraction, but also maintain the balance between accuracy and speed, eliminate the redundant information in the optical flow, further accelerate the calculation of the optical flow vector, can be adapted to domestic AI acceleration cards, and have good versatility, flexibility and portability. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Attached Figure 1 is a schematic diagram of the pose estimation and action prediction method of the present invention.
[0042] Attached Figure 2 is a schematic diagram of the ShuffleNet_CBAM_k5 network structure of the present invention.
[0043] Attached Figure 3 is a schematic diagram of the custom Seq2seq_attention network structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To enable those skilled in the art to better understand the technical solutions in the present invention, the following will, in conjunction with the embodiments of the present invention, clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0045] The action prediction method based on lattice optical flow includes the following steps:
[0046] S1. Use an ordinary RGB (red, green, and blue) camera as the input device of the target detector, and use the target detector to crop the target human body area from the RGB frames in the input video;
[0047] S2. Send the detected single-person area into the human pose estimation network to extract the 2D human joint position information in the input picture, and obtain the depth information corresponding to the joint points of the 2D human joint points through the joint point depth regression network;
[0048] The joint point depth regression network includes two branches: depth estimation based on the local area of the image and depth estimation based on global joint points. The joint point depth information of the two branches is integrated as the corresponding joint point depth information;
[0049] S3. Use the Lucas-Kanade sparse optical flow calculation method to calculate the lattice points around the target human body area, and use the two-dimensional joint positions to average the optical flow around the joint points, calculate the average optical flow vector near the joint points, and obtain the changes of each joint point in the time domain and the correlation between adjacent frames;
[0050] S4. Transmit the joint point depth information and joint optical flow to the Seq2seq_attention network for prediction to generate three-dimensional coordinate pairs to represent the three-dimensional pose.
[0051] In step S1, a custom ShuffleNet_CBAM_k5 network is used to regress the joint positions;
[0052] The custom process is as follows:
[0053] Expand the 3×3 depth convolution in the ShuffleNet v2 convolutional neural network, replace it with a 5×5 depth convolution, modify the padding attribute padding to 2, and apply the channel attention and spatial attention mechanisms in the attention model CBAM to the ShuffleNet v2 convolutional neural network.
[0054] By observing the computational load distribution of ShuffleNet v2, the computational load ratio on depthwise convolution DWConv is actually relatively small. The main computational load lies in the 1×1 convolution. Expanding the convolutional kernel of the 3×3 depthwise convolution will neither increase the computational ratio too much nor improve the effect significantly.
[0055] Compared with the SENet that only focuses on the channel attention mechanism, the CBAM module of the attention model can significantly improve the model performance with a small increase in computational load and number of parameters. The basic module unit with the above two changes is embedded into the ShuffleNet_CBAM_k5 network structure.
[0056] The first part of the ShuffleNet_CBAM_k5 network is the same as the downsampling part of ShuffleNet v2, both are convolutional layers and pooling layers with a stride of 2 and a convolutional kernel size of 3×3. Taking the cropped RGB image with a size of 224×224 as the input, after the first stage, a feature map with a size of 28×28 is obtained and fed into the network stage stacked with the designed ShuffleNet module units. This stage has a total of three sub-stages, and the first module of each sub-stage is a downsampling unit with a stride of 2, as shown in Appendix Figure 2 (b). Finally, after one layer of pooling, the convolutional result is connected to linear regression to obtain the two-dimensional joint position. In this method, 17 joints are defined because the output dimension of pose estimation is 17×2. The output joint positions come from five consecutive frames (the current frame and the previous four frames), and then are passed to the temporal network layer responsible for prediction together with the calculated optical flow data.
[0057] In the step S2, the joint depth regression network is an LSTM (Long Short-Term Memory) network that can be trained end-to-end;
[0058] To solve the problem of the inherent ambiguity in a single view when lifting 2D poses to 3D poses, the joint depth regression network uses 2D joints and local regions of the body parts in the image;
[0059] The two branches of the joint depth regression network are the human skeleton joint mapping network and the image Patch LSTM network;
[0060] The human skeleton joint mapping network learns global depth information based on 2D joints, uses the 2D coordinates of the image as the input and encodes it. The encoding module contains two residual blocks, and each residual block consists of a linear layer, a BN layer, a ReLU activation function, and Dropout;
[0061] The image Patch LSTM network learns using local image information of various parts of the human body, takes a two-dimensional original image as input, and extracts input features through convolution.
[0062] The outputs of the human body skeleton joint point mapping network and the image Patch LSTM network are used as the input of the second-level Seq2seq_attention network through a max-pooling layer. The predicted depth of each joint point is the logarithm space after full connection of the output of the hidden layer of the second-level Seq2seq_attention network.
[0063] In step S3, irrelevant information in the optical flow map is removed first, including optical flow redundant information caused by camera movement and background redundant optical flow information.
[0064] Since two-dimensional joint point estimation runs in parallel and is faster than optical flow calculation in most cases, when the Lucas-Kanade sparse optical flow calculation method obtains the two-dimensional joint point positions, it directly determines a rectangular grid point region with side length d max near the position.
[0065] The Lucas-Kanade sparse optical flow calculation method sets a candidate rectangular calculation region with side length d max for each joint point, then evenly selects grid points within this rectangular region and calculates the optical flow vectors of each grid point. Finally, the average value of the optical flow vectors in this region is taken as the optical flow vector of this joint point.
[0066] In order to remove the optical flow redundant information caused by camera movement, camera movement estimation needs to be carried out first:
[0067] Traverse the global pixels to statistically calculate the optical flow histogram in the background, select the optical flow direction with the highest frequency as the camera movement direction, and obtain the average value of the background optical flow speed magnitude as the camera movement speed magnitude. After obtaining the camera movement vector, the optical flow of each pixel point is corrected.
[0068] Set an optical flow magnitude threshold through experiments. Optical flow information below this threshold is background redundant optical flow information.
[0069] Compared with conventional dense optical flow, the optical flow calculation amount of this Lucas-Kanade sparse optical flow calculation method is reduced more significantly.
[0070] After obtaining the human joint point sequence and its corresponding optical flow vector sequence, the two are taken together as input and passed to the Seq2seq_attention network for prediction, as shown in the appendix Figure 3 shown.
[0071] The Seq2seq_attention network is based on the traditional Seq2seq network structure and consists of two parts: an encoder and a decoder;
[0072] In the encoding stage, the encoder processes the input content and then outputs a vector as an intermediate representation; in the decoding stage, the decoder receives this vector and then generates the predicted output sequence.
[0073] The encoder uses a GRU (Gated Recurrent Neural Network) network, and the decoder uses a GRU network with an attention mechanism and residual connections added; when setting up the model, the size of the vector generated by the encoder is set. It is basically the number of hidden units in the encoder GRU.
[0074] In traditional Seq2seq, only the state of the last hidden layer is passed. After introducing the attention mechanism into the traditional Seq2seq network, the Seq2seq_attention network passes the states of all hidden layers in the encoding stage; in the decoding stage, a hidden layer state h is generated through GRU, and the generated hidden layer state h is processed through the attention mechanism to generate a vector c. After concatenating the hidden layer state h and the vector c to obtain a new vector, it is output through the FC layer; then it continues to repeat at the next time node.
[0075] To solve the problem of discontinuous predicted values for the first frame, residual connections are introduced into the traditional Seq2seq network to learn the speed instead of the human pose itself. That is, the prediction for each frame is equivalent to predicting the change value of the speed, rather than predicting the human pose itself. Therefore, the prediction for the first frame is simplified to a prediction of 0 speed or a speed close to 0. The implementation method is also relatively simple. Just add a connection between the input and output of each GRU network to learn the change in speed.
[0076] The system of the action prediction method based on lattice optical flow includes:
[0077] Human pose estimation module: The load uses an ordinary RGB (red, green, and blue) camera as the input device of the target detector, and uses the target detector to crop the target human body area from the RGB frames in the input video;
[0078] The detected single-person area is sent into the human pose estimation network to extract the 2D human joint position information in the input picture. The obtained 2D human joints obtain the depth information corresponding to the joints through the joint point depth regression network;
[0079] The joint point depth regression network includes two branches: depth estimation based on local image regions and depth estimation based on global joint points. The joint point depth information of the two branches is integrated as the corresponding joint point depth information.
[0080] Optical flow vector calculation module: responsible for calculating the lattice points around the target human body region using the Lucas-Kanade sparse optical flow calculation method, and averaging the optical flow around the joint points using the two-dimensional joint positions to calculate the average optical flow vector near the joint points, obtaining the changes of each joint point in the time domain and the correlation between adjacent frames.
[0081] 3D pose prediction module: responsible for transmitting the joint point depth information and joint optical flow to the Seq2seq_attention network for prediction, generating three-dimensional coordinate pairs to represent the three-dimensional pose.
[0082] Compared with the prior art, the action prediction method and system based on lattice optical flow have the following characteristics:
[0083] First, it can use an ordinary RGB camera to record the movement of the object in real time and infer its actions in the next 0.5s, providing a method for real-time 3D human action prediction.
[0084] Second, by improving the backbone network ShuffleNet for pose estimation, the detection accuracy of human joint points is further improved.
[0085] Third, a sparse optical flow method is used to accelerate the calculation, and the optical flow redundant information brought by the background and motion estimation is eliminated before the calculation. The underlying algorithm is processed using asynchronous and multi-threaded methods, making full use of the computing power resources for inference.
[0086] Fourth, a recurrent neural network structure combined with an attention mechanism is adopted, which can effectively realize the prediction of the subsequent frames of human motion.
[0087] Fifth, the network parameters are quantized with low bits using tools such as domestic acceleration cards, and it can be deployed and run on domestic AI acceleration cards, which can effectively improve the inference speed at the edge side and reduce the consumption of board card resources.
[0088] The above-described embodiments are only one of the specific implementation manners of the present invention. The general changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.
Claims
1. An action prediction method based on lattice optical flow, characterized in that: It includes the following steps: S1. Use an ordinary RGB camera as the input device of the target detector, and use the target detector to crop the target human body area from the RGB frames in the input video; S2. Send the detected single-person area into the human body pose estimation network to extract the 2D human joint position information in the input picture, and the obtained 2D human joints obtain the depth information corresponding to the joints through the joint point depth regression network; The joint point depth regression network includes two branches: depth estimation based on the local area of the image and depth estimation based on global joint points. The depth information of the joints of the two branches is integrated as the depth information corresponding to the joints; In step S2, the joint point depth regression network is an LSTM network that can be trained end-to-end; To solve the problem of the inherent ambiguity in a single view when lifting 2D poses to 3D poses, the joint point depth regression network uses 2D joints and the local areas of the body parts in the image; The two branches of the joint point depth regression network are the human skeleton joint point mapping network and the image PatchLSTM network respectively; The human skeleton joint point mapping network learns the global depth information according to the 2D joints, uses the 2D coordinates of the image as the input and encodes it. The encoding module contains two residual blocks, and each residual block is composed of a linear layer, a BN layer, a ReLU activation function and Dropout; The image Patch LSTM network uses the local image information of each part of the human body for learning, uses the two-dimensional original image as the input, and extracts the input features through convolution; The outputs of the human skeleton joint point mapping network and the image Patch LSTM network are used as the input of the second-level Seq2seq_attention network through the max-pooling layer. The predicted depth of each joint is the logarithmic space after the output of the hidden layer of the second-level Seq2seq_attention network passes through the fully connected layer; S3. Use the Lucas-Kanade sparse optical flow calculation method to calculate the lattice points around the target human body area, and use the two-dimensional joint position to average the optical flow around the joint points, calculate the average optical flow vector near the joint points, and obtain the changes of each joint point in the time domain and the correlation between adjacent frames; S4. Transmit the joint point depth information and joint optical flow into the Seq2seq_attention network for prediction, and generate three-dimensional coordinate pairs to represent the three-dimensional pose.
2. The action prediction method based on lattice optical flow according to claim 1, characterized in that: In step S1, a custom ShuffleNet_CBAM_k5 network is used to regress the joint positions; The custom process is as follows: Expand the convolutional kernel of the 3×3 depth convolution in the ShuffleNet v2 convolutional neural network, replace it with a 5×5 depth convolution, modify the padding attribute padding to 2, and apply the channel attention and spatial attention mechanisms in the CBAM attention model to the ShuffleNet v2 convolutional neural network.
3. The method for action prediction based on lattice optical flow according to claim 1, characterized in that: In step S3, irrelevant information in the optical flow map is removed first, including the optical flow redundant information and background redundant optical flow information caused by camera movement; Since the two-dimensional joint point estimation runs in parallel and is faster than the optical flow calculation in most cases, when the Lucas-Kanade sparse optical flow calculation method obtains the two-dimensional joint point positions, it directly determines a rectangular grid point region with side length d near the positions max ; The Lucas-Kanade sparse optical flow calculation method sets a candidate rectangular calculation area with a side length of d at each joint point max , then uniformly selects grid points within this rectangular area and calculates the optical flow vectors of each grid point. Finally, the average value of the optical flow vectors in this area is taken as the optical flow vector of this joint point.
4. The method for action prediction based on lattice optical flow according to claim 3, characterized in that: In order to remove the optical flow redundant information caused by camera movement, camera movement estimation needs to be carried out first: Traverse the global pixels to statistically calculate the optical flow histogram in the background, select the optical flow direction with the highest frequency as the camera movement direction, and calculate the average of the background optical flow speed magnitude as the camera movement speed magnitude; after obtaining the camera movement vector, correct the optical flow of each pixel point.
5. The method for action prediction based on lattice optical flow according to claim 1, characterized in that: The Seq2seq_attention network is based on the traditional Seq2seq network structure and consists of two parts: an encoder and a decoder; In the encoding stage, the encoder processes the input content and then outputs a vector as an intermediate representation; In the decoding stage, the decoder receives this vector and then generates an output prediction sequence.
6. The method for action prediction based on lattice optical flow according to claim 5, characterized in that: The encoder adopts a GRU network, and the decoder adopts a GRU network with an attention mechanism and residual connection added; when setting the model, set the size of the vector generated by the encoder.
7. The method for action prediction based on lattice optical flow according to claim 6, characterized in that: The Seq2seq_attention network transmits the states of all hidden layers in the encoding stage; in the decoding stage, a hidden layer state h is generated through GRU, and the generated hidden layer state h is processed through the attention mechanism to generate a vector c. After concatenating the hidden layer state h and the vector c to obtain a new vector, it is output through the FC layer; then continue to repeat at the next time node.
8. The method for action prediction based on lattice optical flow according to claim 6, characterized in that: In order to solve the problem of discontinuous prediction values for the first frame, a residual connection is introduced into the traditional Seq2seq network to learn the speed instead of learning the human pose itself. That is, the prediction of each frame is equivalent to predicting the change value of the speed, rather than predicting the human pose itself. Therefore, the prediction of the first frame is simplified to a prediction of 0 speed or a speed close to 0 speed, and a link is added to the input and output of each GRU network to learn the change in speed.
9. A system for the method for action prediction based on lattice optical flow according to any one of claims 1 to 8, It is characterized in that: Including: Human body pose estimation module: The load uses an ordinary RGB camera as the input device of the target detector, and uses the target detector to crop the target human body area from the RGB frames in the input video; The detected single-person area is sent into the human body pose estimation network to extract the 2D human body joint point position information in the input picture, and the obtained 2D human body joint points obtain the depth information corresponding to the joint points through the joint point depth regression network; The joint point depth regression network includes two branches, namely depth estimation based on the local area of the image and depth estimation based on global joint points. The joint point depth information of the two branches is integrated as the corresponding joint point depth information; Optical flow vector calculation module: Responsible for calculating the dot matrix lattice points around the target human body area using the Lucas-Kanade sparse optical flow calculation method, and averaging the optical flow around the joint points using the two-dimensional joint position to calculate the average optical flow vector near the joint points, obtaining the changes of each joint point in the time domain and the correlation between adjacent frames; 3D pose prediction module: Responsible for transmitting the joint point depth information and joint optical flow to the Seq2seq_attention network for prediction, generating three-dimensional coordinate pairs to represent the three-dimensional pose.
Citation Information
Patent Citations
3D human body posture estimation method based on sparsity and depth
CN111046733A
Double-person interaction behavior identification method based on joint point-depth joint attention RGB modal data
CN112668550A