A video stream signal feature extraction method and system
By using the YOLO neural network object detection model and peak search algorithm, the legs and foot information and motion characteristics in the video stream signal are extracted, and the problem of low feature extraction efficiency and accuracy in the prior art is solved, and more efficient and accurate video stream signal feature extraction is achieved.
Patent Information
- Application Number
- CN202210097751.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-01-27
AI Technical Summary
The existing video stream signal feature extraction method is based on end-to-end deep learning modeling, making it difficult to learn key feature information, and the convolutional layer translation invariance problem and the Pooling layer information loss problem are difficult to solve, resulting in low extraction efficiency and accuracy.
The YOLO neural network object detection model is used to obtain the legs and feet information in the video data, delete the multi-frame legs and feet information through the leg confidence, determine the motion trajectory, and use the peak search algorithm to extract the velocity and ratio from the motion trajectory and the leg and feet position as the motion characteristics.
It improves the efficiency and accuracy of video stream signal feature extraction, and can extract relevant motion characteristics in video stream signal more quickly and accurately. It is suitable for Parkinson's disease rehabilitation assessment and other scenarios.
Smart Images

Figure CN114429158B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video signal processing, and in particular to a method and system for extracting features of video stream signals. Background Art
[0002] For feature extraction of video stream signals, currently commonly used methods include TCNN, optical flow method for feature migration or fusion, and using LSTM and other memory methods for feature migration or fusion. However, most existing methods are based on end-to-end deep learning modeling methods, and there is no way to learn key or desired feature information. In addition, if existing networks such as Alexnet and Googlenet are used to extract features, there will be problems with translation invariance of the convolutional layer and information loss in the Pooling layer. How to quickly and accurately extract relevant motion features from video stream signals has become a problem that needs to be solved. Summary of the invention
[0003] The purpose of the present invention is to provide a method and system for extracting features of video stream signals, so as to improve the efficiency and accuracy of extracting features of video stream signals.
[0004] To achieve the above object, the present invention provides the following solutions:
[0005] A method for extracting features of a video stream signal, comprising:
[0006] Obtain the video data to be extracted;
[0007] The video data to be extracted is input into the target detection model to obtain multiple frames of leg and foot information; the leg and foot information includes leg and foot positions, leg and foot labels, and leg and foot confidence; the target detection model is a YOLO neural network;
[0008] Deleting the leg and foot information of multiple frames according to the leg and foot confidences to determine the motion trajectory;
[0009] The motion features are obtained according to the motion trajectory and the leg and foot positions using a peak-finding algorithm, wherein the motion features include speed and ratio; the ratio is the ratio of amplitude to leg length.
[0010] Optionally, the training process of the target detection model specifically includes:
[0011] The YOLO neural network is trained with the labeled video data of the training set as input and the leg and foot positions, leg and foot labels and leg and foot confidence of the training set as output to obtain the target detection model.
[0012] Optionally, deleting the leg and foot information of multiple frames according to the leg and foot confidences to determine the motion trajectory specifically includes:
[0013] Deleting the leg and foot information of multiple frames according to the leg and foot confidences to determine the remaining leg and foot information;
[0014] Calculating the Euclidean distance based on the remaining leg and foot information of adjacent frames;
[0015] Determine a plurality of category sequences according to the Euclidean distance;
[0016] The motion trajectory is determined according to the lengths of the plurality of category sequences; the motion trajectory is the longest category sequence in the same category sequence.
[0017] Optionally, obtaining the motion feature by using a peak-finding algorithm according to the motion trajectory and the leg and foot positions specifically includes:
[0018] According to the motion trajectory, a peak position and a peak number are obtained by using a peak-finding algorithm;
[0019] determining a speed according to the peak position, the number of peaks and the movement time;
[0020] The ratio is determined according to the leg and foot positions in the leg and foot information.
[0021] A video stream signal feature extraction system, comprising:
[0022] An acquisition module, used for acquiring video data to be extracted;
[0023] A leg and foot information determination module is used to input the video data to be extracted into a target detection model to obtain multiple frames of leg and foot information; the leg and foot information includes leg and foot positions, leg and foot labels, and leg and foot confidence; the target detection model is a YOLO neural network;
[0024] A deletion module, used for deleting the leg and foot information of multiple frames according to the leg and foot confidence, to determine the motion trajectory;
[0025] The motion feature determination module is used to obtain the motion feature using a peak-finding algorithm according to the motion trajectory and the leg and foot positions, wherein the motion feature includes speed and ratio; the ratio is the ratio of amplitude to leg length.
[0026] Optionally, the training process of the target detection model specifically includes:
[0027] The YOLO neural network is trained with the labeled video data of the training set as input and the leg and foot positions, leg and foot labels and leg and foot confidence of the training set as output to obtain the target detection model.
[0028] Optionally, the deletion module specifically includes:
[0029] A deleting unit, used for deleting the leg and foot information of multiple frames according to the leg and foot confidence, and determining the remaining leg and foot information;
[0030] A calculation unit, used for calculating the Euclidean distance according to the remaining leg and foot information of adjacent frames;
[0031] A category sequence determination unit, used for determining a plurality of category sequences according to the Euclidean distance;
[0032] The motion trajectory determining unit is used to determine the motion trajectory according to the lengths of the plurality of category sequences; the motion trajectory is the category sequence with the longest length in the same category sequence.
[0033] Optionally, the motion feature determination module specifically includes:
[0034] A peak position and peak number determination unit, used to obtain the peak position and peak number according to the motion trajectory using a peak search algorithm;
[0035] a speed determination unit, configured to determine the speed according to the peak position, the number of peaks and the motion time;
[0036] A ratio determination unit is used to determine the ratio according to the leg and foot positions in the leg and foot information.
[0037] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0038] The present invention provides a method and system for extracting features of video stream signals, which obtain video data to be extracted; input the video data to be extracted into a target detection model to obtain multiple frames of leg and foot information; the leg and foot information includes leg and foot positions, leg and foot labels and leg and foot confidences; the target detection model is a YOLO neural network; multiple frames of leg and foot information are deleted according to the leg and foot confidences to determine the motion trajectory; motion features are obtained according to the motion trajectory and the leg and foot positions using a peak-finding algorithm, and the motion features include speed and ratio; the ratio is the ratio of amplitude to leg length. The leg and foot information is obtained using the YOLO neural network, and then the leg and foot information is processed to obtain the motion features, which can improve the efficiency and accuracy of video stream signal feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0040] Figure 1 A flow chart of the method for extracting features of video stream signals provided by the present invention;
[0041] Figure 2A schematic diagram of the flexibility assessment of the legs of a tester with the knee joint flexed provided by the present invention;
[0042] Figure 3 A diagram of the YOLO model training process provided by the present invention;
[0043] Figure 4 A target detection result diagram of the test subject's leg flexibility assessment in a flexed knee state provided by the present invention;
[0044] Figure 5 A schematic diagram of the motion feature information of the tester extracted by the YOLO target detection provided by the present invention;
[0045] Figure 6 Feature extraction and classification flow chart for YOLO target detection motion feature extraction in Parkinson's disease rehabilitation assessment;
[0046] Figure 7 This is a diagram of the feature extraction process of the peak-finding algorithm;
[0047] Figure 8 Construct a process graph for a random forest classifier;
[0048] Fig. 9 Flowchart of the application of motion feature extraction for YOLO target detection in Parkinson's disease rehabilitation assessment;
[0049] Fig.10 This is the structure diagram of the YOLO neural network model. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0051] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0052] like Figure 1 As shown, the present invention provides a method for extracting features of a video stream signal, comprising:
[0053] Step 101: Obtain video data to be extracted.
[0054] Step 102: Input the video data to be extracted into the target detection model to obtain multiple frames of leg and foot information; the leg and foot information includes leg and foot positions, leg and foot labels, and leg and foot confidence; the target detection model is a YOLO neural network.
[0055] Step 103: Delete the leg and foot information of multiple frames according to the leg and foot confidence to determine the motion trajectory. Step 103 specifically includes:
[0056] The leg and foot information of multiple frames is deleted according to the leg and foot confidence to determine the remaining leg and foot information. The leg and foot information whose leg and foot confidence is less than a set threshold is deleted.
[0057] The Euclidean distance is calculated based on the remaining leg and foot information of adjacent frames.
[0058] A plurality of category sequences are determined according to the Euclidean distance. When the Euclidean distance is greater than a set Euclidean distance threshold, it indicates that the sequence corresponding to the Euclidean distance is a new category.
[0059] The motion trajectory is determined according to the lengths of the plurality of category sequences; the motion trajectory is the longest category sequence in the same category sequence.
[0060] Step 104: obtaining motion features using a peak-finding algorithm according to the motion trajectory and the leg and foot positions, wherein the motion features include speed and ratio; the ratio is the ratio of amplitude to leg length.
[0061] Step 104 specifically includes:
[0062] The peak position and the number of peaks are obtained by using a peak-finding algorithm according to the motion trajectory.
[0063] The speed is determined based on the peak position, the peak number and the movement time.
[0064] The ratio is determined according to the leg and foot positions in the leg and foot information.
[0065] In practical applications, the training process of the target detection model specifically includes:
[0066] The YOLO neural network is trained with the labeled video data of the training set as input and the leg and foot positions, leg and foot labels and leg and foot confidence of the training set as output to obtain the target detection model.
[0067] The present invention also provides a video stream signal feature extraction system, comprising:
[0068] The acquisition module is used to acquire the video data to be extracted.
[0069] The leg and foot information determination module is used to input the video data to be extracted into the target detection model to obtain multiple frames of leg and foot information; the leg and foot information includes leg and foot positions, leg and foot labels and leg and foot confidence; the target detection model is a YOLO neural network.
[0070] A deletion module is used to delete the leg and foot information of multiple frames according to the leg and foot confidence to determine the motion trajectory.
[0071] The motion feature determination module is used to obtain the motion feature using a peak-finding algorithm according to the motion trajectory and the leg and foot positions, wherein the motion feature includes speed and ratio; the ratio is the ratio of amplitude to leg length.
[0072] In practical applications, the training process of the target detection model specifically includes:
[0073] The YOLO neural network is trained with the labeled video data of the training set as input and the leg and foot positions, leg and foot labels and leg and foot confidence of the training set as output to obtain the target detection model.
[0074] In practical applications, the deletion module specifically includes:
[0075] A deleting unit is used to delete the leg and foot information of multiple frames according to the leg and foot confidence level to determine the remaining leg and foot information.
[0076] A calculation unit is used to calculate the Euclidean distance based on the remaining leg and foot information of adjacent frames.
[0077] The category sequence determination unit is used to determine a plurality of category sequences according to the Euclidean distance.
[0078] The motion trajectory determining unit is used to determine the motion trajectory according to the lengths of the plurality of category sequences; the motion trajectory is the category sequence with the longest length in the same category sequence.
[0079] In practical applications, the motion feature determination module specifically includes:
[0080] The peak position and peak number determination unit is used to obtain the peak position and peak number according to the motion trajectory using a peak search algorithm.
[0081] The speed determination unit is used to determine the speed according to the peak position, the peak number and the movement time.
[0082] A ratio determination unit is used to determine the ratio according to the leg and foot positions in the leg and foot information.
[0083] In order to quickly and accurately extract relevant motion features in video stream signals, a method for extracting motion features in video stream signals based on the YOLO target detection algorithm is proposed. The method first obtains motion position or trajectory information based on the target detection model, and then introduces a series of strategies to extract features based on the target position or trajectory, so as to perform the desired analysis, classification and other tasks. The method has great application capabilities, such as rehabilitation assessment of Parkinson's disease, assessment of stroke and other scenarios. The present invention uses the results obtained by target detection, such as position change and other information, and extracts relevant motion features such as amplitude, number, speed and other information after processing.
[0084] like Figure 2 As shown, this is a schematic diagram of the UPDRS (Universal Parkinson's Disease Rating Scale) motor function assessment of the flexibility of the legs in the knee flexed state adopted by the present invention. The test subject lifts his feet about ten centimeters while sitting and taps the ground with his heels. The UPDRS scoring standards are divided into five levels: normal, slow frequency and small amplitude, obvious obstacle, severe obstacle, and almost unable to complete. In this experiment, professional doctors scored the test subjects' performance and divided them into three categories: 0 for complete recovery, 1 for general recovery, and 2 for no recovery. The YOLO target detection test results of a test subject's leg flexibility in the knee flexed state are shown below. Figure 3 shown.
[0085] In order to achieve end-to-end rapid detection, the steps of the present invention are as follows:
[0086] 1. Randomly select 80% of the original video data as the training set, and the rest as the test set. Randomly select two frames from all videos, label the legs and feet, and save the category label, confidence α, the normalized upper left corner coordinates x and y of the box, and the normalized length and width w and h of the box. Use the YOLO neural network to train the leg and foot labels and corresponding pictures to obtain the YOLO target detection model. The YOLO neural network model structure is as follows: Fig.10 As shown. In the test set, the trained YOLO model is used to detect the location of the test subject's legs and feet. The category label and the corresponding α, x, y, w, h are obtained. It should be noted that ideally, two legs and two feet can be detected, but in some cases, due to some reasons, such as the front leg blocking the rear moving leg in the side view, it cannot be detected or because there are other interfering people in the background besides the test subject, so that the number of legs and feet detected is not equal to 2. Save the coordinates and corresponding confidences of all the leg and foot categories of the target detection in the format of [label, α, x, y, w, h].
[0087] 2. Perform abnormal screening on the coordinate data of the legs and feet, that is, remove the category coordinate information with a confidence level lower than 0.6. For the remaining coordinate data, connect the coordinates of each frame by frame classification. The coordinate format is x, y, w, h, which respectively represent the x and y values of the upper left corner of the frame normalized with the upper left corner of the whole image as the origin and the length w and width h of the frame. For the x, y, w, h coordinates of the same category in each next frame, calculate the four-dimensional Euclidean distance. Connect the points with the closest distance in adjacent frames and set the upper limit of the distance. When the upper limit is exceeded, it means that the box is already another individual of the same category. A new sequence needs to be created and saved. The Euclidean distance between the box coordinates of the later frames and the last coordinates of all sequences of the same category is calculated, and the closest ones are connected to construct a category sequence. Taking a sequence with a leg and a foot as an example, it is represented as:
[0088] L1=[(X 11 ,Y 11 ),(X 12 ,Y 12 )…(X 1m ,Y 1m ))]
[0089] S1=[(x 11 ,y 11 ),(x 12 ,y 12 )…(x 1n ,y 1n ))]
[0090] Wherein, L represents the target detection leg coordinate sequence, S represents the foot coordinate sequence, the subscripts of L and S represent the number of leg and foot detections, X, Y and x, y represent the upper left corner coordinates of the normalized back frame of the leg and foot, respectively, and m, n are the initial retained frame numbers of the leg and foot, respectively.
[0091] The sequences with larger amplitude ratio and longer length in the leg and foot sequences are selected as the leg and foot motion trajectories of the video. The specific operation of extracting features from the coordinate data and taking the speed v and amplitude h features as an example is as follows: for example, the peak search algorithm is used to obtain the positions of all peaks on the foot motion trajectory, and then the peak number peaks is accumulated. The peak number represents the number of times a complete leg movement is completed. The speed v is obtained by dividing the peak number peaks by the movement duration T, and the amplitude and leg length ratio h is obtained by dividing the minimum range h1 of 96% of the data containing the foot frame motion coordinate y by the mean h2 of 96% of the data containing the leg frame h.
[0092] In practical applications, we can also use features such as speed v, amplitude h and the corresponding category labels of complete recovery 0, general recovery 1, and non-recovery 2 to perform classification prediction through a classifier. The structure of the classifier is as follows: Figure 8As shown in the figure, the random forest method is used for classification training. The random forest is composed of many decision trees. The input sample needs to be input into each tree for classification. Each decision tree is a classifier. For an input sample, N trees will have N classification results. The random forest integrates all classification voting results and designates the category with the most votes as the final output. In other words, a trained classifier model is finally obtained that can input corresponding features such as v and h to obtain classification results.
[0093]
[0094]
[0095] 3. For new videos, this method is used to automatically extract features to obtain new values of features such as v and h, and then input into the trained classifier model to obtain the prediction results. Experiments have shown that this method has good generalization capabilities for other video data and can quickly and accurately obtain the final prediction results. It proves that the relevant motion feature information extracted by YOLO target detection is more targeted and more effective than the commonly used convolutional neural network method. This method can extract motion features from video stream information more quickly and accurately.
[0096] This method transforms the large-scale rehabilitation assessment of Parkinson's disease, which requires the assistance of professional doctors in hospitals, into a video classification problem. It extracts relevant motion features through the YOLO target detection algorithm for solution, avoiding the traditional end-to-end method, which extracts inaccurate video stream motion features, has poor results, and fails to capture the desired feature information. At the same time, a series of strategies are proposed based on the position information of target detection to ensure the accuracy and speed of this method, so as to achieve the effect of extracting the desired motion feature information such as amplitude and speed. Experiments have shown that the use of this method can better, efficiently and accurately assess the rehabilitation status of Parkinson's patients. Compared with the existing manual judgment method, the improved method has the advantages of speed and convenience, and can be widely used in hospitals, nursing homes or elderly people at home to conduct rehabilitation assessments online.
[0097] Figure 2 This is a schematic diagram of the application of YOLO target detection to extract relevant motion features in the rehabilitation assessment of Parkinson's disease. It is a video stream of a tester's leg flexibility assessment with the knee flexed. The tester is required to face the camera in the center of the camera screen, frontally or at a 45-degree angle, sit with his feet raised about ten centimeters, and tap the ground with his heels. No other people or objects are allowed to block the tester's body, especially the legs and feet, in the picture. The video size is 1080*1920. It is recommended that the test video size should be greater than or equal to this size.
[0098] First, use the trained YOLO neural network to detect the location boxes of all the legs and feet of the tester. The process of training the YOLO model is as follows: Figure 3 As shown, Fig.10 The YOLO neural network model structure diagram is shown below. Frame data with a confidence level below 0.60 is removed, and the coordinates, length and width of the frame are saved frame by frame. The target detection results are shown below: Figure 4 shown.
[0099] Secondly, the trajectory sequences of the legs and feet are screened, and the relevant motion features extracted are as follows Figure 5 As shown, the remaining motion feature data is processed, and the feature extraction and processing process is as follows Figure 6 As shown in Figure 2, the process of feature extraction using the peak-finding algorithm is as follows: Figure 7 As shown in the figure, the first-order difference of the trajectory data is 0, and the position is further differentiated. According to the image, it is judged whether the waveform is horizontal, and each horizontal wave and the corresponding peak point are obtained by grouping. Then, the peak detection distance is used to filter out invalid peak points to obtain the required peak points. The multi-dimensional features are put into the classifier training. The classifier construction process is shown in the figure Figure 8 shown.
[0100] Finally, the test results are obtained and the next step is recommended based on the results. The whole process of applying YOLO target detection to extract relevant motion features in Parkinson's disease rehabilitation assessment is as follows Fig. 9 As shown, firstly, the target detection model is trained after processing the data set, then the motion trajectory of the desired part is obtained by target detection, and then the classification model is trained through feature extraction of the trajectory, so that the newly input trajectory can be classified, and finally the prediction results are tested and opinions are given.
[0101] The extraction of relevant motion features in the video stream is solved through the YOLO target detection related algorithm, and the machine learning algorithm is introduced to ensure speed and accuracy, so as to obtain more effective desired features to facilitate more in-depth research activities.
[0102] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0103] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A method for extracting features of a video stream signal, characterized in that: include: Obtain the video data to be extracted; Inputting the video data to be extracted into the target detection model to obtain multiple frames of leg and foot information; The leg and foot information includes leg and foot positions, leg and foot labels, and leg and foot confidence; The target detection model is a YOLO neural network; Deleting the leg and foot information of multiple frames according to the leg and foot confidences to determine the motion trajectory; A motion feature is obtained by using a peak-finding algorithm according to the motion trajectory and the leg and foot positions, wherein the motion feature includes a speed and a ratio; the ratio is a ratio of the amplitude to the leg length; The sequences with larger amplitude ratio and longer length in the leg and foot sequences are selected as the leg and foot motion trajectories of the video; the coordinate data is feature extracted to obtain the speed v and amplitude h features as an example. The specific operation is as follows: for example, the peak finding algorithm is used to obtain the positions of all peaks on the foot motion trajectory, and then the number of peaks is accumulated to obtain the number of times a complete leg movement is completed. The speed v is obtained by dividing the number of peaks by the movement duration T, and the ratio of amplitude to leg length h is obtained by dividing the minimum range h1 of 96% of the data containing the foot frame motion coordinate y by the mean h2 of 96% of the data containing the leg frame h. The speed v, amplitude h features and the corresponding category labels of complete recovery 0, general recovery 1, and non-recovery 2 are used to perform classification prediction through a classifier.
2. The method for extracting features of a video stream signal according to claim 1, characterized in that: The training process of the target detection model specifically includes: The YOLO neural network is trained with the labeled video data of the training set as input and the leg and foot positions, leg and foot labels and leg and foot confidence of the training set as output to obtain the target detection model.
3. The method for extracting features of a video stream signal according to claim 1, characterized in that: Deleting the leg and foot information of multiple frames according to the leg and foot confidences to determine the motion trajectory specifically includes: Deleting the leg and foot information of multiple frames according to the leg and foot confidences to determine the remaining leg and foot information; Calculating the Euclidean distance based on the remaining leg and foot information of adjacent frames; Determine a plurality of category sequences according to the Euclidean distance; The motion trajectory is determined according to the lengths of the plurality of category sequences; the motion trajectory is the longest category sequence in the same category sequence.
4. The method for extracting features of a video stream signal according to claim 1, characterized in that: The step of obtaining the motion features by using a peak-finding algorithm according to the motion trajectory and the leg and foot positions specifically includes: According to the motion trajectory, a peak position and a peak number are obtained by using a peak-finding algorithm; determining a speed according to the peak position, the number of peaks and the movement time; The ratio is determined according to the leg and foot positions in the leg and foot information.
5. A video stream signal feature extraction system, characterized in that: include: An acquisition module, used for acquiring video data to be extracted; A leg and foot information determination module is used to input the video data to be extracted into a target detection model to obtain leg and foot information of multiple frames; the leg and foot information includes leg and foot positions, leg and foot labels, and leg and foot confidences; The target detection model is a YOLO neural network; A deletion module, used for deleting the leg and foot information of multiple frames according to the leg and foot confidence, to determine the motion trajectory; A motion feature determination module, used to obtain motion features using a peak-finding algorithm according to the motion trajectory and the leg and foot positions, wherein the motion features include speed and ratio; the ratio is the ratio of amplitude to leg length; The sequences with larger amplitude ratio and longer length in the leg and foot sequences are selected as the leg and foot motion trajectories of the video; the coordinate data is feature extracted to obtain the speed v and amplitude h features as an example. The specific operation is as follows: for example, the peak finding algorithm is used to obtain the positions of all peaks on the foot motion trajectory, and then the number of peaks is accumulated to obtain the number of times a complete leg movement is completed. The speed v is obtained by dividing the number of peaks by the movement duration T, and the ratio of amplitude to leg length h is obtained by dividing the minimum range h1 of 96% of the data containing the foot frame motion coordinate y by the mean h2 of 96% of the data containing the leg frame h. The speed v, amplitude h features and the corresponding category labels of complete recovery 0, general recovery 1, and non-recovery 2 are used to perform classification prediction through a classifier.
6. The video stream signal feature extraction system according to claim 5, characterized in that: The training process of the target detection model specifically includes: The YOLO neural network is trained with the labeled video data of the training set as input and the leg and foot positions, leg and foot labels and leg and foot confidence of the training set as output to obtain the target detection model.
7. The video stream signal feature extraction system according to claim 5, characterized in that: The deletion module specifically includes: A deleting unit, used for deleting the leg and foot information of multiple frames according to the leg and foot confidence, and determining the remaining leg and foot information; A calculation unit, used for calculating the Euclidean distance according to the remaining leg and foot information of adjacent frames; A category sequence determination unit, used for determining a plurality of category sequences according to the Euclidean distance; The motion trajectory determining unit is used to determine the motion trajectory according to the lengths of the plurality of category sequences; the motion trajectory is the category sequence with the longest length in the same category sequence.
8. The video stream signal feature extraction system according to claim 5, characterized in that: The motion feature determination module specifically includes: A peak position and peak number determination unit, used to obtain the peak position and peak number according to the motion trajectory using a peak search algorithm; a speed determination unit, configured to determine the speed according to the peak position, the number of peaks and the motion time; A ratio determination unit is used to determine the ratio according to the leg and foot positions in the leg and foot information.
Citation Information
Patent Citations
Video monitoring system and method for target detection and tracking
CN108055501A