A method, system and device for human gait analysis based on video data
By performing facial recognition and deep learning analysis on camera videos, extracting and evaluating the gait data of the elderly, the problem of difficulty in monitoring the status of the elderly in the existing technology is solved, and personalized prediction and active monitoring are achieved.
Patent Information
- Application Number
- CN202510157209.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The existing technology is difficult to effectively cover scenarios of living alone and multiple elderly people, and it is difficult to make personalized predictions based on individual health status, and it is impossible to achieve active monitoring.
By obtaining feedback videos from the camera, using facial recognition technology to determine the identity of the person, and establishing a health relationship file. The deep learning framework and Mask R-CNN model were used to segment the human body area, extract skeleton data, analyze gait characteristics and dynamic data, generate behavioral portraits, and combine physical health data for risk scores.
The analysis and risk score of human gaits are realized, personalized predictions can be made based on video data, and the purpose of active monitoring is achieved, which improves the monitoring efficiency of the elderly's status.
Smart Images

Figure CN119649469B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular, to a method, system, and device for human gait analysis based on video data. Background Art
[0002] In today's society, the number of elderly people living at home has increased. To ensure the status of the elderly at home, installing cameras at home has become one of the ways for children to ensure the status of the elderly at home.
[0003] In the prior art, traditional fall monitoring devices rely on wearable devices or simple sensors, which cannot effectively cover the scenarios of living alone and multiple elderly people, and it is difficult to combine individual health conditions for personalized prediction. In addition, the camera system can only determine the status at home by actively watching, and it is difficult to achieve the purpose of active monitoring. Summary of the Invention
[0004] To solve the above technical problems, this application provides a method, system, and device for human gait analysis based on video data, which is used to analyze the human gait according to the video data, and then confirm the physical status of the person.
[0005] The technical solutions provided in this application are described below:
[0006] The first aspect of this application provides a method for human gait analysis based on video data, including:
[0007] Obtain a feedback video from a camera;
[0008] Determine the identity of the person in the feedback video according to the face recognition technology, and establish a health association file according to the person's identity. The health association file includes the physical health data corresponding to the person's identity;
[0009] Establish a deep learning framework, and segment the feedback video through the Mask R-CNN model in the deep learning framework to determine the human body area;
[0010] Extract the key points of the human body area through a human pose estimation tool to generate skeleton data;
[0011] Perform behavior cutting on the feedback video to obtain an effective video segment as a behavior analysis sample;
[0012] Extract the gait features of the skeleton data in the behavior analysis sample through vision technology;
[0013] Obtain the dynamic data of the gait features through an optical flow algorithm;
[0014] Establish a behavior portrait according to the dynamic data;
[0015] Input the behavior portrait into a deep learning model, and analyze the behavior portrait according to the time series through the deep learning model to obtain an analysis result;
[0016] Combine the physical health data to perform a risk score on the analysis result;
[0017] When the risk score reaches the alarm threshold, execute the corresponding feedback strategy according to the risk score.
[0018] Optionally, the inputting the behavior portrait into a deep learning model, and analyzing the behavior portrait according to the time series through the deep learning model to obtain an analysis result includes:
[0019] Determine the model parameters of the deep learning model corresponding to the person's identity according to the physical health data to obtain a first stability factor;
[0020] Input the behavior data into the target deep learning model, and perform action trend analysis on the behavior portrait according to the time series through the model parameters to obtain a second stability factor;
[0021] Generate an analysis result according to the first stability factor and the second stability factor.
[0022] Optionally, the combining the physical health data to perform a risk score on the analysis result includes:
[0023] Obtain the basic risk parameters according to the physical health data;
[0024] Use the basic risk parameters as the risk score offset, and perform two-way offset on the first stability factor to obtain a safety interval;
[0025] Generate a risk score according to the positional relationship between the second stability factor and the safety interval.
[0026] Optionally, the performing behavior cutting on the feedback video to obtain an effective video segment as a behavior analysis sample includes:
[0027] Obtain a behavior template, and perform behavior marking on the feedback video according to the behavior template;
[0028] Cut the feedback video according to the behavior marking to obtain a behavior analysis sample.
[0029] Optionally, the extracting the gait features of the skeleton data in the behavior analysis sample through vision technology includes:
[0030] Perform data cleaning on the skeleton data through geometric constraints, and extract the gait features of the skeleton data, where the gait features include time domain features, frequency domain features, and spatio-temporal features.
[0031] Optionally, the dynamic data for obtaining the gait features through the optical flow algorithm includes:
[0032] Calculating the motion vectors of the skeleton key points between consecutive frames of the behavior analysis sample by the Lucas-Kanade optical flow method;
[0033] Performing trajectory reconstruction on the motion vectors, and at the same time calculating the change of the joint angles of the skeleton key points over time to obtain dynamic data.
[0034] Optionally, establishing the behavior portrait according to the dynamic data includes:
[0035] Analyzing the behavior components in the dynamic data;
[0036] Generating a sequence of the behavior components and obtaining the duration of each behavior in the behavior components to obtain behavior portrait sample data;
[0037] Reducing the dimension of the behavior portrait sample data based on the feature parameters of the deep learning model to obtain the behavior portrait.
[0038] The second aspect of the present application provides a human gait analysis system based on video data, including:
[0039] A first acquisition unit for acquiring a feedback video from a camera;
[0040] A first determination unit for determining the identity of the person in the feedback video according to the face recognition technology, and establishing a health association file according to the identity of the person, where the health association file contains the physical health data corresponding to the identity of the person;
[0041] A segmentation unit for establishing a deep learning framework and segmenting the feedback video through a Mask R-CNN model in the deep learning framework to determine the human body area;
[0042] A first extraction unit for extracting the key points of the human body area through a human pose estimation tool to generate skeleton data;
[0043] A cutting unit for performing behavior cutting on the feedback video to obtain an effective video segment as a behavior analysis sample;
[0044] A second extraction unit for extracting the gait features of the skeleton data in the behavior analysis sample through vision technology;
[0045] A second acquisition unit for obtaining the dynamic data of the gait features through the optical flow algorithm;
[0046] A building unit, configured to build a behavior portrait according to the dynamic data;
[0047] An analysis unit, configured to input the behavior portrait into a deep learning model, and analyze the behavior portrait according to a time series through the deep learning model to obtain an analysis result;
[0048] A risk scoring unit, configured to perform a risk score on the analysis result in combination with the physical health data;
[0049] An alarm unit, configured to execute a corresponding feedback strategy according to the risk score when the risk score reaches an alarm threshold.
[0050] Optionally, the analysis unit is specifically configured to:
[0051] Determine model parameters of a deep learning model corresponding to the person identity according to the physical health data to obtain a first stability factor;
[0052] Input the behavior data into a target deep learning model, and perform action trend analysis on the behavior portrait according to a time series through the model parameters to obtain a second stability factor;
[0053] Generate an analysis result according to the first stability factor and the second stability factor.
[0054] Optionally, the risk scoring unit is specifically configured to:
[0055] Obtain a basic risk parameter according to the physical health data;
[0056] Use the basic risk parameter as a risk score offset, and perform two-way offset on the first stability factor to obtain a safety interval;
[0057] Generate a risk score according to the position relationship between the second stability factor and the safety interval.
[0058] Optionally, the cutting unit is specifically configured to:
[0059] Obtain a behavior template, and perform behavior marking on the feedback video according to the behavior template;
[0060] Cut the feedback video according to the behavior marking to obtain a behavior analysis sample.
[0061] Optionally, the second extraction unit is specifically configured to:
[0062] Perform data cleaning on the skeleton data through geometric constraints, and extract gait features of the skeleton data, where the gait features include time domain features, frequency domain features, and spatio-temporal features.
[0063] Optionally, the second obtaining unit is specifically configured to:
[0064] Calculate the motion vectors of the skeleton key points between consecutive frames of the behavior analysis sample by using the Lucas-Kanade optical flow method;
[0065] Perform trajectory reconstruction on the motion vectors, and at the same time calculate the change of the joint angles of the skeleton key points over time to obtain dynamic data.
[0066] Optionally, the establishing unit is specifically configured to:
[0067] Analyze the behavior components in the dynamic data;
[0068] Generate a sequence of the behavior components and obtain the duration of each behavior in the behavior components to obtain behavior portrait sample data;
[0069] Reduce the dimension of the behavior portrait sample data based on the feature parameters of the deep learning model to obtain a behavior portrait.
[0070] A third aspect of the present application provides a human gait analysis device based on video data, and the device includes:
[0071] A processor, a memory, an input / output unit, and a bus;
[0072] The processor is connected to the memory, the input / output unit, and the bus;
[0073] The memory stores a program, and the processor calls the program to execute the method of the first aspect and any optional method in the first aspect.
[0074] A fourth aspect of the present application provides a computer-readable storage medium, and a program is stored on the computer-readable storage medium, and when the program is executed on a computer, it executes the method of the first aspect and any optional method in the first aspect.
[0075] It can be seen from the above technical solutions that the present application has the following advantages:
[0076] The present application analyzes the people in the feedback video, establishes a health association file for the identified person after determining the person, and confirms the skeleton of the person in the feedback video, so as to analyze the dynamic behavior of the person through the skeleton data, perform a risk score on the dynamic behavior of the person, and analyze the behavior state of the person in the video according to the risk score to achieve the purpose of risk warning. Description of the Drawings
[0077] To more clearly illustrate the technical solutions in this application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0078] Figure 1 It is a schematic flowchart of an embodiment of the human gait analysis method based on video data in this application;
[0079] Figure 2a It is a schematic flowchart of an embodiment of the first stage of the human gait analysis method based on video data in this application;
[0080] Figure 2b It is a schematic flowchart of an embodiment of the second stage of the human gait analysis method based on video data in this application;
[0081] Figure 3 It is a schematic structural diagram of an embodiment of the human gait analysis system based on video data in this application;
[0082] Figure 4 It is a schematic structural diagram of an embodiment of the human gait analysis device based on video data in this application. Detailed implementation manners
[0083] It should be noted that the human gait analysis method based on video data provided in this application can be applied to a terminal, a system, or a server. For example, the terminal can be a smart phone, a computer, a tablet computer, a smart TV, a smart watch, a portable computer terminal, or a fixed terminal such as a desktop computer. For the convenience of description, this application takes the terminal as the execution subject for example.
[0084] The following will clearly and completely describe the technical solutions in this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts belong to the scope of protection of this application.
[0085] Please refer to Figure 1 , this application first provides an embodiment of human gait analysis based on video data, and this embodiment includes:
[0086] S101. Obtain the feedback video from the camera;
[0087] The acquisition source of the feedback video is the video stream captured by the household cameras in families with the elderly or the cameras in nursing homes and welfare institutions. In actual situations, there is no age limit for the people included in the video stream captured by the cameras. Therefore, for all people appearing in the video stream, as long as there is a need for predicting or monitoring their physical states, subsequent analysis and processing can be carried out using the feedback video obtained through this solution. Specifically, the acquisition of the content of the video stream requires user permission, so the acquisition source of the feedback video is compliant.
[0088] In actual situations, the feedback video is the video stream in which there are people in the video content. When there are no people appearing in the video, no feedback video will be generated based on the video stream to reduce the storage pressure on the server.
[0089] S102. Determine the identities of the people in the feedback video according to the face recognition technology, and establish a health association file based on the identities of the people. The health association file contains the physical health data corresponding to the identities of the people.
[0090] The face recognition technology is a technology that uses artificial intelligence algorithms to identify and verify individual identities. The health association file contains electronic records of personal health information and is used for health management and analysis. Among them, the health association file is directly associated with the medical system or the data of the health association file is obtained through user upload. Specifically, the current cases stored in the hospital are generally electronic cases. As the patient himself / herself, the cases can be exported or called through the hospital's external system, so that the terminal can establish a health association file according to the case content and relevant physical examination parameters. At the same time, the health association file contains the physical health data of the people, and the physical health data is the content uploaded in advance according to the user's needs. Therefore, in the actual process of establishing the health association file, the corresponding physical health data can surely be obtained according to the identities of the people.
[0091] Specifically, the terminal uses face recognition algorithms (such as FaceNet or ArcFace) to detect and recognize the faces of the people in the feedback video. The recognition result is matched with the pre-established health file through face comparison to determine the identities of the people, and the health data related to the identified individuals is extracted from the health file according to the identities of the people.
[0092] S103. Establish a deep learning framework, and segment the feedback video through the Mask R-CNN model in the deep learning framework to determine the human body area.
[0093] The deep learning framework is a software platform used to build and train deep learning models, such as TensorFlow or PyTorch; the Mask R-CNN model is a deep learning model for instance segmentation, which can identify and segment multiple objects in an image.
[0094] In this embodiment, the Mask R-CNN model is used to remove the background. Therefore, the Mask R-CNN model mainly trains the content of the indoor scene and the corresponding background of the camera shooting angle to improve the background segmentation accuracy and efficiency of the Mask R-CNN model for the feedback video, thereby improving the accuracy of the human body area. After the camera is installed, except for disassembly or being affected by external forces, the shooting field of view of the camera is determined. The Mask R-CNN model can train the scene obtained by the camera to generate a complete area to achieve a more accurate background segmentation effect.
[0095] S104. Extract the key points of the human body area through a human pose estimation tool to generate skeleton data;
[0096] The human pose estimation tool is a tool for detecting human joints and limb parts, including but not limited to: OpenPose or MediaPipe, etc., which is a tool for generating skeleton data based on the joint key points of the human body area.
[0097] OpenPose is an open-source project written in C++. It was initially developed by Carnegie Mellon University. It focuses on real-time multi-person pose estimation and can detect the key points of the human body, face, and hands. When using OpenPose to generate skeleton data, the human body area is input into OpenPose in the state of video data or image data. In OpenPose, a pre-trained deep neural network (such as the CPM network) can be obtained to detect the approximate position and key points of the human body. For the case where there are multiple people in the video data, for each detected human figure, multiple convolutional neural networks are built into OpenPose, enabling OpenPose to accurately locate the key points of each person in the video. After determining the key points, OpenPose uses a top-down method to first detect the body center and then gradually refine to other key points, and outputs the skeleton key point coordinates and confidence scores of each detected person.
[0098] MediaPipe is a cross-platform framework developed by Google for building application-level multimedia processing pipelines, including pose estimation, facial recognition, gesture tracking, and more. When using MediaPipe to generate skeleton data, the human body area is input into MediaPipe in the form of video data or image data. In actual situations, MediaPipe provides a variety of solutions, such as Pose, Face, Hands, etc., to select the appropriate solution for different tasks. MediaPipe uses a series of predefined processing modules to process input data. For pose estimation, MediaPipe uses a 3D pose estimation model to detect key points of the human body. MediaPipe is optimized for real-time applications, including the use of GPU acceleration and model compression. The output includes key point coordinates, confidence scores, and optional 3D pose information.
[0099] In actual situations, both of the above solutions can generate skeleton data. Generally, the skeleton generation solution is selected according to the complexity of the scene. MediaPipe provides a simpler API and richer solutions, which are suitable for rapid integration and development. OpenPose provides more customization options, which is suitable for scenes that require deep customization.
[0100] S105, performing behavior segmentation on the feedback video to obtain valid video clips as behavior analysis samples;
[0101] Behavior cutting refers to the process of editing video clips according to specific behavior patterns. In this embodiment, the effective video clips refer to the video clips in which the coordinates of the skeleton data of the characters in the video change. The change in the skeleton data coordinates indicates that the human body of the character in the video is performing an action, that is, generating a behavior. The video clips in which the human body generates behavior can be used as behavior analysis samples of the character corresponding to the skeleton data.
[0102] The behavior analysis sample is used to analyze the behavior of a person in the video. The person is identified through face recognition in the previous step. If the face is not recognized, the person will not be analyzed. Only after it is determined that the face is the corresponding health-related file, the behavior analysis sample of the person will be obtained.
[0103] S106, extracting gait features of skeleton data in behavior analysis samples by using visual technology;
[0104] Specifically, the process of gait feature extraction includes: extracting basic gait parameters such as gait cycle, step length, and walking speed through the state of the skeleton data in the feedback video. After determining the basic gait parameters, frequency features of the gait are extracted through methods such as Fourier transform, and then combined with time and space information to extract dynamic features of the gait, such as the speed and acceleration of key points; calculating the angles between adjacent key points, such as the angles of the knee joint and hip joint; finally, comparing the gait features of the left and right limbs to extract symmetry features.
[0105] Specifically, the time and space information can be obtained through the background data of the camera itself and the time stamps of the feedback video.
[0106] S107. Obtain the dynamic data of gait features through the optical flow algorithm;
[0107] The optical flow algorithm is a computer vision technique used to estimate the motion of objects in an image. In gait analysis, the optical flow algorithm can help us estimate the motion of the skeleton key points from consecutive video frames, thereby obtaining the dynamic data of the gait.
[0108] Specifically, the optical flow algorithm obtains the dynamic features (the speed and acceleration of the skeleton data) from the gait features, and tracks the motion trajectories generated by the key points of the skeleton data moving over time according to the dynamic features, and calculates the changes in joint angles over time based on the key points of the skeleton data recognized as joint positions, such as the knee joint and hip joint, to quantify the motion parameters of the skeleton data and obtain the dynamic data.
[0109] S108. Establish a behavior portrait based on the dynamic data;
[0110] A behavior portrait refers to an individual behavior feature description constructed by analyzing the motion changes of the key points of the skeleton data over time. The behavior portrait is used to capture the spatial and temporal features of an individual in a specific behavior (such as walking, running, sitting down, etc.), and then determine the behavior patterns and habits of the person corresponding to the behavior portrait.
[0111] After determining the dynamic data of the gait features, the terminal will determine the actions to be confirmed according to the behavior template, that is, establish a behavior portrait corresponding to the actions, such as: walking, bending down, sitting down, standing up, etc. After determining the behavior template, the terminal extracts the corresponding behavior from the dynamic data according to the behavior template, and parameterizes the normal state of the behavior pattern of the task through the dynamic data, so that the subsequent terminal can analyze the state of the person in the camera through the behavior portrait.
[0112] S109. Input the behavior portrait into a deep learning model, and analyze the behavior portrait according to the time series through the deep learning model to obtain an analysis result;
[0113] The terminal sorts the generated behavior portraits in chronological order according to the time series. The purpose of sorting the behavior portraits by the time series is to analyze the behavior habits of an individual according to time, and identify the patterns and habits of the individual's behavior, such as which activities are more frequent and which behaviors occur within a specific time period. The information output from the above analysis is the behavior habit of the individual. After determining the behavior habit of the individual, the terminal will analyze the change trend of the individual over time according to the output behavior habit, and obtain an analysis result, which is used to predict the future behavior of the individual to formulate intervention measures.
[0114] S110. Perform a risk score on the analysis result in combination with the physical health data;
[0115] Specifically, the scoring criteria for the risk score are determined according to the health status of the individual. Specifically, the terminal determines the health status of the individual according to the health-related file corresponding to the individual, and offsets the risk score of the individual's health status with the disease history in the health-related file as an offset to obtain the risk scoring criteria for the currently analyzed individual, and performs a risk score on the corresponding individual's behavior in the subsequent input or real-time acquired feedback video according to this risk scoring criteria.
[0116] In actual situations, different diseases have different impacts on human behavior. The deep learning model can determine the impact of the corresponding disease on the behavior pattern of people based on a sufficient number of samples of the same disease. After the deep learning model completes the analysis of the behavior pattern of a single disease, these diseases will be vectorized as labels, and then an analysis model for determining the behavior habit of a person according to the disease state of the person will be generated. This model is used to assist in performing a risk score on the analysis result of human behavior and provides a relatively reliable comparison data for the risk score.
[0117] S111. When the risk score reaches the alarm threshold, execute the corresponding feedback strategy according to the risk score.
[0118] Specifically, the alarm threshold is a preset value. Generally, the risk score is set in the range of 0-100 points. According to different score intervals, the terminal has different feedback strategies for the risk score. When generating the risk scoring criteria for the behavior habit of the individual, the terminal generally refers to the state of the individual corresponding to this scoring criteria as the standard state. The risk score of the standard state is generally set in the range of 30-40 points. When the health status of the individual is in the recovery period (recovered or under conditioning), the standard state generally sets a relatively high value relative to the base value. When the health status of the individual is in a general state, the base value will be directly used for setting. Therefore, the risk score of the standard state is reflected as an interval for different health statuses of the individual.
[0119] When the risk score obtained in real time is higher than the standard state, it indicates that the risk value of the person's health status is relatively high. A feedback will be generated based on the percentage ratio exceeding the standard state. The feedback strategies are divided into strong feedback and weak feedback. For weak feedback, users who do not need to forcibly receive data do not need to confirm the feedback information. When the terminal generates strong feedback, users who receive the feedback information forcibly need to confirm the information to ensure that users can make corresponding strategies even if they do so.
[0120] This application analyzes the person in the feedback video, establishes a health association file for the person after identifying the person, and confirms the skeleton of the person in the feedback video, so as to analyze the dynamic behavior of the person through the skeleton data, perform a risk score on the dynamic behavior of the person, and analyze the behavior status of the person in the video according to the risk score to achieve the purpose of risk warning.
[0121] Please refer to Figure 2a and Figure 2b , another embodiment of the human gait analysis based on video data is provided in the embodiment of the present application. This embodiment includes:
[0122] S201. Obtain a feedback video from a camera;
[0123] S202. Determine the identity of the person in the feedback video according to the face recognition technology, and establish a health association file according to the person's identity. The health association file contains the physical health data corresponding to the person's identity;
[0124] S203. Establish a deep learning framework, and segment the feedback video through the Mask R-CNN model in the deep learning framework to determine the human body area;
[0125] S204. Extract the key points of the human body area through a human pose estimation tool to generate skeleton data;
[0126] Steps S201 to S204 in this embodiment are similar to steps S101 to S104 in the foregoing embodiment, and will not be elaborated here specifically.
[0127] S205. Obtain a behavior template, and perform behavior marking on the feedback video according to the behavior template;
[0128] The content of the behavior template is a predefined behavior pattern, which is used to identify and classify specific behaviors. The behavior marking is to obtain the behavior matching the behavior template from the feedback video and mark the time axis position where the behavior appears.
[0129] Specifically, the feedback video is a continuous video resource. All-day monitoring makes it extremely easy for the feedback video to contain invalid information. However, in actual situations, video files take up a large amount of memory on the server or the terminal itself. To avoid excessive occupation during subsequent data analysis, when the terminal obtains the feedback video, it will mark the action in the feedback video according to the behavior template.
[0130] S206. Cut the feedback video according to the behavior mark to obtain the behavior analysis sample.
[0131] This cutting behavior only occurs after the marking behavior. The essence of cutting is to crop the video. In actual situations, the feedback video may be synchronized in real time and has a limited retention duration. To ensure that the memory can obtain new video content, the old video data will be cleared after a sufficient duration. Therefore, after determining the mark, the terminal will immediately crop and archive the video content corresponding to the mark. These videos obtained through cropping and archiving are the behavior analysis samples.
[0132] S207. Clean the skeleton data through geometric constraints and extract the gait features of the skeleton data. The gait features include time-domain features, frequency-domain features, and spatio-temporal features.
[0133] Geometric constraints are used to constrain the key points corresponding to the skeleton data generated in step S204 according to the anatomical orientation of the human body, thereby correcting unreasonable key point positions, such as the reasonable ranges of limb lengths and joint angles.
[0134] Specifically, to perform geometric constraints on the skeleton data, it is first necessary to specify the rules of geometric constraints. In actual situations, the feedback video can generate a unit length based on the actual human data obtained, so as to determine the length ranges of the human limbs and torso according to this unit length. After completing the rule construction, determine the main joints and relative position relationships of the key points according to the human relationship, so as to determine the actual joints corresponding to each key point in the skeleton data. At the same time, perform preliminary correction on the obvious error detection results, such as removing extreme values that exceed the human body ratio.
[0135] After determining the constraint rules, the terminal will correct the key points at or beyond the limit value according to the constraint rules. For the key points close to the limit value, the terminal will perform a rationality analysis on the key points in combination with the current posture of the person. If the key points are determined to be reasonable, the key points will be retained.
[0136] After determining the reasonable key points, the terminal obtains the parameters of the specific action content in the analysis sample, so as to determine the time-domain features, frequency-domain features, and spatio-temporal features in the gait features according to the vector values of the actions (initial speed, acceleration, and offset angle of the actions) and fixed parameters (time, key point coordinate values) included in the analysis sample, and based on the obtained vector values and fixed parameters.
[0137] Specifically, the time-domain features should at least include gait cycle, gait symmetry, and gait variability. Among them, the gait cycle is the time length and travel distance of each step when a person walks in the measured skeleton data; gait symmetry is to compare the gait cycles of the left and right legs to evaluate the symmetry of the gait; gait variability is to calculate the variability of consecutive gait cycles to reflect the stability of the gait.
[0138] The frequency-domain features are obtained by performing a Fourier transform on the gait signal generated from the waveform corresponding to the gait of the analysis sample. The terminal performs a Fourier transform on the gait cycle signal, extracts the frequency components, and after determining the frequency, the terminal calculates the power spectral density of the analyzed gait signal to identify the main frequency components.
[0139] The spatio-temporal features are features determined by calculating the key-point trajectory analysis, joint angle changes, and speed and acceleration changes of the analysis sample. That is, through spatio-temporal feature analysis, the key-point trajectory, key-point angle, and dynamic features of the speed and acceleration of the key points during movement of the skeleton data can be obtained.
[0140] The data set containing the above three features is called gait features.
[0141] S208. Calculate the motion vectors of the skeleton key points between consecutive frames of the behavior analysis sample by the Lucas-Kanade optical flow method;
[0142] The Lucas-Kanade optical flow method is a classic optical flow estimation technique used to estimate the motion of feature points in an image sequence. The Lucas-Kanade optical flow method estimates the motion vectors of the key points by assuming that the motion of the feature points is consistent within a small neighborhood and solving a system of linear equations.
[0143] Specifically, to calculate the motion vectors, it is first necessary to determine the feature points to be calculated. These feature points are the feature points included in the skeleton data, and the number of these feature points can be multiple. For each determined feature point, an optical flow equation is established according to the Lucas-Kanade optical flow method, and the optical flow equation is solved to obtain the motion vectors of each key point between consecutive frames.
[0144] After calculating the motion vectors of the key points, to ensure the reliability of the data, after determining the vectors, the terminal will optimize the motion vectors. The specific optimization process will be carried out through vector smoothing and vector correction. Among them, vector smoothing is used to reduce the noise and outliers in the motion vectors; vector correction is used to screen out unreasonable motion vectors.
[0145] The finally obtained data is the motion vectors of the analysis sample.
[0146] S209. Reconstruct the trajectory of the motion vector, and at the same time calculate the change of the joint angle of the skeleton key points over time to obtain dynamic data.
[0147] Trajectory reconstruction refers to reconstructing the motion trajectory of the skeleton key points in the video sequence based on the motion vector.
[0148] Specifically, first, after the terminal determines the key points to be analyzed, it initializes an empty trajectory list based on the key points. At each time step, it updates the position of the key points according to the motion vector and adds them to the trajectory list to obtain the motion trajectory of the key points.
[0149] After determining the motion trajectory of the key points, the terminal will smooth the motion trajectory to reduce the influence of noise and outliers. Common methods include moving average filtering or Kalman filtering.
[0150] After completing the data smoothing of the recorded motion trajectory, the terminal will reconstruct the human skeleton corresponding to the key points according to the obtained motion trajectory, obtain the angle change between adjacent key points, and generate the coordinates of the key points on each unit frame at the same time. Finally, the dynamic data of the analysis sample corresponding to the key points is determined through coordinate calculation and angle calculation.
[0151] S210. Analyze the behavior components in the dynamic data;
[0152] Behavior components refer to the basic action units that make up complex behaviors. In human behavior analysis, identifying and classifying these basic action units is the premise for understanding complex behavior patterns. Behavior components are mainly divided into upper limb action behaviors, lower limb action behaviors, and body posture states. After the terminal obtains the skeleton data, it determines the key points for determining upper limb actions and lower limb actions according to the positions of the key points of the skeleton data, and disassembles the dynamic data according to the position states of these key points. After determining the upper limb actions and lower limb actions, it corrects the upper limb actions and lower limb actions according to the current body posture of the skeleton data to disassemble the actions in detail.
[0153] In this embodiment, it is mainly to analyze the human gait. Therefore, the unit actions for disassembling complex actions that the terminal needs to determine include but are not limited to stepping, standing, lifting the foot, heel touching the ground, and toe leaving the ground. The above unit actions can simply form a complete gait, so that the terminal can perform in-depth analysis on the human gait subsequently, and then determine the behavior habits of the person.
[0154] S211. Generate a sequence of behavior components and obtain the duration of each behavior in the behavior components to obtain behavior portrait sample data;
[0155] After determining the behavioral components, the terminal inputs the behavioral components into a convolutional network to classify these behaviors. At the same time, a sequence is generated with the time stamp within the interval from the start time to the end time of the action execution, and the actual lengths of the action components are sorted one by one on the interval to form a behavioral component sequence.
[0156] After determining the behavioral component sequence, the terminal records the duration of each behavioral component and integrates the durations of each behavioral component to integrate the behavioral component sequence and duration data into a complete data set to form behavioral portrait sample data. The behavioral portrait sample data not only includes the types and orders of the behavioral components, but also contains the duration of each behavior, thus providing a multi-dimensional behavioral description.
[0157] S212. Reduce the dimension of the behavioral portrait sample data based on the characteristic parameters of the deep learning model to obtain a behavioral portrait.
[0158] Reducing the dimension of the data is to reduce the data dimension. Specifically, the terminal determines the parameters that need to be retained from the behavioral portrait sample data according to the preset characteristic parameters in the deep learning model. In the actual usage state, the deep learning model is a learning model that reaches reasonable convergence after being trained with a sufficient amount of samples. The deep learning model can analyze the input parameters. Therefore, in the actual situation, the terminal needs to determine the input parameters for the deep learning model from the behavioral portrait sample data. These input parameters for inputting into the deep learning model are the necessary data for the behavioral portrait. In the actual situation, when there is a need to input the behavioral portrait into different deep learning models, the characteristic parameters of all deep learning models will be integrated to retain the data in the behavioral portrait sample data.
[0159] S213. Determine the model parameters of the deep learning model corresponding to the person's identity based on the physical health data to obtain the first stability factor;
[0160] The first stability factor is obtained by the deep learning model through targeted training on all the feedback videos and physical health data corresponding to this person that can be obtained from the person's health-related file for the basic risk scoring standard. Generally, after being generated, the first stability factor will be stored in the physical health data and is a data that the user cannot actively modify. The first stability factor is the health status of the individual person determined by the physical health data. Among them, the first stability factor is a person's health score generated by a health data template with the current physical condition (such as medical history or chronic diseases) as a reference factor.
[0161] S214. Input the behavioral data into the target deep learning model, and perform action trend analysis on the behavioral portrait according to the time series through the model parameters to obtain the second stability factor;
[0162] The second stability factor is the health status corresponding to the latest behavioral data of an individual generated based on input data. The terminal inputs the behavioral data obtained through preprocessing into the target deep learning model. The target deep learning model is a model trained and optimized according to the model parameters of the deep learning model corresponding to the person's identity, and can accurately analyze the behavioral characteristics of a specific person. The model performs action trend analysis on the behavioral portrait according to the time series, that is, analyzes the behavioral patterns and change trends of the person in different time periods. The deep learning model analyzes the input data to capture the behavioral stability of the person in the current health status, thereby obtaining the second stability factor. The second stability factor reflects the behavioral stability of the person in the current health status and provides an important basis for subsequent risk scoring.
[0163] Among them, the second stability factor can determine the current living status of an individual through gait data, and at the same time adjusts the second stability factor with behavioral stability (gait quality) as an influencing factor to improve the reliability of the second stability factor.
[0164] S215. Generate an analysis result based on the first stability factor and the second stability factor.
[0165] Specifically, the actual output values of the first stability factor and the second stability factor are specific numerical values, that is, the values for scoring the health status of an individual. The purpose of the deep learning model to analyze the person's gait is to assist in adjusting the person's health status according to the person's behavior. Therefore, when the terminal obtains the second stability factor, it calculates and generates an analysis result by combining the second stability factor with the first stability factor according to a preset rule.
[0166] Specifically, the preset rule can be the difference between the first stability factor and the second stability factor or the parameter change amount of the first stability factor and the second stability factor, and specific details are not limited here.
[0167] S216. Obtain the basic risk parameter according to the physical health data;
[0168] The basic risk parameter is the offset value of the initial interval obtained from a healthy individual, that is, the safety interval of a healthy individual is an interval obtained by offsetting the first stability factor as the middle quantity by the basic risk parameter to the left and right.
[0169] Among them, a healthy individual needs to meet the conditions of having no medical history and the weight of the influence of the behavior portrait on the first stability factor being less than a preset value. Specifically, the first stability factor is obtained through training historical data. In actual situations, the behavior portrait will enhance the influence weight of the analysis result of the historical data's behavior portrait on the first stability factor according to the impact of physical defects or physical traumas on the individual's behavior. Therefore, when this weight value is too high, it indicates that although the individual has no disease history, there are traumas or discomforts, and thus does not belong to a healthy individual.
[0170] S217. Use the basic risk parameter as the risk score offset to perform a two-way offset on the first stability factor to obtain a safety interval.
[0171] After obtaining the basic risk parameter, use it as the offset of the risk score to perform a two-way offset on the first stability factor, thereby obtaining a safety interval. The two-way offset means that on the basis of the first stability factor, a certain value is offset upward and downward respectively to form a safety interval. The range of this safety interval will change according to the size of the basic risk parameter. For example, if the basic risk parameter is relatively high, then the range of the safety interval will be relatively small, indicating that there may be risks when the individual's behavior changes within a relatively small range; if the basic risk parameter is relatively low, then the range of the safety interval will be relatively large, indicating that the individual is still in a safe state within a relatively large range of behavior changes. The determination of the safety interval provides a clear reference standard for subsequent risk scoring.
[0172] S218. Generate a risk score according to the positional relationship between the second stability factor and the safety interval.
[0173] After determining the safety interval, generate a risk score according to the positional relationship between the second stability factor and the safety interval. If the second stability factor falls within the safety interval, then the risk score will be relatively low, indicating that the individual's behavior is in a safe state; if the second stability factor falls outside the safety interval, then the risk score will be relatively high, indicating that there are certain risks in the individual's behavior. In addition, the risk score can be further refined according to the distance between the second stability factor and the safety interval. For example, the farther away from the safety interval, the higher the risk score. Thus, the behavior risk of the individual can be evaluated more accurately, providing a reliable basis for subsequent feedback strategies.
[0174] S219. When the risk score reaches the alarm threshold, execute the corresponding feedback strategy according to the risk score.
[0175] Step S219 in this embodiment is similar to step S211 in the foregoing embodiment, and details are not described herein again.
[0176] In this embodiment, through detailed action analysis of the behaviors of the people in the video, combined with the joint analysis of multiple models, in a multi-modal environment, behavior prediction is performed on the skeleton data of the people's gambling wins to achieve the purpose of monitoring behaviors. At the same time, when abnormal behaviors occur, prompts for the abnormal behaviors are given according to the preset states, so as to achieve the situation where an alarm can be actively issued when the camera monitors that the task has fallen, and the trigger state of the fall can be prompted through the people's behaviors, so as to reduce the situation where no one discovers the fall.
[0177] The above has described in detail the method for human gait analysis based on video data in the embodiments of the present application. Next, the human gait analysis system and device based on video data will be described in detail.
[0178] Please refer to Figure 3 , an embodiment of the human gait analysis system based on video data is provided in the embodiments of the present application. This embodiment includes:
[0179] The first acquisition unit 301 is used to acquire the feedback video from the camera;
[0180] The first determination unit 302 is used to determine the identity of the people in the feedback video according to the face recognition technology, and establish a health association file according to the identity of the people. The health association file contains the physical health data corresponding to the identity of the people;
[0181] The segmentation unit 303 is used to establish a deep learning framework, and segment the feedback video through the Mask R-CNN model in the deep learning framework to determine the human body area;
[0182] The first extraction unit 304 is used to extract the key points of the human body area through the human pose estimation tool to generate skeleton data;
[0183] The cutting unit 305 is used to perform behavior cutting on the feedback video to obtain an effective video segment as a behavior analysis sample;
[0184] The second extraction unit 306 is used to extract the gait features of the skeleton data in the behavior analysis sample through vision technology;
[0185] The second acquisition unit 307 is used to obtain the dynamic data of the gait features through the optical flow algorithm;
[0186] The establishment unit 308 is used to establish a behavior portrait according to the dynamic data;
[0187] The analysis unit 309 is used to input the behavior portrait into the deep learning model, and analyze the behavior portrait according to the time series through the deep learning model to obtain the analysis result;
[0188] A risk scoring unit 310, configured to perform a risk score on the analysis result in combination with physical health data;
[0189] An alarm unit 311, configured to execute a corresponding feedback strategy according to the risk score when the risk score reaches an alarm threshold.
[0190] In this embodiment, the analysis unit 309 is specifically configured to:
[0191] Determine model parameters of a deep learning model corresponding to a person's identity based on physical health data to obtain a first stability factor;
[0192] Input the behavior data into the target deep learning model, and perform an action trend analysis on the behavior portrait according to the time series through the model parameters to obtain a second stability factor;
[0193] Generate an analysis result based on the first stability factor and the second stability factor.
[0194] In this embodiment, the risk scoring unit 310 is specifically configured to:
[0195] Obtain a basic risk parameter according to the physical health data;
[0196] Use the basic risk parameter as a risk score offset to perform a two-way offset on the first stability factor to obtain a safety interval;
[0197] Generate a risk score according to the positional relationship between the second stability factor and the safety interval.
[0198] In this embodiment, the cutting unit 305 is specifically configured to:
[0199] Obtain a behavior template, and perform behavior marking on the feedback video according to the behavior template;
[0200] Cut the feedback video according to the behavior marking to obtain a behavior analysis sample.
[0201] In this embodiment, the second extraction unit 306 is specifically configured to:
[0202] Perform data cleaning on the skeleton data through geometric constraints, and extract gait features of the skeleton data, where the gait features include time domain features, frequency domain features, and spatio-temporal features.
[0203] In this embodiment, the second acquisition unit 307 is specifically configured to:
[0204] Calculate the motion vectors of the skeleton key points between consecutive frames of the behavior analysis sample through the Lucas-Kanade optical flow method;
[0205] Perform trajectory reconstruction on the motion vectors, and at the same time calculate the change of the joint angles of the skeleton key points over time to obtain dynamic data.
[0206] In this embodiment, the establishment unit 308 is specifically configured to:
[0207] Analyze the behavior components in the dynamic data;
[0208] Generate a sequence of behavior components and obtain the duration of each behavior in the behavior components to obtain behavior portrait sample data;
[0209] Reduce the dimension of the behavior portrait sample data based on the feature parameters of the deep learning model to obtain a behavior portrait.
[0210] In this embodiment, the functions of each unit correspond to the steps in the foregoing Figure 1 , Figure 2a , Figure 2b The corresponding steps in the illustrated embodiments will not be described in detail here.
[0211] Please refer to Figure 4 , another embodiment provided by the embodiment of the present application includes:
[0212] A processor 401, a memory 402, an input / output unit 403, and a bus 404;
[0213] The processor 401 is connected to the memory 402, the input / output unit 403, and the bus 404;
[0214] The processor 401 specifically executes the operations corresponding to the steps in Figure 1 , Figure 2a , Figure 2b The corresponding operations in the method will not be described in detail here.
[0215] The present application also relates to a computer-readable storage medium on which a program is stored. When the program runs on a computer, the computer is caused to execute any of the foregoing methods.
[0216] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.
[0217] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0218] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0219] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0220] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
Claims
1. A human gait analysis method based on video data, characterized in that: The human body gait analysis method comprises: Get the video feed from the camera; Determine the identity of the person in the feedback video according to face recognition technology, and establish a health-related file according to the identity of the person, wherein the health-related file contains physical health data corresponding to the identity of the person; Establishing a deep learning framework, and segmenting the feedback video using a Mask R-CNN model in the deep learning framework to determine the human body area; Extract key points of the human body region by using a human posture estimation tool to generate skeleton data; Performing behavior segmentation on the feedback video to obtain valid video clips as behavior analysis samples; Extracting gait features of the skeleton data in the behavior analysis sample by visual technology, wherein the gait features include time domain features, frequency domain features, and spatiotemporal features; Acquiring dynamic data of the spatiotemporal features through an optical flow algorithm; Establishing a behavior profile according to the dynamic data, wherein the behavior profile is a description of individual behavior characteristics constructed by analyzing the movement changes of key points of the skeleton data over time; Inputting the behavior profile into a deep learning model, and analyzing the behavior profile according to a time series by the deep learning model to obtain an analysis result; Performing a risk score on the analysis result in combination with the physical health data; When the risk score reaches an alarm threshold, a corresponding feedback strategy is executed according to the risk score.
2. The human gait analysis method according to claim 1, characterized in that: The step of inputting the behavior profile into a deep learning model and analyzing the behavior profile according to a time series by the deep learning model to obtain an analysis result includes: Determine the model parameters of the deep learning model corresponding to the person's identity according to the physical health data, and obtain a first stability factor, wherein the first stability factor is a basic risk scoring standard obtained by the deep learning model through targeted training based on all feedback videos and physical health data corresponding to the person that can be obtained in the health-related file of the person; Input the behavior data into a target deep learning model, and perform action trend analysis on the behavior portrait according to a time series using the model parameters to obtain a second stability factor, wherein the second stability factor is generated based on the input data and corresponds to the health status of the individual character's latest behavior data; An analysis result is generated based on the first stabilization factor and the second stabilization factor.
3. The human gait analysis method according to claim 2, characterized in that: The step of performing risk scoring on the analysis result in combination with the physical health data comprises: Obtaining basic risk parameters according to the physical health data; Using the basic risk parameter as a risk score offset, bidirectionally offset the first stability factor to obtain a safety interval; A risk score is generated according to the positional relationship between the second stability factor and the safety interval.
4. The human gait analysis method according to claim 1, characterized in that: The step of performing behavior segmentation on the feedback video to obtain valid video segments as behavior analysis samples includes: Acquire a behavior template, and perform behavior marking on the feedback video according to the behavior template; The feedback video is cut according to the behavior marker to obtain a behavior analysis sample.
5. The human gait analysis method according to claim 4, characterized in that: The step of extracting the gait features of the skeleton data in the behavior analysis sample by using visual technology includes: The skeleton data is cleaned by geometric constraints, and gait features of the skeleton data are extracted, wherein the gait features include time domain features, frequency domain features and spatiotemporal features.
6. The human gait analysis method according to claim 5, characterized in that: The step of obtaining the dynamic data of the gait feature by using an optical flow algorithm includes: Calculate the motion vector of the skeleton key point between consecutive frames of the behavior analysis sample by Lucas-Kanade optical flow method; The motion vector is reconstructed by trajectory, and the change of the joint angle of the skeleton key point over time is calculated to obtain dynamic data. The trajectory reconstruction refers to reconstructing the motion trajectory of the skeleton key point in the video sequence according to the motion vector.
7. The human gait analysis method according to any one of claims 1 to 6, characterized in that: The establishing of a behavior profile according to the dynamic data includes: parsing behavioral components in the dynamic data; Generate a sequence of the behavior components and obtain the duration of each behavior in the behavior components to obtain behavior profile sample data, wherein the behavior profile sample data includes the type and sequence of the behavior components and the duration of each behavior; The behavior portrait sample data is reduced in dimension based on the characteristic parameters of the deep learning model to obtain a behavior portrait.
8. A human gait analysis system based on video data, characterized in that: The human body gait analysis system comprises: A first acquisition unit is used to acquire a feedback video from a camera; A first determining unit is used to determine the identity of a person in the feedback video according to face recognition technology, and to establish a health-related file according to the identity of the person, wherein the health-related file contains physical health data corresponding to the identity of the person; A segmentation unit, used to establish a deep learning framework, and segment the feedback video through a Mask R-CNN model in the deep learning framework to determine a human body area; A first extraction unit, configured to extract key points of the human body region by using a human body posture estimation tool to generate skeleton data; A cutting unit, used for performing behavior cutting on the feedback video to obtain valid video clips as behavior analysis samples; A second extraction unit, configured to extract gait features of the skeleton data in the behavior analysis sample by using visual technology, wherein the gait features include time domain features, frequency domain features, and spatiotemporal features; A second acquisition unit, used for acquiring dynamic data of the spatiotemporal features through an optical flow algorithm; An establishing unit, configured to establish a behavior profile according to the dynamic data, wherein the behavior profile is a description of individual behavior characteristics constructed by analyzing the movement changes of key points of the skeleton data over time; An analysis unit, configured to input the behavior profile into a deep learning model, and analyze the behavior profile according to a time series through the deep learning model to obtain an analysis result; a risk scoring unit, configured to perform risk scoring on the analysis result in combination with the physical health data; An alarm unit is used to execute a corresponding feedback strategy according to the risk score when the risk score reaches an alarm threshold.
9. A human gait analysis device based on video data, characterized in that: The human gait analysis device comprises: Processor, memory, input-output unit, and bus; The processor is connected to the memory, the input and output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, and when the program is executed on a computer, the method according to any one of claims 1 to 7 is performed.
Citation Information
Patent Citations
Method and device for building human body 3D feature identity information database
CN106599785A
Monitoring method, device and system, recognition method and apparatus
CN114049681A